Image encoding / decoding method and device, and recording medium on which bitstream is stored
The method improves motion vector correction in video encoding/decoding by employing block-based and sub-block-based techniques, enhancing accuracy and reducing complexity for high-resolution video processing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- LG ELECTRONICS INC
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-23
AI Technical Summary
Existing video compression technologies face challenges in accurately correcting motion vectors, leading to inefficiencies in encoding and decoding high-resolution, high-quality video content.
A method and apparatus for correcting motion vectors using block-based and sub-block-based motion vector correction techniques, including bilateral matching (BM) and bi-directional optical flow (BDOF), with adaptive interweaved subblock-based corrections to reduce complexity and improve accuracy.
Enhances motion vector accuracy and reduces encoding/decoding complexity by adaptively applying motion vector corrections, minimizing information requirements and streamlining pipeline stages.
Smart Images

Figure KR2025016357_23042026_PF_FP_ABST
Abstract
Description
Video encoding / decoding method and device, and a recording medium storing a bitstream
[0001] The present invention relates to a video encoding / decoding method and apparatus, and a recording medium storing a bitstream.
[0002] Recently, the demand for high-resolution, high-quality video, such as HD (High Definition) and UHD (Ultra High Definition) video, has been increasing across various application fields, and accordingly, high-efficiency video compression technologies are being discussed.
[0003] Various image compression technologies exist, such as inter-prediction technology that predicts pixel values in the current picture from previous or subsequent pictures, intra-prediction technology that predicts pixel values in the current picture using pixel information within the current picture, and entropy coding technology that assigns short codes to values with high frequency and long codes to values with low frequency; by utilizing these image compression technologies, image data can be effectively compressed for transmission or storage.
[0004] The present disclosure provides a motion vector correction method and apparatus.
[0005] The present disclosure provides a method and apparatus for correcting motion vectors based on mutually intersecting subblocks.
[0006] The present disclosure provides a block-based motion vector correction method and apparatus comprising one or more steps.
[0007] The present disclosure provides a subblock-based motion vector correction method and apparatus comprising one or more steps.
[0008] The present disclosure provides a signaling method and apparatus for syntactic elements for motion vector correction.
[0009] An image decoding method and apparatus according to the present disclosure may derive a motion vector for a current block, correct the motion vector for the current block based on a pre-defined method, generate a predicted block of the current block based on the corrected motion vector, and restore the current block based on the predicted block. The pre-defined method may include at least one of block-based motion vector correction or sub-block-based motion vector correction.
[0010] In the image decoding method and apparatus according to the present disclosure, the block-based motion vector correction may include at least one of BM (bilateral matching)-based motion vector correction or BDOF (bi-directional optical flow)-based motion vector correction.
[0011] In the image decoding method and apparatus according to the present disclosure, the subblock-based motion vector correction may include at least one of BM (bilateral matching)-based motion vector correction or BDOF (bi-directional optical flow)-based motion vector correction.
[0012] In the image decoding method and apparatus according to the present disclosure, the subblock for BM-based motion vector correction may be a first subblock, and the subblock for BDOF-based motion vector correction may be a second subblock.
[0013] In the image decoding method and apparatus according to the present disclosure, the block-based motion vector correction may be performed based on bilateral matching (BM), and the sub-block-based motion vector correction may be performed based on bi-directional optical flow (BDOF).
[0014] In the image decoding method and apparatus according to the present disclosure, the subblock-based motion vector correction may be performed by selectively using either BM (bilateral matching) or BDOF (bi-directional optical flow) depending on whether a predetermined condition is satisfied.
[0015] In the image decoding method and apparatus according to the present disclosure, the subblock for motion vector correction based on the subblock may be determined as either a first subblock or a second subblock depending on whether a predetermined condition is satisfied.
[0016] In the image decoding method and apparatus according to the present disclosure, the previously defined method may further include motion vector correction based on interweaved subblocks.
[0017] In the image decoding method and apparatus according to the present disclosure, the interweaved subblock-based motion vector correction can be performed in parallel or independently of the subblock-based motion vector correction.
[0018] In the image decoding method and apparatus according to the present disclosure, whether to perform motion vector correction based on the interweaved subblock may be determined based on the size of the current block.
[0019] The image encoding method and apparatus according to the present disclosure may derive a motion vector for a current block, correct the motion vector for the current block based on a predefined method, generate a prediction block of the current block based on the corrected motion vector, generate a residual block of the current block based on the prediction block, derive transformation coefficients of the current block based on the residual block, and encode residual information regarding the transformation coefficients. The predefined method may include at least one of block-based motion vector correction or sub-block-based motion vector correction.
[0020] A computer-readable digital storage medium is provided that stores encoded video / image information that causes an image decoding method to be performed by a decoding device according to the present disclosure.
[0021] A computer-readable digital storage medium is provided that stores video / image information generated according to the image encoding method according to the present disclosure.
[0022] A method and apparatus for transmitting video / image information generated according to the image encoding method according to the present disclosure are provided.
[0023] The accuracy of the motion vector can be improved through motion vector correction according to the present disclosure.
[0024] Through the interweaved subblock-based motion vector correction according to the present disclosure, step dependency of motion vector correction can be reduced and the pipeline stages of encoding / decoding can be reduced, thereby reducing the complexity of encoding / decoding.
[0025] By adaptively performing motion vector correction based on interweaved subblocks according to the present disclosure, the complexity of encoding / decoding can be reduced.
[0026] According to the present disclosure, the complexity of encoding / decoding can be reduced by adaptively utilizing the number of steps for configuring motion vector correction considering image characteristics and subblock-based motion vector correction.
[0027] According to the present disclosure, accuracy can be improved while minimizing the amount of information required by adaptively signaling syntactic elements regarding motion vector correction.
[0028] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0029] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.
[0030] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.
[0031] FIG. 4 illustrates a decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0032] FIG. 5 illustrates an example (500) in which a 16x16 block is divided into 4x4 sub-blocks and an example (510) in which a corresponding interweaved type of sub-block is configured.
[0033] FIG. 6 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0034] FIG. 7 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0035] FIG. 8 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0036] FIG. 9 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0037] The present disclosure is susceptible to various modifications and may have various embodiments; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present disclosure to specific embodiments, and it should be understood that it includes all modifications, equivalents, and substitutions that fall within the spirit and scope of the present disclosure. Similar reference numerals have been used for similar components in the description of each drawing.
[0038] Terms such as "first," "second," etc., may be used to describe various components, but said components should not be limited by said terms. Such terms are used solely for the purpose of distinguishing one component from another. For example, without departing from the scope of the present disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term "and / or" includes a combination of a plurality of related described items or any of a plurality of related described items.
[0039] When it is stated that one component is "connected" or "connected" to another component, it should be understood that while it may be directly connected or connected to that other component, there may also be other components in between. On the other hand, when it is stated that one component is "directly connected" or "directly connected" to another component, it should be understood that there are no other components in between.
[0040] The terms used in this application are used merely to describe specific embodiments and are not intended to limit the disclosure. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, terms such as “comprising” or “having” are intended to specify the presence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0041] The present disclosure relates to video / video coding. For example, the methods / embodiments disclosed herein may be applied to methods disclosed in the VVC (versatile video coding) standard. Additionally, the methods / embodiments disclosed herein may be applied to methods disclosed in the EVC (essential video coding) standard, AV1 (AOMedia Video 1) standard, AVN2 (2nd generation of audio video coding standard), or next-generation video / video coding standards (e.g., H.267 or H.268).
[0042] This specification presents various embodiments regarding video / image coding, and unless otherwise noted, said embodiments may be performed in combination with one another.
[0043] In this specification, "video" may refer to a set of images over time. "Picture" generally refers to a unit representing a single image of a specific time period, and "slice" or "tile" is a unit that constitutes a part of a picture in coding. A slice or tile may contain one or more coding tree units (CTUs). A picture may consist of one or more slices or tiles. A tile is a rectangular area composed of multiple CTUs within a specific tile column and a specific tile row of a picture. A tile column is a rectangular area of CTUs having a height equal to the height of the picture and a width specified by the syntax requirements of the picture parameter set. A tile row is a rectangular area of CTUs having a height specified by the picture parameter set and a width equal to the width of the picture. CTUs within a tile are arranged continuously according to the CTU raster scan, whereas tiles within a picture may be arranged continuously according to the tile's raster scan. A single slice may include an integer number of complete tiles or an integer number of consecutive complete CTU rows within a tile of a picture that can be exclusively contained in a single NAL unit. Meanwhile, a single picture may be divided into two or more subpictures. A subpicture may be a rectangular area of one or more slices within a picture.
[0044] A pixel, or pel, can refer to the smallest unit that constitutes a picture (or image). Additionally, the term 'sample' may be used as a counterpart to pixel. A sample generally represents a pixel or its value, and it may represent only the pixel / pixel value of the luminance (luma) component or only the pixel / pixel value of the chroma component.
[0045] A unit may represent a basic unit of image processing. A unit may include at least one of a specific area of a picture and information related to that area. A unit may include one luminance block and two chroma (e.g., cb, cr) blocks. Depending on the case, the term unit may be used interchangeably with terms such as block or area. In general, an MxN block may include samples (or sample arrays) or a set (or array) of transform coefficients consisting of M columns and N rows.
[0046] In this specification, "A or B" may mean "only A," "only B," or "both A and B." Alternatively, in this specification, "A or B" may be interpreted as "A and / or B." For example, in this specification, "A, B or C" may mean "only A," "only B," "only C," or "any combination of A, B and C."
[0047] A slash ( / ) or a comma used in this specification may mean "and / or." For example, "A / B" may mean "A and / or B." Accordingly, "A / B" may mean "only A," "only B," or "both A and B." For example, "A, B, C" may mean "A, B or C."
[0048] In this specification, "at least one of A and B" may mean "only A," "only B," or "both A and B." Additionally, in this specification, the expressions "at least one of A or B" or "at least one of A and / or B" may be interpreted as synonymous with "at least one of A and B."
[0049] Additionally, in this specification, "at least one of A, B and C" may mean "only A," "only B," "only C," or "any combination of A, B and C." Also, "at least one of A, B or C" or "at least one of A, B and / or C" may mean "at least one of A, B and C."
[0050] Additionally, parentheses used in this specification may mean "for example." Specifically, where indicated as "prediction (intra-prediction)," "intra-prediction" may be proposed as an example of "prediction." In other words, "prediction" in this specification is not limited to "intra-prediction," and "intra-prediction" may be proposed as an example of "prediction." Furthermore, even where indicated as "prediction (i.e., intra-prediction)," "intra-prediction" may be proposed as an example of "prediction."
[0051] Technical features described individually within a single drawing in this specification may be implemented individually or simultaneously.
[0052] FIG. 1 illustrates a video / image coding system according to the present disclosure.
[0053] Referring to FIG. 1, the video / image coding system may include a first device (source device) and a second device (receiving device).
[0054] A source device can transmit encoded video / image information or data in the form of a file or streaming to a receiving device via a digital storage medium or a network. The source device may include a video source, an encoding device, and a transmission unit. The receiving device may include a receiver, a decoding device, and a renderer. The encoding device may be referred to as a video / image encoding device, and the decoding device may be referred to as a video / image decoding device. A transmitter may be included in the encoding device. A receiver may be included in the decoding device. The renderer may include a display unit, and the display unit may be composed of a separate device or an external component.
[0055] A video source may acquire video / images through processes such as video / image capture, synthesis, or generation. The video source may include a video / image capture device and / or a video / image generation device. A video / image capture device may include one or more cameras, a video / image archive containing previously captured video / images, etc. A video / image generation device may include a computer, a tablet, a smartphone, etc., and may generate video / images (electronically). For example, a virtual video / image may be generated through a computer, etc., in which case the video / image capture process may be replaced by a process in which related data is generated.
[0056] The encoding device can encode input video / images. The encoding device can perform a series of procedures, such as prediction, transformation, and quantization, for compression and coding efficiency. The encoded data (encoded video / image information) can be output in the form of a bitstream.
[0057] The transmission unit can transmit encoded video / image information or data output in the form of a bitstream to the receiving unit of a receiving device in the form of a file or streaming via a digital storage medium or a network. The digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit may include elements for creating a media file through a predetermined file format and elements for transmission via a broadcasting / communication network. The receiving unit can receive / extract the bitstream and transmit it to a decoding device.
[0058] The decoding device can decode video / images by performing a series of procedures such as inverse quantization, inverse transform, and prediction corresponding to the operation of the encoding device.
[0059] The renderer can render the decoded video / image. The rendered video / image can be displayed through the display unit.
[0060] FIG. 2 shows a schematic block diagram of an encoding device to which an embodiment of the present disclosure can be applied and to which encoding of a video / image signal is performed.
[0061] Referring to FIG. 2, the encoding device (200) may be configured to include an image partitioner (210), a predictor (220), a residual processor (230), an entropy encoder (240), an adder (250), a filter (260), and a memory (270). The predictor (220) may include an inter-predictor (221) and an intra-predictor (222). The residual processor (230) may include a transformer (232), a quantizer (233), a dequantizer (234), and an inverse transformer (235). The residual processor (230) may further include a subtractor (231). The addition unit (250) may be referred to as a reconstructor or a reconstructed block generator. The above-described image segmentation unit (210), prediction unit (220), residual processing unit (230), entropy encoding unit (240), addition unit (250), and filtering unit (260) may be configured by one or more hardware components (e.g., an encoding device chipset or processor) according to the embodiment. Additionally, the memory (270) may include a decoded picture buffer (DPB) and may be configured by a digital storage medium. The hardware component may further include the memory (270) as an internal / external component.
[0062] The image segmentation unit (210) can divide an input image (or picture, frame) input to an encoding device (200) into one or more processing units. For example, the processing unit may be called a coding unit (CU). In this case, the coding unit may be recursively divided from a coding tree unit (CTU) or a largest coding unit (LCU) according to a QTBTTT (Quad-tree binary-tree ternary-tree) structure.
[0063] For example, a single coding unit may be divided into multiple coding units with a deeper depth based on a quad tree structure, a binary tree structure, and / or a terrestrial structure. In this case, for example, the quad tree structure may be applied first and the binary tree structure and / or terrestrial structure may be applied later. Alternatively, the binary tree structure may be applied before the quad tree structure. A coding procedure according to the present specification may be performed based on a final coding unit that is no longer divided. In this case, based on coding efficiency according to image characteristics, the maximum coding unit may be used directly as the final coding unit, or, if necessary, the coding unit may be recursively divided into coding units of a lower depth so that a coding unit of the optimal size may be used as the final coding unit. Here, the term "coding procedure" may include procedures such as prediction, transformation, and restoration described below.
[0064] As another example, the processing unit may further include a Prediction Unit (PU) or a Transform Unit (TU). In this case, the Prediction Unit and the Transform Unit may each be divided or partitioned from the aforementioned final coding unit. The Prediction Unit may be a unit for sample prediction, and the Transform Unit may be a unit for deriving transformation coefficients and / or a unit for deriving a residual signal from transformation coefficients.
[0065] The term "unit" may be used interchangeably with terms such as "block" or "area" depending on the context. In general, an MxN block may represent a set of samples or transform coefficients consisting of M columns and N rows. A sample may generally represent a pixel or a pixel value, and may represent only the pixel / pixel value of the luminance component or only the pixel / pixel value of the chroma component. A sample may be used to refer to a single picture (or image) as a term corresponding to a pixel or pel.
[0066] The encoding device (200) can generate a residual signal (residual block, residual sample array) by subtracting a prediction signal (prediction block, prediction sample array) output from an inter prediction unit (221) or an intra prediction unit (222) from an input video signal (original block, original sample array), and the generated residual signal is transmitted to a conversion unit (232). In this case, the unit that subtracts the prediction signal (prediction block, prediction sample array) from the input video signal (original block, original sample array) within the encoding device (200) may be called a subtraction unit (231).
[0067] The prediction unit (220) performs a prediction for a block to be processed (hereinafter referred to as the current block) and can generate a predicted block containing prediction samples for the current block. The prediction unit (220) can determine whether intra prediction is applied or inter prediction is applied at the current block or CU level. The prediction unit (220) can generate various information regarding the prediction, such as prediction mode information, as described below in the description of each prediction mode, and transmit it to the entropy encoding unit (240). The information regarding the prediction can be encoded by the entropy encoding unit (240) and output in the form of a bitstream.
[0068] The intra prediction unit (222) can predict the current block by referencing samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block, depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and multiple directional modes. The non-directional mode may include at least one DC mode or a planar mode. The directional mode may include 33 directional modes or 65 directional modes depending on the degree of fineness of the prediction direction. However, this is merely an example, and depending on the settings, more or fewer directional modes may be used. The intra prediction unit (222) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.
[0069] The inter prediction unit (221) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between the template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. The reference picture containing the reference block and the reference picture containing the temporal template area may be the same or different. The above temporal template area may be referred to by names such as collocated reference block, collocated CU (colCU), etc., and the reference picture containing the above temporal template area may be referred to as a collocated picture (colPic). For example, the inter prediction unit (221) may construct a list of motion information candidates based on the template areas and generate information indicating which candidate is used to derive the motion vector and / or reference picture index of the current block. Inter prediction may be performed based on various prediction modes, for example, in the case of skip mode and merge mode, the inter prediction unit (221) may use the motion information of the template area as the motion information of the current block. In the case of skip mode, unlike merge mode, a residual signal may not be transmitted.In the case of motion vector prediction (MVP) mode, the motion vector of the template area is used as a motion vector predictor, and the motion vector of the current block can be indicated by signaling the motion vector difference.
[0070] The prediction unit (220) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in screen content coding (SCC) for games. IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When the palette mode is applied, sample values within the picture can be signaled based on information regarding the palette table and palette index. The prediction signal generated through the prediction unit (220) can be used to generate a restoration signal or to generate a residual signal.
[0071] The transformation unit (232) can generate transform coefficients by applying a transformation technique to a residual signal. For example, the transformation technique may include at least one of a Discrete Cosine Transform (DCT), a Discrete Sine Transform (DST), a Karhunen-Loeve Transform (KLT), a Graph-Based Transform (GBT), or a Conditionally Non-linear Transform (CNT). Here, GBT refers to a transformation obtained from a graph when the relationship information between pixels is represented as a graph. CNT refers to a transformation obtained based on a prediction signal generated using all previously restored pixels. Additionally, the transformation process may be applied to a pixel block of the same size in a square, or to a block of variable size that is not square.
[0072] The quantization unit (233) quantizes the transformation coefficients and transmits them to the entropy encoding unit (240), and the entropy encoding unit (240) can encode the quantized signal (information regarding the quantized transformation coefficients) and output it as a bitstream. The information regarding the quantized transformation coefficients may be called residual information. The quantization unit (233) can rearrange the block-shaped quantized transformation coefficients into a one-dimensional vector form based on the coefficient scan order, and can also generate information regarding the quantized transformation coefficients based on the one-dimensional vector-shaped quantized transformation coefficients.
[0073] The entropy encoding unit (240) can perform various encoding methods such as exponential Golomb, CAVLC (context-adaptive variable length coding), CABAC (context-adaptive binary arithmetic coding), etc. The entropy encoding unit (240) may encode information required for video / image restoration (e.g., values of syntax elements, etc.) together or separately, in addition to quantized transform coefficients.
[0074] Encoded information (e.g., encoded video / image information) may be transmitted or stored in the form of a bitstream at the level of a Network Abstraction Layer (NAL) unit. The video / image information may further include information regarding various parameter sets, such as an Adaptation Parameter Set (APS), a Picture Parameter Set (PPS), a Sequence Parameter Set (SPS), or a Video Parameter Set (VPS). Additionally, the video / image information may further include general constraint information. In this specification, information and / or syntax elements transmitted / signaled from an encoding device to a decoding device may be included in the video / image information. The video / image information may be encoded through the encoding procedure described above and included in the bitstream. The bitstream may be transmitted over a network or stored on a digital storage medium. Here, the network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. A transmission unit (not shown) that transmits the signal output from the entropy encoding unit (240) and / or a storage unit (not shown) that stores it may be configured as internal / external elements of the encoding device (200), or the transmission unit may be included in the entropy encoding unit (240).
[0075] Quantized transformation coefficients output from the quantization unit (233) can be used to generate a prediction signal. For example, a residual signal (residual block or residual samples) can be restored by applying inverse quantization and inverse transformation to the quantized transformation coefficients through the inverse quantization unit (234) and the inverse transformation unit (235). An adder (250) can generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the restored residual signal to the prediction signal output from the inter-prediction unit (221) or the intra-prediction unit (222). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the predicted block can be used as the reconstructed block. The adder (250) may be called a reconstruction unit or a reconstruction block generation unit. The generated restoration signal can be used for intra prediction of the next block to be processed within the current picture, and can also be used for inter prediction of the next picture after filtering as described below. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture encoding and / or restoration process.
[0076] The filtering unit (260) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (260) can generate a modified restored picture by applying various filtering methods to the restored picture, and can store the modified restored picture in memory (270), specifically in the DPB of memory (270). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc. The filtering unit (260) can generate various information regarding filtering and transmit it to the entropy encoding unit (240). The information regarding filtering can be encoded in the entropy encoding unit (240) and output in the form of a bitstream.
[0077] The modified restored picture transmitted to the memory (270) can be used as a reference picture in the inter-prediction unit (221). Through this, when inter-prediction is applied, the encoding device can avoid prediction mismatches between the encoding device (200) and the decoding device, and can also improve encoding efficiency.
[0078] The DPB of the memory (270) can store the modified restored picture to be used as a reference picture in the inter-prediction unit (221). The memory (270) can store motion information of blocks from which motion information is derived (or encoded) within the current picture and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter-prediction unit (221) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (270) can store restoration samples of blocks restored within the current picture and transmit them to the intra-prediction unit (222).
[0079] FIG. 3 shows a schematic block diagram of a decoding device to which an embodiment of the present disclosure can be applied and to which decoding of a video / image signal is performed.
[0080] Referring to FIG. 3, the decoding device (300) may be configured to include an entropy decoder (310), a residual processor (320), a predictor (330), an adder (340), a filter (350), and a memory (360). The predictor (330) may include an inter-predictor (332) and an intra-predictor (331). The residual processor (320) may include a dequantizer (321) and an inverse transformer (321).
[0081] The aforementioned entropy decoding unit (310), residual processing unit (320), prediction unit (330), addition unit (340), and filtering unit (350) may be configured by a single hardware component (e.g., a decoding device chipset or processor) according to an embodiment. Additionally, the memory (360) may include a DPB (decoded picture buffer) and may be configured by a digital storage medium. The hardware component may further include the memory (360) as an internal / external component.
[0082] When a bitstream containing video / image information is input, the decoding device (300) can restore the image in correspondence with the process in which the video / image information is processed by the encoding device of FIG. 2. For example, the decoding device (300) can derive units / blocks based on block division information obtained from the bitstream. The decoding device (300) can perform decoding using a processing unit applied by the encoding device. Accordingly, the processing unit for decoding may be a coding unit, and the coding unit may be divided from a coding tree unit or a maximum coding unit according to a quad tree structure, a binary tree structure, and / or a binary tree structure. One or more conversion units may be derived from the coding unit. And, the restored image signal decoded and output through the decoding device (300) can be played back through a playback device.
[0083] The decoding device (300) can receive a signal output from the encoding device of FIG. 2 in the form of a bitstream, and the received signal can be decoded through an entropy decoding unit (310). For example, the entropy decoding unit (310) can parse the bitstream to derive information (e.g., video / image information) necessary for image restoration (or picture restoration). The video / image information may further include information regarding various parameter sets, such as an adaptation parameter set (APS), a picture parameter set (PPS), a sequence parameter set (SPS), or a video parameter set (VPS). Additionally, the video / image information may further include general constraint information. The decoding device can decode the picture based on information regarding the parameter sets and / or the general constraint information. The signaling / receiving information and / or syntax elements described below in this specification may be decoded through the decoding procedure and obtained from the bitstream. For example, the entropy decoding unit (310) can decode information within the bitstream based on coding methods such as exponential chord coding, CAVLC, or CABAC, and output the values of syntax elements required for image restoration and the quantized values of transformation coefficients regarding residuals. More specifically, the CABAC entropy decoding method can receive a bin corresponding to each syntax element in the bitstream, determine a context model using information on the syntax element to be decoded and decoding information of surrounding and decoding target blocks or information on symbols / bins decoded in the previous step, predict the probability of occurrence of the bin according to the determined context model, and perform arithmetic decoding of the bin to generate a symbol corresponding to the value of each syntax element.At this time, the CABAC entropy decoding method can update the context model using the decoded symbol / bin information for the context model of the next symbol / bin after determining the context model. Among the information decoded in the entropy decoding unit (310), information regarding prediction is provided to the prediction unit (inter prediction unit (332) and intra prediction unit (331)), and the residual value for which entropy decoding was performed in the entropy decoding unit (310), i.e., quantized transformation coefficients and related parameter information, can be input to the residual processing unit (320). The residual processing unit (320) can derive residual signals (residual blocks, residual samples, residual sample array). Additionally, among the information decoded in the entropy decoding unit (310), information regarding filtering can be provided to the filtering unit (350). Meanwhile, a receiving unit (not shown) that receives a signal output from an encoding device may be further configured as an internal / external element of the decoding device (300), or the receiving unit may be a component of the entropy decoding unit (310).
[0084] Meanwhile, the decoding device according to the present specification may be called a video / image / picture decoding device, and the decoding device may be divided into an information decoding device (video / image / picture information decoding device) and a sample decoding device (video / image / picture sample decoding device). The information decoding device may include the entropy decoding unit (310), and the sample decoding device may include at least one of the inverse quantization unit (321), inverse transform unit (322), adder (340), filtering unit (350), memory (360), inter prediction unit (332), and intra prediction unit (331).
[0085] In the inverse quantization unit (321), the quantized transformation coefficients can be inversely quantized to output transformation coefficients. The inverse quantization unit (321) can rearrange the quantized transformation coefficients into a two-dimensional block form. In this case, the rearrangement can be performed based on the coefficient scan order performed by the encoding device. The inverse quantization unit (321) can perform inverse quantization on the quantized transformation coefficients using quantization parameters (e.g., quantization step size information) and obtain transformation coefficients.
[0086] In the inverse conversion unit (322), the conversion coefficients are inversely converted to obtain a residual signal (residual block, residual sample array).
[0087] The prediction unit (320) can perform a prediction for the current block and generate a predicted block containing prediction samples for the current block. The prediction unit (320) can determine whether an intra prediction or an inter prediction is applied to the current block based on information regarding the prediction output from the entropy decoding unit (310), and can determine a specific intra / inter prediction mode.
[0088] The prediction unit (320) can generate a prediction signal based on various prediction methods described below. For example, the prediction unit (320) may apply intra prediction or inter prediction for prediction of a single block, and may also apply intra prediction and inter prediction simultaneously. This may be called a combined inter and intra prediction (CIIP) mode. Additionally, the prediction unit may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content video / video coding, such as in games, such as SCC (screen content coding). IBC basically performs prediction within the current picture, but it may be performed similarly to inter prediction in that it derives a reference block within the current picture. That is, IBC may utilize at least one of the inter prediction techniques described in this specification. The palette mode can be viewed as an example of intra coding or intra prediction. When palette mode is applied, information regarding the palette table and palette index can be included in the above video / image information and signaled.
[0089] The intra prediction unit (331) can predict the current block by referring to samples within the current picture. The referenced samples may be located near the current block or at a certain distance from the current block depending on the prediction mode. In intra prediction, the prediction modes may include one or more non-directional modes and a plurality of directional modes. The intra prediction unit (331) may determine the prediction mode applied to the current block by using the prediction mode applied to the template area.
[0090] The inter prediction unit (332) can derive a prediction block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. At this time, to reduce the amount of motion information transmitted in the inter prediction mode, motion information can be predicted in blocks, sub-blocks, or samples based on the correlation of motion information between a template area and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction information (L0 prediction, L1 prediction, Bi prediction, etc.). In the case of inter prediction, the template area may include a spatial template area (spatial neighboring block) existing within the current picture and a temporal template area (temporal neighboring block) existing in the reference picture. For example, the inter prediction unit (332) may construct a motion information candidate list based on the template areas and derive the motion vector and / or reference picture index of the current block based on the received candidate selection information. Inter-prediction can be performed based on various prediction modes, and information regarding the prediction may include information indicating the inter-prediction mode for the current block.
[0091] The adder (340) can generate a restoration signal (restoration picture, restoration block, restoration sample array) by adding the acquired residual signal to the prediction signal (prediction block, prediction sample array) output from the prediction unit (including the inter prediction unit (332) and / or the intra prediction unit (331)). In cases where there is no residual for the block to be processed, such as when a skip mode is applied, the prediction block can be used as the restoration block.
[0092] The addition unit (340) may be called a restoration unit or a restoration block generation unit. The generated restoration signal may be used for intra-predicting the next block to be processed within the current picture, may be output after filtering as described below, or may be used for inter-predicting the next picture. Meanwhile, LMCS (luma mapping with chroma scaling) may be applied during the picture decoding process.
[0093] The filtering unit (350) can improve subjective / objective image quality by applying filtering to the restored signal. For example, the filtering unit (350) can generate a modified restored picture by applying various filtering methods to the restored picture, and can transmit the modified restored picture to memory (360), specifically to the DPB of memory (360). The various filtering methods may include deblocking filtering, sample adaptive offset, adaptive loop filter, bilateral filter, etc.
[0094] The (modified) restored picture stored in the DPB of the memory (360) can be used as a reference picture in the inter prediction unit (332). The memory (360) can store motion information of blocks from which motion information within the current picture has been derived (or decoded) and / or motion information of blocks within the picture that have already been restored. The stored motion information can be transmitted to the inter prediction unit (332) to be used as motion information in a spatial template area or motion information in a temporal template area. The memory (360) can store restoration samples of blocks restored within the current picture and transmit them to the intra prediction unit (331).
[0095] In this specification, the embodiments described in the filtering unit (260), inter prediction unit (221), and intra prediction unit (222) of the encoding device (200) may be applied to the filtering unit (350), inter prediction unit (332), and intra prediction unit (331) of the decoding device (300) in the same or corresponding manner.
[0096] FIG. 4 illustrates an image decoding method performed by a decoding device (300) as an embodiment according to the present disclosure.
[0097] Referring to FIG. 4, a motion vector can be derived for the current block (S400).
[0098] The above motion vector can be derived based on a predetermined inter-prediction mode. The predetermined inter-prediction mode may be a merge mode or an AMVP (advanced motion vector prediction) mode. However, it is not limited thereto, and the motion vector may be understood as being replaced by a motion vector predictor.
[0099] Referring to FIG. 4, the motion vector for the current block can be corrected based on a pre-defined method (S410).
[0100] The above-defined method is for decoder side motion vector refinement (DMVR) and may include at least one of block-based motion vector refinement or subblock-based motion vector refinement.
[0101] Block-based motion vector correction is a method of correcting motion vectors on a block-by-block basis, where a block unit may refer to a block having the same width and height as the current block.
[0102] Block-based motion vector correction can be performed based on bilateral matching (BM). For example, multiple candidate vectors may be defined for the current block. Each candidate vector may include an L0 vector and an L1 vector. For each of the multiple candidate vectors, an L0 block can be obtained from an L0 reference picture list based on the L0 vector, and an L1 block can be obtained from an L1 reference picture list based on the L1 vector. The bilateral matching cost (BM cost) between the L0 block and the L1 block can be calculated. The motion vector of the current block can be corrected using the candidate vector that minimizes the bilateral matching cost. Alternatively, block-based motion vector correction may be performed based on bi-directional optical flow (BDOF).
[0103] Sub-block-based motion vector correction is a method of correcting motion vectors on a sub-block basis, wherein a sub-block may refer to a block in which at least one of the width or height is smaller than the current block.
[0104] The sub-block-based motion vector correction may include at least one of the first sub-block-based motion vector correction, the second sub-block-based motion vector correction, or the third sub-block-based motion vector correction.
[0105] Motion vector correction based on the first subblock can be performed based on bidirectional matching or BDOF. Here, the first subblock may refer to a block of a predefined size such as 4x4, 8x8, 16x16, or 32x32. Alternatively, the first subblock may refer to a block of a predefined size such as 4x8, 8x4, 8x16, 16x8, 16x32, or 32x16. Alternatively, the first subblock may be adaptively determined based on the size of the current block. For example, the first subblock may have a subblock size obtained by N-dividing the current block. N may be an integer of 2, 4, 8, or greater.
[0106] Motion vector correction based on the second subblock can be performed based on bidirectional matching or BDOF. Here, the second subblock may refer to a block of a predefined size such as 4x4, 8x8, or 16x16. Alternatively, the second subblock may refer to a block of a predefined size such as 4x8, 8x4, 8x16, or 16x8. Alternatively, the second subblock may be adaptively determined based on the size of the current block or the first subblock. For example, the second subblock may have a subblock size obtained by M-dividing the current block or the first subblock. M may be an integer of 2, 4, 8, or greater. However, at least one of the width or height of the second subblock may be smaller than that of the first subblock.
[0107] Motion vector correction based on the third subblock may be performed based on bidirectional matching or BDOF. Motion vector correction based on the third subblock may be performed adaptively depending on whether the size or shape of the current block satisfies pre-defined conditions. The size of the current block may be defined as at least one of width, height, the product of width and height, the sum of width and height, the ratio of width to height, or the minimum / maximum value of width and height. The third subblock may be a square block such as 4x4 or 8x8. Alternatively, the third subblock may be a non-square block such as 4x8, 8x4, 8x16, or 16x8. At least one of the size or shape of the third subblock may be adaptively determined based on at least one of the size or shape of the current block. Alternatively, the third subblock may be set as a subblock having the same size as the aforementioned first subblock or second subblock. Alternatively, the third sub-block may be smaller than either the first sub-block or the second sub-block.
[0108] Below, we will examine an embodiment of a DMVR composed of multiple steps.
[0109] Example 1
[0110] The motion vector of the current block can be corrected through block-based motion vector correction (S11). Here, block-based motion vector correction can be performed based on BM.
[0111] The corrected motion vector in S11 can be corrected through motion vector correction based on the first sub-block (S12). Here, motion vector correction based on the first sub-block can be performed based on BM. The first sub-block can be a 16x16 block. However, motion vector correction based on the first sub-block may not be performed depending on the size of the current block. For example, if the first sub-block is a 16x16 block, or if at least one of the width or height of the current block to which decoder-side motion vector correction is applied is smaller than 16, motion vector correction based on the first sub-block may be omitted.
[0112] The corrected motion vector in S11 or S12 can be corrected through motion vector correction based on the second sub-block (S13). Here, motion vector correction based on the second sub-block can be performed based on BM. The second sub-block can be an 8x8 block. However, motion vector correction based on the second sub-block may not be performed depending on the size of the current block. For example, if the second sub-block is an 8x8 block, or if at least one of the width or height of the current block to which decoder-side motion vector correction is applied is less than 8, motion vector correction based on the second sub-block may be omitted.
[0113] The corrected motion vector in S11, S12, or S13 can be corrected through motion vector correction based on a third sub-block (S14). Here, the motion vector correction based on a third sub-block can be performed based on BDOF. The third sub-block can be a 4x4 block or an 8x8 block. However, the motion vector correction based on a third sub-block can be performed adaptively depending on whether the size of the current block satisfies a pre-defined condition. For example, if the size of the current block is greater than a threshold value defined equally in the encoding device and the decoding device, the motion vector correction based on a third sub-block can be performed, and if not, the motion vector correction based on a third sub-block can not be performed.
[0114] The above-described embodiment corresponds to a case where decoder-side motion vector correction is performed in the order of block-based motion vector correction, first sub-block-based motion vector correction, second sub-block-based motion vector correction, and third sub-block-based motion vector correction. However, this is not limited thereto, and the third sub-block-based motion vector correction may be performed before the first sub-block-based motion vector correction and / or the second sub-block-based motion vector correction.
[0115] Example 2
[0116] Decoder side motion vector correction can be performed based on subblocks configured in an interweaved form.
[0117] FIG. 5 illustrates an example (500) of dividing a 16x16 block into 4x4 sub-blocks and an example (510) of configuring the sub-blocks in a corresponding interweaved form. Referring to FIG. 5, the interweaved division according to the present disclosure may mean dividing such that the division boundary is positioned at a location shifted by a predetermined offset in at least one of the x-axis direction or the y-axis direction so as not to overlap with the division boundary according to the existing division. At this time, the offset may vary depending on the size of the current block and / or sub-block. For example, the offset may be 1 / 2 or 1 / 4 of the width or height of the sub-block.
[0118] The present disclosure proposes a method for performing conventional subblock-based motion vector correction and interweaved subblock-based motion vector correction in parallel. Through this, the step dependency of motion vector correction can be reduced, and the complexity of encoding / decoding can be reduced by reducing the pipeline stages of encoding / decoding.
[0119] Specifically, the motion vector of the current block can be corrected through block-based motion vector correction (S21). Block-based motion vector correction is as described above.
[0120] The corrected motion vector in S21 can be corrected through sub-block-based motion vector correction (S22). The sub-block-based motion vector correction may include at least one of a first sub-block-based motion vector correction, a second sub-block-based motion vector correction, or a third sub-block-based motion vector, as previously discussed.
[0121] The corrected motion vector in S21 can be corrected through motion vector correction based on an interweaved subblock (S23). Motion vector correction based on an interweaved subblock can be performed in parallel with or independently of motion vector correction based on a subblock. Here, the interweaved subblock may have a size corresponding to the subblock used in S22.
[0122] A final corrected motion vector for the current block can be derived based on the corrected motion vector in S22 and the corrected motion vector in S23 (S24). For example, a final corrected motion vector for the current block can be derived through a blending process (e.g., average, weighted sum, etc.) between the corrected motion vector in S22 and the corrected motion vector in S23.
[0123] Example 3
[0124] The present disclosure proposes a method for performing conventional subblock-based motion vector correction and interweaved subblock-based motion vector correction in parallel, while additionally performing BDOF-based motion vector correction. Through this, the step dependency of motion vector correction can be reduced, and the complexity of encoding / decoding can be reduced by reducing the pipeline stages of encoding / decoding.
[0125] Specifically, the motion vector of the current block can be corrected through block-based motion vector correction (S31). Block-based motion vector correction is as described above.
[0126] The corrected motion vector in S31 can be corrected through at least one of the aforementioned first sub-block-based motion vector correction or second sub-block-based motion vector correction (S32).
[0127] The corrected motion vector in S31 can be corrected through motion vector correction based on an interweaved subblock (S33). Motion vector correction based on an interweaved subblock can be performed in parallel with or independently of the motion vector correction in S32. Here, the interweaved subblock may have a size corresponding to the subblock used in S32. For example, in S32, the motion vector can be corrected based on a 16x16 subblock, and correspondingly, in S33, the motion vector can be corrected based on a 16x16 interweaved subblock.
[0128] A blending process (e.g., average, weighted sum, etc.) can be performed based on the corrected motion vector in S32 and the corrected motion vector in S33 (S34).
[0129] The motion vector derived through the blending process of S34 can be corrected through motion vector correction based on the third sub-block (S35). Motion vector correction based on the third sub-block can be performed based on a method different from the motion vector correction in S33. For example, motion vector correction in S33 can be performed based on BM, and motion vector correction in S35 can be performed based on BDOF. Alternatively, motion vector correction in S33 can be performed based on BDOF, and motion vector correction in S35 can be performed based on BM.
[0130] Example 4
[0131] The present disclosure proposes a method for performing conventional subblock-based motion vector correction and interweaved subblock-based motion vector correction in parallel. Through this, the step dependency of motion vector correction can be reduced, and the complexity of encoding / decoding can be reduced by reducing the pipeline stages of encoding / decoding.
[0132] Specifically, the motion vector of the current block can be corrected through block-based motion vector correction (S41). Block-based motion vector correction is as described above.
[0133] The corrected motion vector in S41 can be corrected through motion vector correction based on the first sub-block (S42). Here, motion vector correction based on the first sub-block can be performed based on BM. The first sub-block may be smaller than the current block. For example, the first sub-block may be a 16x16 block. However, this is merely an example, and the first sub-block may be defined as a block of a different size or shape as described above.
[0134] The corrected motion vector in S42 can be corrected through motion vector correction based on the second subblock (S43). Here, motion vector correction based on the second subblock can be performed based on BDOF. The second subblock may be smaller than or equal to the first subblock used in S42. For example, the second subblock may be an 8x8 block. However, this is merely an example, and the second subblock may be defined as a block of a different size or shape as described above.
[0135] The corrected motion vector in S41 can be corrected through motion vector correction based on the first interweaved subblock (S44). Here, the motion vector correction based on the first interweaved subblock can be performed based on BM. The first interweaved subblock can have a size corresponding to the first subblock used in S42.
[0136] The corrected motion vector in S44 can be corrected through motion vector correction based on the second interweaved subblock (S45). Here, the motion vector correction based on the second interweaved subblock can be performed based on BDOF. The second interweaved subblock can have a size corresponding to the second subblock used in S43.
[0137] The motion vector correction of S44 and S45 can be performed in parallel with or independently of the motion vector correction of S42 and S43.
[0138] A final corrected motion vector for the current block can be derived based on the corrected motion vector in S43 and the corrected motion vector in S45 (S46). For example, a final corrected motion vector for the current block can be derived through a blending process (e.g., average, weighted sum, etc.) between the corrected motion vector in S43 and the corrected motion vector in S45.
[0139] Example 5
[0140] The present disclosure proposes a method for performing conventional subblock-based motion vector correction and interweaved subblock-based motion vector correction in parallel, while additionally performing BDOF-based motion vector correction. Through this, the step dependency of motion vector correction can be reduced, and the complexity of encoding / decoding can be reduced by reducing the pipeline stages of encoding / decoding.
[0141] Specifically, the motion vector of the current block can be corrected through block-based motion vector correction (S51). Block-based motion vector correction is as described above.
[0142] The corrected motion vector in S51 can be corrected through motion vector correction based on the first sub-block (S52). Here, motion vector correction based on the first sub-block can be performed based on BM. The first sub-block may be smaller than the current block. For example, the first sub-block may be a 16x16 block. However, this is merely an example, and the first sub-block may be defined as a block of a different size or shape as described above.
[0143] The corrected motion vector in S52 can be corrected through motion vector correction based on the second subblock (S53). Here, motion vector correction based on the second subblock can be performed based on BDOF. The second subblock may be smaller than or equal to the first subblock used in S52. For example, the second subblock may be an 8x8 block. However, this is merely an example, and the second subblock may be defined as a block of a different size or shape as described above.
[0144] The corrected motion vector in S51 can be corrected through motion vector correction based on the first interweaved subblock (S54). Here, the motion vector correction based on the first interweaved subblock can be performed based on BM. The first interweaved subblock can have a size corresponding to the first subblock used in S52.
[0145] The corrected motion vector in S54 can be corrected through motion vector correction based on the second interweaved subblock (S55). Here, the motion vector correction based on the second interweaved subblock can be performed based on BDOF. The second interweaved subblock can have a size corresponding to the second subblock used in S53.
[0146] The motion vector correction of S54 and S55 can be performed in parallel with or independently of the motion vector correction of S52 and S53.
[0147] A blending process (e.g., average, weighted sum, etc.) can be performed based on the corrected motion vector in S53 and the corrected motion vector in S55 (S56).
[0148] The motion vector derived through the blending process of S56 can be corrected through motion vector correction based on the third sub-block (S57). Motion vector correction based on the third sub-block can be performed based on any one of the correction methods used in motion vector correction of S52 to S55. For example, if BM and BDOF were used in motion vector correction of S52 to S55, motion vector correction based on the third sub-block can be performed based on BM or BDOF. The third sub-block may be smaller than or equal to the first sub-block and / or second sub-block used previously. For example, the third sub-block may be a 4x4 block or an 8x8 block. However, this is merely an example, and the third sub-block may be defined as a block of a different size or shape as described above.
[0149] Whether to perform motion vector correction based on the aforementioned interweaved subblock can be determined based on at least one of the size of the current block, the size of the subblock for the DMVR of the current block, or a motion vector correction method (e.g., BM, BDOF).
[0150] For example, whether motion vector correction based on interweaved subblocks is performed can be determined based on the size of the current block. Alternatively, whether motion vector correction based on interweaved subblocks is performed can be determined based on a comparison between the size of the current block and a threshold. The threshold may be pre-defined identically in both the encoding device and the decoding device. If the size of the current block is smaller than the threshold, motion vector correction based on interweaved subblocks may not be performed. In other words, if the current block is determined to be a small block, constructing interweaved subblocks to perform motion vector correction may lead to increased complexity. Therefore, motion vector correction based on interweaved subblocks may be performed only when the size of the current block is greater than or equal to the threshold.
[0151] The size of the current block can be defined as the product of the width and height of the current block. In this case, the threshold can be any one of 256, 512, 1024, or 2048. If the product of the width and height of the current block is less than the threshold, motion vector correction based on interweaved subblocks may not be performed, and otherwise, motion vector correction based on interweaved subblocks may be performed.
[0152] Alternatively, the size of the current block may be defined as the minimum value between the width and height of the current block. In this case, the threshold may be any one of 16, 32, or 64. If the minimum value between the width and height of the current block is smaller than the threshold, motion vector correction based on interweaved subblocks may not be performed, and otherwise, motion vector correction based on interweaved subblocks may be performed.
[0153] Based on the size of the subblock used in subblock-based motion vector correction, it can be determined whether interweaved subblock-based motion vector correction is performed. Interweaved subblock-based motion vector correction can be performed only when the size of the subblock for subblock-based motion vector correction corresponds to a pre-defined size.
[0154] If the subblock size for subblock-based motion vector correction is greater than or equal to 16x16, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed. Alternatively, if the subblock size is greater than or equal to 32x32, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed. Alternatively, if the subblock size is greater than or equal to 64x64, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed. Alternatively, if the minimum value between the width and height of the subblock is greater than or equal to 16, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed. Alternatively, if the minimum value between the width and height of the subblock is greater than or equal to 32, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed. Alternatively, if the minimum value between the width and height of the subblock is greater than or equal to 64, interweaved subblock-based motion vector correction may be performed, and if not, interweaved subblock-based motion vector correction may not be performed.
[0155] Based on the motion vector correction method, it may be determined whether interweaved subblock-based motion vector correction is performed. Among a plurality of motion vector correction steps, interweaved subblock-based motion vector correction may be performed only for the motion vector correction step that uses a pre-defined motion vector correction method. When the motion vector correction methods of BM and BDOF are used for the current block, interweaved subblock-based motion vector correction may be performed in response to motion vector correction using BM, and interweaved subblock-based motion vector correction may not be performed in response to motion vector correction using BDOF. Alternatively, when the motion vector correction methods of BM and BDOF are used for the current block, interweaved subblock-based motion vector correction may be performed in response to motion vector correction using BDOF, and interweaved subblock-based motion vector correction may not be performed in response to motion vector correction using BM.
[0156] Depending on the image characteristics, at least one of the number of motion vector correction steps or whether subblock-based motion vector correction is performed may be determined.
[0157] Motion vector correction can be divided into block-based motion vector correction using BM, block-based motion vector correction using BDOF, sub-block-based motion vector correction using BM, and sub-block-based motion vector correction using BDOF, depending on the application unit and correction method.
[0158] Motion vector correction with multiple stages can be configured by combining at least two of the four motion vector corrections described above. In this case, a motion vector with higher accuracy can be derived, but the complexity of the decoding / encoding unit increases due to the multiple stages performed serially, and additional latency may be induced in the motion vector derivation process. Furthermore, depending on image characteristics in addition to the size and shape of the current block, motion vector correction with multiple stages may increase unnecessary complexity of the decoding / encoding unit relative to encoding efficiency. Accordingly, the present disclosure proposes a method for determining the number of stages for configuring motion vector correction and whether sub-block-based motion vector correction is performed, taking into account image characteristics.
[0159] Example A
[0160] The motion vector of the current block can be corrected based on block-based motion vector correction without sub-block-based motion vector correction. Here, block-based motion vector correction may include BM-based motion vector correction and BDOF-based motion vector correction. In this case, block-based motion vector correction may be performed in the order of BM-based motion vector correction followed by BDOF-based motion vector correction. Alternatively, block-based motion vector correction may be performed in the order of BDOF-based motion vector correction followed by BM-based motion vector correction. In this way, the accuracy of motion vector correction can be increased by applying different correction methods while reducing the number of steps constituting motion vector correction. Alternatively, block-based motion vector correction may be performed based on a single predefined correction method. The single predefined correction method may be BM or BDOF. Alternatively, either BM or BDOF may be selectively used depending on whether a predetermined condition is satisfied. For example, if the current block satisfies a predetermined condition, block-based motion vector correction can be performed based on BM, and if not, block-based motion vector correction can be performed based on BDOF.
[0161] Example B
[0162] The motion vector of the current block can be corrected based on block-based motion vector correction and sub-block-based motion vector correction. This can be performed in the order of block-based motion vector correction followed by sub-block-based motion vector correction. Here, block-based motion vector correction can be performed based on a pre-defined correction method. The pre-defined correction method may be BM or BDOF. Sub-block-based motion vector correction may include BM-based motion vector correction and BDOF-based motion vector correction. That is, sub-block-based motion vector correction may consist of two steps using different correction methods. In this case, sub-block-based motion vector correction may be performed in the order of BM-based motion vector correction followed by BDOF-based motion vector correction. In this case, the sub-block for BM-based motion vector correction may be the aforementioned first sub-block (e.g., 16x16 block), and the sub-block for BDOF-based motion vector correction may be the aforementioned second sub-block (e.g., 8x8 block). Alternatively, sub-block-based motion vector correction may be performed in the order of BDOF-based motion vector correction and BM-based motion vector correction. In this case, the sub-block for BDOF-based motion vector correction may be the aforementioned first sub-block (e.g., 8x8 block), and the sub-block for BM-based motion vector correction may be the aforementioned second sub-block (e.g., 4x4 block, or pixel unit).
[0163] Example C
[0164] The motion vector of the current block can be corrected based on block-based motion vector correction and sub-block-based motion vector correction. This can be performed in the order of block-based motion vector correction followed by sub-block-based motion vector correction. Here, sub-block-based motion vector correction may utilize a correction method different from block-based motion vector correction (e.g., BM or BDOF). That is, sub-block-based motion vector correction may consist of a single step corresponding to BM or BDOF-based motion vector correction. For example, block-based motion vector correction may be performed based on BM, and sub-block-based motion vector correction may be performed based on BDOF. Alternatively, block-based motion vector correction may be performed based on BDOF, and sub-block-based motion vector correction may be performed based on BM. In this case, the sub-block for motion vector correction may be the aforementioned first sub-block (e.g., an 8x8 block).
[0165] Example D
[0166] The motion vector of the current block can be corrected based on block-based motion vector correction and sub-block-based motion vector correction. This can be performed in the order of block-based motion vector correction followed by sub-block-based motion vector correction. Here, block-based motion vector correction can be performed based on a pre-defined correction method. The pre-defined correction method may be BM or BDOF. Sub-block-based motion vector correction can be performed by selectively using either BM or BDOF depending on whether a predetermined condition is satisfied. That is, sub-block-based motion vector correction can be composed of a single step corresponding to BM or BDOF-based motion vector correction. For example, if the current block satisfies a predetermined condition, sub-block-based motion vector correction can be performed based on BM, and if not, sub-block-based motion vector correction can be performed based on BDOF. At this time, the sub-block for motion vector correction may be the aforementioned first sub-block (e.g., an 8x8 block).
[0167] Example E
[0168] The motion vector of the current block can be corrected based on block-based motion vector correction and sub-block-based motion vector correction. This can be performed in the order of block-based motion vector correction followed by sub-block-based motion vector correction. Here, block-based motion vector correction can be performed based on a pre-defined correction method. The pre-defined correction method may be BM or BDOF. Sub-block-based motion vector correction can be performed based on a correction method different from block-based motion vector correction. For example, block-based motion vector correction can be performed based on BM, and sub-block-based motion vector correction can be performed based on BDOF. Alternatively, block-based motion vector correction can be performed based on BDOF, and sub-block-based motion vector correction can be performed based on BM. The sub-block for motion vector correction can be determined as either a first sub-block or a second sub-block depending on whether it satisfies a predetermined condition. That is, sub-block-based motion vector correction can be composed of a single step corresponding to motion vector correction based on the first sub-block or the second sub-block. For example, if the current block satisfies a predetermined condition, motion vector correction based on the first sub-block may be performed, and if not, motion vector correction based on the second sub-block may be performed. Here, the first sub-block and the second sub-block are as described above, and redundant descriptions are omitted here.
[0169] The method for configuring multiple steps for motion vector correction is as follows.
[0170] It may include one or more of BM-based motion vector correction and BDOF-based motion vector correction. That is, it may be composed only of BM-based motion vector correction, or it may be composed only of BDOF-based motion vector correction. Alternatively, it may be composed of a mixture of BM-based motion vector correction and BDOF-based motion vector correction.
[0171] Depending on whether explicit or implicit conditions are satisfied, either BM-based motion vector correction or BDOF-based motion vector correction can be adaptively selected.
[0172] At least one of the BM-based motion vector correction or BDOF-based motion vector correction may include block-based motion vector correction and one or more sub-block-based motion vector corrections.
[0173] A subblock for motion vector correction can be defined with any size within a range that does not exceed the size of the current block. For example, the size of the subblock can be 16x16, 8x8, or 4x4.
[0174] In subblock-based motion vector correction, either BM-based motion vector correction or BDOF-based motion vector correction can be adaptively selected depending on whether explicit or implicit conditions are satisfied.
[0175] In subblock-based motion vector correction, the size of the subblock can be adaptively selected depending on whether explicit or implicit conditions are satisfied.
[0176] Various implicit conditions, such as the size of the current block, may be utilized as conditions for determining the number of motion vector correction steps and / or whether sub-block-based motion vector correction is performed. Additionally, the encoding device may derive information for determining the number of motion vector correction steps and / or whether sub-block-based motion vector correction is performed based on image characteristics, and may explicitly signal this to the decoding device.
[0177] The number of motion vector correction steps can be adjusted based on the size of the current block. The number of motion vector correction steps can be configured differently based on a comparison between the size of the current block and a threshold. The threshold may be pre-defined identically in the encoding device and the decoding device. For example, if the size of the current block is smaller than the threshold, the number of motion vector correction steps can be reduced. If the size of the current block is greater than or equal to the threshold, motion vector correction consisting of four steps can be performed as in Example 1 described above. On the other hand, if the size of the current block is smaller than the threshold, sub-block-based motion vector correction can be omitted, and block-based motion vector correction consisting of two steps can be performed, as in Example A described above. That is, if the size of the current block is small, the number of motion vector correction steps can be reduced to decrease the hardware pipeline increase that may be caused by motion vector correction consisting of multiple steps.
[0178] Alternatively, if the current block size is greater than or equal to the threshold, the number of motion vector correction steps may be increased.
[0179] Alternatively, if the current block is determined to be a small block (e.g., if the size of the current block is smaller than a threshold), the number of motion vector correction steps may be reduced, but only one method of BM-based motion vector correction or BDOF-based motion vector correction may be applied. That is, for the current block, block-based motion vector correction using BM and sub-block-based motion vector correction using BM are permitted, but block-based motion vector correction using BDOF and sub-block-based motion vector correction using BDOF may not be permitted. In this case, the size of the current block may be defined based on at least one of the width or height of the current block. For example, the values of the width and / or height of the current block themselves may be compared with the threshold, or the product of the width and height of the current block may be compared with the threshold. Even when motion vector correction consisting of two steps is applied as in Example A, if the motion vector obtained by the motion vector correction corresponding to the first step is less than or equal to a predetermined threshold, the motion vector correction corresponding to the second step may be omitted. This allows the complexity of encoding / decoding to be further reduced.
[0180] Whether subblock-based motion vector correction is performed can be determined based on the size of the current block. The decision to perform subblock-based motion vector correction can be made based on a comparison between the size of the current block and a threshold. The threshold may be a value that is pre-defined identically in both the encoding device and the decoding device. For example, if the size information of the current block is smaller than the pre-defined threshold, subblock-based motion vector correction may not be performed; otherwise, subblock-based motion vector correction may be performed. In other words, since performing subblock-based motion vector correction for small-sized blocks may unnecessarily increase complexity, subblock-based motion vector correction may be performed only when the size of the current block is greater than or equal to the threshold.
[0181] Here, the size of the current block can be defined by the width and / or height of the current block. In this case, the threshold can be 64, 32, 16, 8, or 4. The size of the current block can be defined by the product of the width and height of the current block. In this case, the threshold can be 256, 512, 1024, or 2048. The size of the current block can be defined by the sum of the width and height of the current block. In this case, the threshold can be 12, 16, 20, 24, 32, 48, or 64.
[0182] The number of motion vector correction steps and / or whether sub-block-based motion vector correction is performed may be determined based on the quantization parameter (QP) applied to the current block, the number of motion vector correction steps applied to the surrounding blocks of the current block, whether sub-block-based motion vector correction is performed in the surrounding blocks of the current block, and whether the correction method applied to the current block / surrounding blocks is BM or BDOF. The surrounding blocks may be surrounding blocks that are spatially and / or temporally adjacent to the current block.
[0183] Whether the correction method applied to the current block is BM or BDOF may be implicitly determined based on the information of the current block. For example, either BM or BDOF may be selected based on the temporal correlation between the L0 / L1 reference picture and the current picture according to the prediction direction of the current block.
[0184] Information indicating the number of motion vector correction steps, whether to perform subblock-based motion vector correction, and the correction method can be configured as follows.
[0185] Information regarding the total number of motion vector correction steps can be signaled.
[0186] Information regarding the motion vector correction method can be signaled. Information regarding the number of steps constituting the block-based motion vector correction can be signaled.
[0187] Information regarding the correction method used in block-based motion vector correction can be signaled.
[0188] Information regarding whether subblock-based motion vector correction is performed can be signaled.
[0189] Information regarding the number of steps constituting subblock-based motion vector correction can be signaled.
[0190] Information regarding the number of steps using BM and information regarding the number of steps using BDOF can be signaled.
[0191] Based on at least one of the aforementioned configurations, the number of motion vector correction steps, whether sub-block-based motion vector correction is performed, and the correction method can be explicitly signaled. For example, information regarding the correction method for block-based motion vector correction and information regarding the number of steps constituting block-based motion vector correction can be defined identically in the encoding device and the decoding device without separate signaling. On the other hand, information regarding whether sub-block-based motion vector correction is performed can be signaled separately. Information regarding the number of steps constituting sub-block-based motion vector correction can be signaled only when sub-block-based motion vector correction is performed. Through this, accuracy can be efficiently increased while minimizing the amount of information transmitted.
[0192] The number of motion vector correction steps, the signaling location of information regarding whether subblock-based motion vector correction is performed, and the application unit can be defined as follows.
[0193] pic_parameter_set_rbsp( ) {Descriptor...pps_num_dmvr_stage_minus2ue(v)...}
[0194] Table 1 shows an example in which information regarding the number of motion vector correction stages (pps_num_dmvr_stage_minus2) is defined in the picture parameter set (PPS). In this case, the minimum value of the number of motion vector correction stages may be 2. In this case, for efficient signaling, the encoding device may derive a value obtained by subtracting 2 from the number of motion vector correction stages and transmit it to the decoding device. The decoding device may determine the number of motion vector correction stages by adding 2 to the value transmitted from the encoding device. However, this is merely an example, and the present disclosure is not limited to cases where the minimum value of the number of motion vector correction stages is 2.
[0195] pic_parameter_set_rbsp( ) {Descriptor...pps_num_cu_dmvr_stage_minus2ue(v)pps_num_subcu_dmvr_stage_minus2ue(v)...}
[0196] Table 1 defines the number of motion vector correction stages as a single syntax element, but as shown in Table 2, information regarding the number of block-based motion vector correction stages (pps_num_cu_dmvr_stage_minus2) and information regarding the number of sub-block-based motion vector correction stages (pps_num_subcu_dmvr_stage_minus2) may be defined separately. Here, it is assumed that the minimum value that the number of block-based motion vector correction stages can have and the minimum value that the number of sub-block-based motion vector correction stages can have are each 2. However, this is merely an example, and any minimum value may be commonly defined and used in the encoding / decoding device.
[0197] Meanwhile, as shown in Table 1, the number of motion vector correction steps can be defined as a single syntax element, but it can be commonly defined and used in encoding / decoding devices so that its meaning can be interpreted in various ways. For example, the single syntax element may represent the total number of motion vector correction steps. Alternatively, the single syntax element may represent the number of block-based motion vector correction steps. Alternatively, the single syntax element may represent the number of sub-block-based motion vector correction steps.
[0198] If a single syntax element refers to the number of partial correction steps (e.g., the number of block or subblock-based motion vector correction steps) rather than the total number of motion vector correction steps, the number of remaining steps not determined by a single syntax element may be commonly defined and used by the encoding / decoding device. Alternatively, the remaining steps not determined by a single syntax element may be determined not to be performed separately.
[0199] Information regarding the number of motion vector correction steps, information regarding whether sub-block-based motion vector correction is performed, etc., may be signaled in the PPS. This may mean that the information defined in the PPS is applied to picture units affected by the PPS. However, the location where information regarding the number of motion vector correction steps and information regarding whether sub-block-based motion vector correction is performed is not limited to the PPS, and may be signaled at least one of the VPS, SPS, PPS, Picture Header, Slice Header, CTU, or CU. The scope of application of information regarding the number of motion vector correction steps and information regarding whether sub-block-based motion vector correction is performed may be determined at the higher level where the information is defined (e.g., sequence, picture, slice) or at a specific encoding / decoding unit.
[0200] At least one of the information for determining whether to perform motion vector correction based on subblocks, or information regarding the number of motion vector correction steps, can be signaled based on information regarding whether to activate motion vector correction with variable steps.
[0201] seq_parameter_set_rbsp() {Descriptor...sps_adaptive_stage_dmvr_enabled_flagu(1)...}
[0202] pic_parameter_set_rbsp( ) {Descriptor...if( sps_adaptive_stage_dmvr_enabled_flag )pps_num_dmvr_stage_minus2ue(v)...}
[0203] Referring to Table 3, information regarding whether the motion vector correction with variable stages is enabled in SPS (sps_adaptive_stage_dmvr_enabled_flag) can be signaled. If the value of sps_adaptive_stage_dmvr_enabled_flag is 1, it may indicate that the motion vector correction with variable stages is enabled, and if the value of sps_adaptive_stage_dmvr_enabled_flag is 0, it may indicate that the motion vector correction with variable stages is not enabled.
[0204] Referring to Table 4, if the value of sps_adaptive_stage_dmvr_enabled_flag is 1, information regarding the number of motion vector correction stages in the PPS may be signaled, and if not, information regarding the number of motion vector correction stages in the PPS may not be signaled. In addition, based on the value of sps_adaptive_stage_dmvr_enabled_flag being 1, information regarding whether subblock-based motion vector correction is performed may be signaled.
[0205] In Table 3, information regarding whether motion vector correction with variable steps is enabled is signaled in the SPS, and in Table 4, information regarding the number of motion vector correction steps is signaled in the PPS, but this is merely an example.
[0206] At least one of the information regarding whether the motion vector correction with variable steps described above is enabled, the information regarding the number of motion vector correction steps, or the information regarding whether subblock-based motion vector correction is performed may be signaled in at least one of VPS, SPS, PPS, Picture Header, Slice Header, CTU, or CU. In this case, the information regarding whether the motion vector correction with variable steps is enabled may be signaled in the same high-level syntax (HLS) or specific decoding unit (e.g., CTU, CU) as the information regarding the number of motion vector correction steps. Alternatively, the information regarding whether the motion vector correction with variable steps is enabled may be signaled in a higher HLS or decoding unit than the information regarding the number of motion vector correction steps.
[0207] Referring to FIG. 4, a prediction block of the current block can be generated based on the corrected motion vector (S420).
[0208] For example, the final motion vector of the current block can be derived based on the motion vector corrected through the aforementioned process, and the predicted block of the current block can be generated based on the final motion vector.
[0209] Alternatively, the above final motion vector may be used as an additional motion vector for multiple predictions of the current block.
[0210] For example, a first prediction block of the current block can be generated based on a motion vector derived from S400, and a second prediction block of the current block can be generated based on the final motion vector. A prediction block of the current block can be generated based on the weighted sum of the first and second prediction blocks. Alternatively, a first prediction block of the current block can be generated based on a motion vector obtained by block-based motion vector correction, and a second prediction block can be generated based on the final motion vector. A prediction block of the current block can also be generated based on the weighted sum of the first and second prediction blocks.
[0211] The current block can be restored based on the predicted block of the current block (S430).
[0212] Transform coefficients can be derived based on residual information signaled through a bitstream. A residual block can be generated by applying at least one of inverse quantization or inverse transformation to the derived transformation coefficients. A restoration block of the current block can be generated based on the prediction block of the current block and the residual block.
[0213] FIG. 6 illustrates a schematic configuration of a decoding device (300) that performs a decoding method according to the present disclosure.
[0214] Referring to FIG. 6, the decoding device (300) may include a motion information induction unit (600), a prediction block generation unit (610), and a restoration unit (620). The motion information induction unit (600) and the prediction block generation unit (610) may be provided in the inter prediction unit (332) of FIG. 3.
[0215] The motion information induction unit (600) can perform the process of inducing a motion vector according to S400. Additionally, the motion information induction unit (600) can perform the process of correcting a motion vector according to S410. The prediction block generation unit (610) can perform the process of generating a prediction block according to S420. The restoration unit (620) can perform the process of restoring the current block according to S430.
[0216] FIG. 7 illustrates an encoding method performed by an encoding device (200) as an embodiment according to the present disclosure.
[0217] Referring to FIG. 7, a motion vector can be derived for the current block (S700).
[0218] The above motion vector can be derived based on a predetermined inter-prediction mode. The predetermined inter-prediction mode may be a merge mode or an AMVP (advanced motion vector prediction) mode. However, it is not limited thereto, and the motion vector may be understood as being replaced by a motion vector predictor.
[0219] Referring to FIG. 7, the motion vector for the current block can be corrected based on a pre-defined method (S710). The method for correcting the motion vector is as described with reference to FIG. 4, and a redundant explanation is omitted here.
[0220] Referring to FIG. 7, a prediction block of the current block can be generated based on the corrected motion vector (S720). The method for generating the prediction block is as described with reference to FIG. 4.
[0221] Referring to FIG. 7, the transformation coefficients of the current block can be derived based on the residual block of the current block (S730). The residual block of the current block can be generated based on the prediction block generated in S720. The transformation coefficients can be derived by performing at least one of transformation or quantization on the residual block.
[0222] Referring to FIG. 7, a bitstream can be generated by encoding residual information regarding the conversion coefficients of the current block (S740).
[0223] FIG. 8 illustrates a schematic configuration of an encoding device (200) that performs an encoding method according to the present disclosure.
[0224] Referring to FIG. 8, the encoding device (200) may include a motion information induction unit (800), a prediction block generation unit (810), a transformation coefficient induction unit (820), and a residual information encoding unit (830).
[0225] The motion information induction unit (800) and the prediction block generation unit (810) may be provided in the inter prediction unit (221) of FIG. 2. The transformation coefficient induction unit (820) may be provided in the residual processing unit (230) of FIG. 2. The residual information encoding unit (830) may be provided in the entropy encoding unit (240).
[0226] The motion information derivation unit (800) can perform the process of deriving a motion vector according to S700. Additionally, the motion information derivation unit (800) can perform the process of correcting a motion vector according to S710. The prediction block generation unit (810) can perform the process of generating a prediction block according to S720. The transformation coefficient derivation unit (820) can perform the process of deriving a transformation coefficient according to S730. The residual information encoding unit (830) can perform the encoding process of residual information according to S740.
[0227] In the embodiments described above, methods are described based on flowcharts as a series of steps or blocks; however, the embodiments are not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps as described above. Furthermore, those skilled in the art will understand that the steps shown in the flowcharts are not exclusive, and other steps may be included, or one or more steps of the flowcharts may be omitted without affecting the scope of the embodiments of this document.
[0228] The method according to the embodiments of the present document described above may be implemented in the form of software, and the encoding device and / or decoding device according to the present document may be included in a device that performs image processing, such as a TV, computer, smartphone, set-top box, display device, etc.
[0229] When the embodiments described in this document are implemented in software, the method described above may be implemented as a module (process, function, etc.) that performs the function described above. The module may be stored in memory and executed by a processor. The memory may be located inside or outside the processor and may be connected to the processor by various well-known means. The processor may include an application-specific integrated circuit (ASIC), other chipsets, logic circuits, and / or data processing devices. The memory may include read-only memory (ROM), random access memory (RAM), flash memory, memory cards, storage media, and / or other storage devices. That is, the embodiments described in this document may be implemented and executed on a processor, microprocessor, controller, or chip. For example, the functional units illustrated in each figure may be implemented and executed on a computer, processor, microprocessor, controller, or chip. In this case, information on instructions or algorithms for implementation may be stored on a digital storage medium.
[0230] In addition, the decoding device and encoding device to which the embodiment(s) of the present specification are applied may be included in multimedia broadcasting transmission and reception devices, mobile communication terminals, home cinema video devices, digital cinema video devices, surveillance cameras, video conversation devices, real-time communication devices such as video communication, mobile streaming devices, storage media, camcorders, Video on Demand (VoD) service providers, Over-the-top video (OTT) devices, internet streaming service providers, 3D video devices, virtual reality (VR) devices, augmented reality (AR) devices, video phone video devices, transportation terminals (e.g., vehicle terminals (including autonomous vehicles), airplane terminals, ship terminals, etc.), and medical video devices, and may be used to process video signals or data signals. For example, Over-the-top video (OTT) devices may include game consoles, Blu-ray players, internet-connected TVs, home theater systems, smartphones, tablet PCs, Digital Video Recorders (DVRs), etc.
[0231] Additionally, the processing method to which the embodiment(s) of this specification are applied may be produced in the form of a program that is executed by a computer and may be stored on a computer-readable recording medium. Multimedia data having a data structure according to the embodiment(s) of this specification may also be stored on a computer-readable recording medium. The computer-readable recording medium includes all types of storage devices and distributed storage devices in which computer-readable data is stored. The computer-readable recording medium may include, for example, a Blu-ray disc (BD), a Universal Serial Bus (USB), a ROM, a PROM, an EPROM, an EEPROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device. Additionally, the computer-readable recording medium includes a medium implemented in the form of a carrier wave (e.g., transmission over the Internet). Furthermore, a bitstream generated by an encoding method may be stored on a computer-readable recording medium or transmitted via a wired or wireless communication network.
[0232] Additionally, the embodiments of this specification may be implemented as a computer program product by program code, and said program code may be executed on a computer by the embodiments of this specification. said program code may be stored on a carrier readable by a computer.
[0233] FIG. 9 shows an example of a content streaming system to which embodiments of the present disclosure can be applied.
[0234] Referring to FIG. 9, a content streaming system to which the embodiment(s) of the present specification are applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0235] The above encoding server compresses content input from multimedia input devices, such as smartphones, cameras, and camcorders, into digital data to generate a bitstream and transmits it to the streaming server. As another example, if multimedia input devices, such as smartphones, cameras, and camcorders, generate the bitstream directly, the encoding server may be omitted.
[0236] The bitstream above may be generated by an encoding method or a bitstream generation method to which the embodiment(s) of the present specification are applied, and the streaming server may temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0237] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server acts as a medium to inform the user of available services. When a user requests a desired service from the web server, the web server transmits it to the streaming server, and the streaming server transmits the multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server plays the role of controlling commands and responses between each device within the content streaming system.
[0238] The streaming server may receive content from a media storage and / or an encoding server. For example, when receiving content from the encoding server, the content may be received in real time. In this case, to provide a seamless streaming service, the streaming server may store the bitstream for a certain period of time.
[0239] Examples of the above user devices may include mobile phones, smartphones, laptop computers, digital broadcasting terminals, PDAs (personal digital assistants), PMPs (portable multimedia players), navigation systems, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs (head-mounted displays)), digital TVs, desktop computers, digital signage, etc.
[0240] Each server within the above-mentioned content streaming system can be operated as a distributed server, and in this case, data received from each server can be processed in a distributed manner.
[0241] The claims described in this specification may be combined in various ways. For example, the technical features of the method claims in this specification may be combined to be implemented as a device, and the technical features of the device claims in this specification may be combined to be implemented as a method. Furthermore, the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a device, and the technical features of the method claims and the technical features of the device claims in this specification may be combined to be implemented as a method.
Claims
1. A step for deriving a motion vector for the current block; A step of correcting the motion vector for the current block based on a pre-defined method; A step of generating a prediction block of the current block based on the corrected motion vector; and The method includes the step of restoring the current block based on the prediction block, The above-defined method comprises at least one of block-based motion vector correction or sub-block-based motion vector correction.
2. In Paragraph 1, A method in which the above block-based motion vector correction comprises at least one of BM (bilateral matching)-based motion vector correction or BDOF (bi-directional optical flow)-based motion vector correction.
3. In Paragraph 1, A method in which the above subblock-based motion vector correction comprises at least one of BM (bilateral matching)-based motion vector correction or BDOF (bi-directional optical flow)-based motion vector correction.
4. In Paragraph 3, The subblock for motion vector correction based on the above BM is the first subblock, and The subblock for motion vector correction based on the above BDOF is a second subblock, a method.
5. In Paragraph 1, The above block-based motion vector correction is performed based on BM (bilateral matching), and A method in which the above-described subblock-based motion vector correction is performed based on BDOF (bi-directional optical flow).
6. In Paragraph 1, The above-described subblock-based motion vector correction is performed by selectively using either BM (bilateral matching) or BDOF (bi-directional optical flow) depending on whether a predetermined condition is satisfied.
7. In Paragraph 1, A method in which a subblock for motion vector correction based on the above subblock is determined as either a first subblock or a second subblock depending on whether a predetermined condition is satisfied.
8. In Paragraph 1, The above-defined method further comprises motion vector correction based on interweaved subblocks.
9. In Paragraph 8, A method in which the above-mentioned interweaved subblock-based motion vector correction is performed in parallel or independently of the above-mentioned subblock-based motion vector correction.
10. In Paragraph 9, A method in which whether to perform motion vector correction based on the above interweaved subblock is determined based on the size of the current block.
11. Step of deriving the motion vector for the current block; A step of correcting the motion vector for the current block based on a pre-defined method; A step of generating a prediction block of the current block based on the corrected motion vector; A step of generating a residual block of the current block based on the prediction block above; A step of deriving transformation coefficients of the current block based on the above residual block; and The method includes the step of encoding residual information regarding the above-mentioned transformation coefficients, The above-defined method comprises at least one of block-based motion vector correction or sub-block-based motion vector correction.
12. A computer-readable storage medium for storing a bitstream generated by the method according to paragraph 11.
13. A step of acquiring a bitstream for image information; wherein the bitstream is generated based on the steps of deriving a motion vector for a current block, correcting the motion vector for the current block based on a pre-defined method, generating a prediction block of the current block based on the corrected motion vector, generating a residual block of the current block based on the prediction block, deriving transformation coefficients of the current block based on the residual block, and encoding residual information regarding the transformation coefficients, and The method includes the step of transmitting data including the above bitstream, The above-defined method comprises at least one of block-based motion vector correction or sub-block-based motion vector correction.
Citation Information
Patent Citations
A laser pointer system on outside mirrors working in conjunction with a turn signal light
KR1020240108243A
Mobile robot with variable driving machanism and driving control method of the same
KR1020250171620A
KR20210137463A
KR20230123946A
KR20240025058A