Image encoding / decoding method and device, and recording medium on which bitstream is stored
The method enhances video encoding/decoding efficiency by deriving fusion merge candidates and generating prediction blocks, addressing inefficiencies in high-resolution image processing and reducing associated costs.
Patent Information
- Application Number
- PCT/KR2025/003875
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-03-26
- Publication Date
- 2025-10-02
AI Technical Summary
Existing video encoding/decoding technologies face inefficiencies in high-resolution and high-quality image processing due to suboptimal merge candidate selection, leading to low prediction accuracy and increased transmission and storage costs.
A method for deriving fusion merge candidates and generating prediction blocks by adding fused merge candidates to the merge candidate list, based on motion information of neighboring blocks and blocks decoded before the current block, with motion vector weighting and rearrangement considering similarity and distance, to enhance encoding/decoding efficiency.
Improves prediction accuracy and encoding/decoding efficiency, reducing the costs associated with high-resolution and high-quality image data transmission and storage.
Smart Images

Figure KR2025003875_02102025_PF_FP_ABST
Abstract
Description
Video encoding / decoding method, device, and recording medium storing bitstream
[0001] The present invention relates to a video encoding / decoding method, device, and recording medium storing a bitstream. Specifically, the present invention relates to a video encoding / decoding method, device, and recording medium storing a bitstream based on an improved pairwise average merge candidate derivation method and a fusion merge candidate derivation method.
[0002] Recently, the demand for high-resolution, high-quality images, such as UHD (Ultra High Definition) images, is increasing across various application fields. As image data becomes higher in resolution and quality, the relative amount of data increases compared to conventional image data. Therefore, transmitting image data using existing media such as wired or wireless broadband lines or storing it using existing storage media leads to increased transmission and storage costs. To address these issues arising from the increasing resolution and quality of image data, high-efficiency image encoding / decoding technologies for higher-resolution and higher-quality images are required.
[0003] In merge mode, the motion information of one merge candidate selected from the merge candidate list is used as is for the prediction of the current block, which may lower the prediction accuracy.
[0004] The pair average merge candidate derivation method may not provide the optimal merge candidate because it uses merge candidates at fixed positions, which may result in low prediction accuracy.
[0005] The purpose of the present invention is to provide a video encoding / decoding method and device with improved encoding / decoding efficiency.
[0006] In addition, the present invention aims to provide a recording medium storing a bitstream generated by an image decoding method or device according to the present invention.
[0007] In addition, the present invention aims to provide a method for deriving a fusion merge candidate and a method for generating a prediction block to solve the problems of the existing merge mode as described above.
[0008] In addition, the present invention aims to provide an improved pair average merge candidate derivation method for solving the problems of the above pair average merge candidate derivation method.
[0009] A video decoding method according to one embodiment of the present invention includes the steps of generating a merge candidate list based on at least one of motion information of a neighboring block of a current block and motion information of a block decoded before the current block, adding a fused merge candidate to the merge candidate list when the number of merge candidates in the merge candidate list is less than a maximum number of merge candidates, and generating a prediction block of the current block based on the merge candidate list, wherein the fused merge candidate is derived based on at least two merge candidates selected from the merge candidate list, and the fused merge candidate can be added to the merge candidate list until the number of merge candidates in the merge candidate list becomes a predefined number.
[0010] In the above image decoding method, the fusion merge candidate may be added to the merge candidate list if it has motion information that does not overlap with the merge candidates in the merge candidate list.
[0011] In the above image decoding method, the position of the merge candidate selected from the merge candidate list within the merge candidate list may be any position.
[0012] In the above image decoding method, the motion vector of the fusion merge candidate can be determined by weighting the motion vectors of the merge candidates selected from the merge candidate list.
[0013] In the above image decoding method, the weight of the motion vector can be determined based on the size of the motion vector.
[0014] In the above image decoding method, the step of generating the merge candidate list can be performed by rearranging the merge candidates in the merge candidate list.
[0015] In the above image decoding method, when the merge candidates in the merge candidate list are reordered, the reordering may be based on the similarity between the template of the reference block indicated by the merge candidate and the template of the current block.
[0016] In the above image decoding method, the merge candidate selected from the merge candidate list can be selected based on the similarity between the template of the reference block indicated by the merge candidate in the merge candidate list and the template of the current block.
[0017] In the above image decoding method, the merge candidate selected from the merge candidate list can be selected based on the distance between the reference picture of the merge candidate in the merge candidate list and the current picture.
[0018] In the above image decoding method, a merge candidate selected from the merge candidate list may be a merge candidate having the same reference picture among the merge candidates in the merge candidate list.
[0019] According to one embodiment of the present invention, a video encoding method includes the steps of generating a merge candidate list based on at least one of motion information of a neighboring block of a current block and motion information of a block encoded before the current block, adding a fused merge candidate to the merge candidate list when the number of merge candidates in the merge candidate list is less than a maximum number of merge candidates, and generating a prediction block of the current block based on the merge candidate list, wherein the fused merge candidate is derived based on at least two merge candidates selected from the merge candidate list, and the fused merge candidate can be added to the merge candidate list until the number of merge candidates in the merge candidate list becomes a predefined number.
[0020] A non-transitory computer-readable recording medium according to one embodiment of the present invention can store a bitstream generated by the image encoding method.
[0021] A bitstream transmission method according to one embodiment of the present invention can transmit a bitstream generated by the image encoding method.
[0022] The features briefly summarized above regarding the present disclosure are merely exemplary aspects of the detailed description of the present disclosure that follows and do not limit the scope of the present disclosure.
[0023] According to the present invention, a video encoding / decoding method and device with improved encoding / decoding efficiency can be provided.
[0024] Additionally, according to the present invention, a method for generating a merge candidate list based on derivation of fusion merge candidates can be provided.
[0025] Additionally, according to the present invention, a method for generating a prediction block based on an initial prediction block weighted sum can be provided.
[0026] Additionally, according to the present invention, prediction accuracy can be improved.
[0027] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains from the description below.
[0028] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.
[0029] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.
[0030] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.
[0031] FIG. 4 is a diagram for explaining a pair average merge candidate derivation method according to one embodiment of the present invention.
[0032] FIG. 5 is a drawing for explaining a method for deriving a fusion merge candidate according to one embodiment of the present invention.
[0033] Figure 6 is a flowchart illustrating an image decoding method according to one embodiment of the present invention.
[0034] FIG. 7 is a drawing exemplarily showing a content streaming system to which an embodiment according to the present invention can be applied.
[0035] A video decoding method according to one embodiment of the present invention includes the steps of generating a merge candidate list based on at least one of motion information of a neighboring block of a current block and motion information of a block decoded before the current block, adding a fused merge candidate to the merge candidate list when the number of merge candidates in the merge candidate list is less than a maximum number of merge candidates, and generating a prediction block of the current block based on the merge candidate list, wherein the fused merge candidate is derived based on at least two merge candidates selected from the merge candidate list, and the fused merge candidate can be added to the merge candidate list until the number of merge candidates in the merge candidate list becomes a predefined number.
[0036] The present invention is susceptible to various modifications and embodiments. Therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the present invention to specific embodiments, but rather to encompass all modifications, equivalents, and substitutes falling within the spirit and scope of the present invention. In the drawings, similar reference numerals designate the same or similar functions throughout. The shape and size of elements in the drawings may be provided by way of example only for clarity. The detailed description of the exemplary embodiments described below refers to the accompanying drawings, which illustrate specific embodiments. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that the various embodiments, while different, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Furthermore, it should be understood that the location or arrangement of individual components within each disclosed embodiment may be modified without departing from the spirit and scope of the embodiment. Accordingly, the detailed description set forth below is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments, if properly described, is defined only by the appended claims, along with the full scope equivalents to which such claims are entitled.
[0037] In the present invention, terms such as first, second, etc. may be used to describe various components, but the components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the second component, and similarly, the second component may also be referred to as the first component. The term "and / or" includes a combination of multiple related described items or any of multiple related described items.
[0038] The components shown in the embodiments of the present invention are independently depicted to represent different characteristic functions, and do not mean that each component is composed of separate hardware or a single software component. That is, each component is listed and included as a separate component for convenience of explanation, and at least two components among each component may be combined to form a single component, or a single component may be divided into multiple components to perform a function, and such integrated and separate embodiments of each component are also included in the scope of the present invention as long as they do not deviate from the essence of the present invention.
[0039] The terminology used herein is merely used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In addition, some components of the present invention are not essential components that perform essential functions in the present invention and may be optional components merely for performance enhancement. The present invention may be implemented by including only components essential to realizing the essence of the present invention, excluding components used only for performance enhancement, and a structure including only essential components, excluding optional components used only for performance enhancement, is also within the scope of the present invention.
[0040] In embodiments, the term "at least one" may mean one of a number greater than or equal to 1, such as 1, 2, 3, and 4. In embodiments, the term "a plurality of" may mean one of a number greater than or equal to 2, such as 2, 3, and 4.
[0041] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In describing the embodiments of this specification, if it is determined that a detailed description of a related known configuration or function may obscure the gist of this specification, the detailed description will be omitted. The same reference numerals will be used for identical components in the drawings, and duplicate descriptions of identical components will be omitted.
[0042] Glossary of Terms
[0043] Hereinafter, “video” may mean a single picture constituting a video, or may refer to the video itself. For example, “encoding and / or decoding of a video” may mean “encoding and / or decoding of a video,” or may mean “encoding and / or decoding of one of the videos constituting the video.”
[0044] Hereinafter, the terms "video" and "movie" may be used interchangeably and have the same meaning. Furthermore, the target image may be an encoding target image, which is the target of encoding, and / or a decoding target image, which is the target of decoding. Furthermore, the target image may be an input image input to an encoding device, or an input image input to a decoding device. Here, the target image may have the same meaning as the current image.
[0045] Hereinafter, the terms encoder and image encoding device may be used interchangeably and have the same meaning.
[0046] Hereinafter, the terms decoder and image decoding device may be used interchangeably and have the same meaning.
[0047] Hereinafter, “image”, “picture”, “frame” and “screen” may be used with the same meaning and may be used interchangeably.
[0048] Hereinafter, the term "target block" may refer to an encoding target block, which is the target of encoding, and / or a decoding target block, which is the target of decoding. Furthermore, the target block may refer to a current block, which is the target of current encoding and / or decoding. For example, the terms "target block" and "current block" may be used interchangeably and have the same meaning.
[0049] Hereinafter, "block" and "unit" may be used with the same meaning and may be used interchangeably. In addition, "unit" may mean including a luminance component block and a corresponding chroma component block to distinguish it from a block. For example, a coding tree unit (CTU) may be composed of one luma component (Y) coding tree block (CTB) and two chroma component (Cb, Cr) coding tree blocks associated with it.
[0050] Hereinafter, the terms “sample,” “pixel,” and “pixel” may be used interchangeably and have the same meaning. Here, a sample may represent a basic unit that constitutes a block.
[0051] Hereinafter, “inter” and “between screens” may be used interchangeably and have the same meaning.
[0052] Hereinafter, “intra” and “within screen” may be used interchangeably and have the same meaning.
[0053]
[0054] Figure 1 is a block diagram showing the configuration according to one embodiment of an encoding device to which the present invention is applied.
[0055] The encoding device (100) may be an encoder, a video encoding device, or an image encoding device. A video may include one or more images. The encoding device (100) may sequentially encode one or more images.
[0056] Referring to FIG. 1, the encoding device (100) may include an image segmentation unit (110), an intra prediction unit (120), a motion prediction unit (121), a motion compensation unit (122), a switch (115), a subtractor (113), a transformation unit (130), a quantization unit (140), an entropy encoding unit (150), an inverse quantization unit (160), an inverse transformation unit (170), an adder (117), a filter unit (180), and a reference picture buffer (190).
[0057] Additionally, the encoding device (100) can generate a bitstream including encoded information through encoding an input image and output the generated bitstream. The generated bitstream can be stored in a computer-readable recording medium or can be streamed via a wired / wireless transmission medium.
[0058] The video segmentation unit (110) can segment the input video into various forms to increase the efficiency of video encoding / decoding. That is, the input video is composed of multiple pictures, and one picture can be hierarchically segmented and processed for compression efficiency, parallel processing, etc. For example, one picture can be segmented into one or more tiles or slices, which can then be segmented into multiple Coding Tree Units (CTUs). Alternatively, one picture can first be segmented into multiple sub-pictures defined as groups of rectangular slices, and each sub-picture can then be segmented into the tiles / slices. Here, the sub-pictures can be utilized to support the function of partially independently encoding / decoding and transmitting the picture. Since multiple sub-pictures can each be individually restored, there is an advantage of easy editing in applications that configure multi-channel input into a single picture. In addition, tiles can be segmented horizontally to generate bricks. Here, a brick can be utilized as the basic unit of intra-picture parallel processing. In addition, one CTU can be recursively split into a quadtree (QT), and the terminal node of the split can be defined as a coding unit (CU). The CU can be split into a prediction unit (PU) and a transformation unit (TU), and prediction and splitting can be performed. Meanwhile, the CU can be utilized as a prediction unit and / or a transformation unit itself. Here, for flexible splitting, each CTU can be recursively split into a multi-type tree (MTT) as well as a quadtree (QT). Splitting of a CTU into a multi-type tree can start from the terminal node of a QT, and the MTT can be composed of a binary tree (BT) and a triple tree (TT).For example, the MTT structure can be divided into vertical binary split mode (SPLIT_BT_VER), horizontal binary split mode (SPLIT_BT_HOR), vertical ternary split mode (SPLIT_TT_VER), and horizontal ternary split mode (SPLIT_TT_HOR). In addition, the minimum block size (MinQTSize) of the quad tree of the luminance block during splitting can be set to 16x16, the maximum block size (MaxBtSize) of the binary tree can be set to 128x128, and the maximum block size (MaxTtSize) of the triple tree can be set to 64x64. In addition, the minimum block size (MinBtSize) of the binary tree and the minimum block size (MinTtSize) of the triple tree can be set to 4x4, and the maximum depth (MaxMttDepth) of the multi-type tree can be set to 4. Additionally, to improve the encoding efficiency of the I slice, a dual tree can be applied that uses different CTU partition structures for luminance and chrominance components. On the other hand, in the P and B slices, the luminance and chrominance CTBs (Coding Tree Blocks) within the CTU can be partitioned into a single tree that shares the coding tree structure.
[0059] The encoding device (100) may perform encoding on the input image in intra mode and / or inter mode. Alternatively, the encoding device (100) may perform encoding on the input image in a third mode (e.g., IBC mode, Palette mode, etc.) other than the intra mode and inter mode. However, if the third mode has functional characteristics similar to the intra mode or inter mode, it may be classified as intra mode or inter mode for convenience of explanation. In the present invention, the third mode will be classified and described separately only when a specific description is required.
[0060] When the intra mode is used as the prediction mode, the switch (115) can be switched to intra, and when the inter mode is used as the prediction mode, the switch (115) can be switched to inter. Here, the intra mode can mean an intra-screen prediction mode, and the inter mode can mean an inter-screen prediction mode. The encoding device (100) can generate a prediction block for an input block of an input image. In addition, after the prediction block is generated, the encoding device (100) can encode a residual block using a residual of the input block and the prediction block. The input image can be referred to as a current image that is currently a target of encoding. The input block can be referred to as a current block that is currently a target of encoding or an encoding target block.
[0061] When the prediction mode is intra mode, the intra prediction unit (120) can use samples of blocks already encoded / decoded around the current block as reference samples. The intra prediction unit (120) can perform spatial prediction on the current block using the reference samples, and can generate prediction samples for the input block through spatial prediction. Here, intra prediction can mean prediction within the screen.
[0062] As an intra prediction method, non-directional prediction modes such as DC mode and Planar mode, as well as directional prediction modes (e.g., 65 directions) can be applied. Here, the intra prediction method can be expressed as an intra prediction mode or an intra-screen prediction mode.
[0063] When the prediction mode is inter mode, the motion prediction unit (121) can search for an area that best matches the input block from the reference image during the motion prediction process and derive a motion vector using the searched area. At this time, the area can be used as a search area. The reference image can be stored in the reference picture buffer (190). Here, when encoding / decoding for the reference image is processed, it can be stored in the reference picture buffer (190).
[0064] The motion compensation unit (122) can generate a prediction block for the current block by performing motion compensation using a motion vector. Here, inter prediction may mean inter-screen prediction or motion compensation.
[0065] The above motion prediction unit (121) and motion compensation unit (122) can generate a prediction block by applying an interpolation filter to a portion of an area within a reference image when the value of the motion vector does not have an integer value. In order to perform inter-screen prediction or motion compensation, it is possible to determine whether the motion prediction and motion compensation method of the prediction unit included in the corresponding encoding unit is one of Skip Mode, Merge Mode, Advanced Motion Vector Prediction (AMVP) mode, and Intra Block Copy (IBC) mode based on the encoding unit, and perform inter-screen prediction or motion compensation according to each mode.
[0066] In addition, based on the above inter-screen prediction method, the AFFINE mode of sub-PU based prediction, the SbTMVP (Subblock-based Temporal Motion Vector Prediction) mode, and the MMVD (Merge with MVD) mode and the GPM (Geometric Partitioning Mode) mode of PU based prediction can be applied. In addition, in order to improve the performance of each mode, the HMVP (History based MVP), the PAMVP (Pairwise Average MVP), the CIIP (Combined Intra / Inter Prediction), the AMVR (Adaptive Motion Vector Resolution), the BDOF (Bi-Directional Optical-Flow), the BCW (Bi-predictive with CU Weights), the LIC (Local Illumination Compensation), the TM (Template Matching), and the OBMC (Overlapped Block Motion Compensation) can be applied.
[0067] Among these, AFFINE mode is a technology that is used in both AMVP and MERGE modes and also has high encoding efficiency. In the existing video coding standard, since MC (Motion Compensation) is performed by considering only the parallel translation of the block, there was a disadvantage in that it could not properly compensate for motions that occur in reality, such as zoom in / out and rotation. To supplement this, a 4-parameter affine motion model using two control point motion vectors (CPMV) and a 6-parameter affine motion model using three control point motion vectors can be applied to inter prediction. Here, CPMV is a vector representing the affine motion model of one of the upper left, upper right, and lower left of the current block.
[0068] The subtractor (113) can generate a residual block using the difference between the input block and the predicted block. The residual block may also be referred to as a residual signal. The residual signal may refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal may be a signal generated by transforming, quantizing, or transforming and quantizing the difference between the original signal and the predicted signal. The residual block may be a residual signal in block units.
[0069] The transform unit (130) can perform a transform on the residual block to generate a transform coefficient and output the generated transform coefficient. Here, the transform coefficient may be a coefficient value generated by performing a transform on the residual block. When the transform skip mode is applied, the transform unit (130) may also skip the transform on the residual block.
[0070] Quantized levels can be generated by applying quantization to transform coefficients or residual signals. In the following embodiments, quantized levels may also be referred to as transform coefficients.
[0071] For example, a 4x4 luminance residual block generated through within-screen prediction can be transformed using a basis vector based on DST (Discrete Sine Transform), and the remaining residual blocks can be transformed using a basis vector based on DCT (Discrete Cosine Transform). In addition, through RQT (Residual Quad Tree) technology, the transform block is divided into a quad tree shape for one block, and after performing transformation and quantization on each transform block divided through RQT, a coded block flag (cbf) can be transmitted to increase encoding efficiency when all coefficients become 0.
[0072] Another alternative is to apply Multiple Transform Selection (MTS) technology, which selectively performs transformation using multiple transformation bases. That is, instead of dividing CUs into TUs via RQT, a Sub-block Transform (SBT) technology can perform a function similar to TU division. Specifically, SBT is applied only to inter-screen prediction blocks, and unlike RQT, it can divide the current block into ½ or ¼ blocks vertically or horizontally, and then perform transformation on only one of the blocks. For example, in a vertically divided block, the transformation can be performed on the leftmost or rightmost block, and in a horizontally divided block, the transformation can be performed on the topmost or bottommost block.
[0073] Additionally, LFNST (Low Frequency Non-Separable Transform), a secondary transform technique that further transforms the residual signal converted to the frequency domain through DCT or DST, can be applied. LFNST additionally performs a transform on the low-frequency region of 4x4 or 8x8 in the upper left, which allows the residual coefficients to be concentrated in the upper left.
[0074] The quantization unit (140) can generate a quantized level by quantizing a transform coefficient or residual signal according to a quantization parameter (QP), and can output the generated quantized level. At this time, the quantization unit (140) can quantize the transform coefficient using a quantization matrix.
[0075] For example, a quantizer with QP values of 0 to 51 can be used. Alternatively, if the image size is larger and high encoding efficiency is required, a QP of 0 to 63 can be used. In addition, a Dependent Quantization (DQ) method that uses two quantizers instead of a single quantizer can be applied. DQ performs quantization using two quantizers (e.g., Q0 and Q1), but even without signaling information about the use of a specific quantizer, the quantizer to be used for the next transform coefficient can be selected based on the current state through a state transition model.
[0076] The entropy encoding unit (150) can generate a bitstream by performing entropy encoding according to a probability distribution on values produced by the quantization unit (140) or coding parameter values produced during the encoding process, and can output the bitstream. The entropy encoding unit (150) can perform entropy encoding on information about image samples and information for decoding the image. For example, the information for decoding the image can include syntax elements, etc.
[0077] When entropy encoding is applied, a small number of bits are allocated to symbols with a high occurrence probability, and a large number of bits are allocated to symbols with a low occurrence probability, thereby representing the symbols, whereby the size of the bit string for the symbols to be encoded can be reduced. The entropy encoding unit (150) can use an encoding method such as exponential Golomb, Context-Adaptive Variable Length Coding (CAVLC), or Context-Adaptive Binary Arithmetic Coding (CABAC) for entropy encoding. For example, the entropy encoding unit (150) can perform entropy encoding using a Variable Length Coding / Code (VLC) table. In addition, the entropy encoding unit (150) may perform arithmetic encoding using the binarization method, probability model, and context model derived from the binarization method of the target symbol and the probability model of the target symbol / bin.
[0078] In this regard, when applying CABAC, the table probability update method can be changed to a simple formula-based table update method to reduce the size of the probability table stored in the decryption device. Furthermore, two different probability models can be used to obtain more accurate symbol probability values.
[0079] The entropy encoding unit (150) can change a two-dimensional block form coefficient into a one-dimensional vector form through a transform coefficient scanning method to encode a transform coefficient level (quantized level).
[0080] Coding parameters may include not only information (flags, indexes, etc.) encoded in an encoding device (100) and signaled to a decoding device (200), such as syntax elements, but also information derived during an encoding process or a decoding process, and may mean information necessary when encoding or decoding an image.
[0081] Here, signaling a flag or index may mean that the encoder entropy encodes the flag or index and includes it in the bitstream, and that the decoder entropy decodes the flag or index from the bitstream.
[0082] The encoded current image can be used as a reference image for other images to be processed later. Accordingly, the encoding device (100) can reconstruct or decode the encoded current image again and store the reconstructed or decoded image as a reference image in the reference picture buffer (190).
[0083] The quantized level can be dequantized in the dequantization unit (160) and inversely transformed in the inverse transformation unit (170). The dequantized and / or inversely transformed coefficients can be combined with a prediction block through an adder (117), and a reconstructed block can be generated by combining the dequantized and / or inversely transformed coefficients and the prediction block. Here, the dequantized and / or inversely transformed coefficients refer to coefficients on which at least one of dequantization and inverse transformation has been performed, and may refer to a reconstructed residual block. The dequantization unit (160) and the inverse transformation unit (170) can be performed in the reverse process of the quantization unit (140) and the transformation unit (130).
[0084] The restoration block may pass through a filter unit (180). The filter unit (180) may apply a deblocking filter, a sample adaptive offset (SAO), an adaptive loop filter (ALF), a bilateral filter (BIF), a Luma Mapping with Chroma Scaling (LMCS), etc. as a filtering technique, in whole or in part, to the restoration sample, restoration block, or restoration image. The filter unit (180) may also be referred to as an in-loop filter. In this case, the in-loop filter is also used as a name excluding LMCS.
[0085] A deblocking filter can remove block distortion that occurs at the boundaries between blocks. Whether to apply a deblocking filter to the current block can be determined based on the samples contained in several columns or rows within the block. When applying a deblocking filter to a block, different filters can be applied depending on the required deblocking filtering strength.
[0086] Sample adaptive offset can be used to compensate for encoding errors by adding an appropriate offset value to sample values. Sample adaptive offset can compensate for the offset from the original image on a sample-by-sample basis for deblocked images. This can be done by dividing the samples contained in the image into a fixed number of regions, determining the regions to be offset, and applying the offset to those regions. Alternatively, the offset can be applied by considering the edge information of each sample.
[0087] Bilateral filter (BIF) can also compensate for the offset from the original image on a sample-by-sample basis for the deblocked image.
[0088] An adaptive loop filter can perform filtering based on a comparison between a reconstructed image and the original image. By dividing the samples contained in the image into predetermined groups and determining the filter to be applied to each group, filtering can be performed differentially for each group. Information regarding whether to apply an adaptive loop filter can be signaled for each coding unit (CU), and the shape and filter coefficients of the adaptive loop filter applied to each block can vary.
[0089] In LMCS (Luma Mapping with Chroma Scaling), luma mapping (LM) refers to remapping luminance values through a piece-wise linear model, and chroma scaling (CS) refers to a technique that scales the residual values of chrominance components according to the average luminance value of the prediction signal. In particular, LMCS can be utilized as an HDR correction technique that reflects the characteristics of HDR (High Dynamic Range) images.
[0090] The restored block or restored image that has passed through the filter unit (180) may be stored in the reference picture buffer (190). The restored block that has passed through the filter unit (180) may be a part of the reference image. In other words, the reference image may be a restored image composed of restored blocks that have passed through the filter unit (180). The stored reference image may be used for inter-screen prediction or motion compensation thereafter.
[0091] Figure 2 is a block diagram showing the configuration according to one embodiment of a decryption device to which the present invention is applied.
[0092] The decoding device (200) may be a decoder, a video decoding device, or an image decoding device.
[0093] Referring to FIG. 2, the decoding device (200) may include an entropy decoding unit (210), an inverse quantization unit (220), an inverse transformation unit (230), an intra prediction unit (240), a motion compensation unit (250), an adder (201), a switch (203), a filter unit (260), and a reference picture buffer (270).
[0094] The decoding device (200) can receive a bitstream output from the encoding device (100). The decoding device (200) can receive a bitstream stored in a computer-readable recording medium, or a bitstream streamed through a wired / wireless transmission medium. The decoding device (200) can perform decoding on the bitstream in intra mode or inter mode. In addition, the decoding device (200) can generate a restored image or a decoded image through decoding, and can output the restored image or the decoded image.
[0095] If the prediction mode used for decryption is intra mode, the switch (203) can be switched to intra. If the prediction mode used for decryption is inter mode, the switch (203) can be switched to inter.
[0096] The decoding device (200) can decode the input bitstream to obtain a reconstructed residual block and generate a prediction block. Once the reconstructed residual block and the prediction block are obtained, the decoding device (200) can generate a reconstructed block to be decoded by adding the reconstructed residual block and the prediction block. The block to be decoded may be referred to as a current block.
[0097] The entropy decoding unit (210) can generate symbols by performing entropy decoding according to a probability distribution for the bitstream. The generated symbols may include symbols in the form of quantized levels. Here, the entropy decoding method may be the reverse process of the entropy encoding method described above.
[0098] The entropy decoding unit (210) can change a one-dimensional vector-shaped coefficient into a two-dimensional block-shaped coefficient through a transform coefficient scanning method to decode a transform coefficient level (quantized level).
[0099] The quantized level can be inversely quantized in the inverse quantization unit (220) and inversely transformed in the inverse transformation unit (230). The quantized level can be generated as a restored residual block as a result of performing inverse quantization and / or inverse transformation. At this time, the inverse quantization unit (220) can apply a quantization matrix to the quantized level. The inverse quantization unit (220) and inverse transformation unit (230) applied to the decoding device can apply the same technology as the inverse quantization unit (160) and inverse transformation unit (170) applied to the encoding device described above.
[0100] When intra mode is used, the intra prediction unit (240) can generate a predicted block by performing spatial prediction on the current block using sample values of already decoded blocks surrounding the block to be decoded. The intra prediction unit (240) applied to the decoding device can apply the same technology as the intra prediction unit (120) applied to the encoding device described above.
[0101] When the inter mode is used, the motion compensation unit (250) can generate a prediction block by performing motion compensation using a motion vector and a reference image stored in the reference picture buffer (270) on the current block. The motion compensation unit (250) can generate a prediction block by applying an interpolation filter to a portion of the reference image when the value of the motion vector does not have an integer value. In order to perform motion compensation, it is possible to determine whether the motion compensation method of the prediction unit included in the corresponding encoding unit is skip mode, merge mode, AMVP mode, or current picture reference mode based on the encoding unit, and motion compensation can be performed according to each mode. The motion compensation unit (250) applied to the decoding device can apply the same technology as the motion compensation unit (122) applied to the encoding device described above.
[0102] The adder (201) can add the restored residual block and the predicted block to generate a restored block. The filter unit (260) can apply at least one of an Inverse-LMCS, a deblocking filter, a sample adaptive offset, and an adaptive loop filter to the restored block or restored image. The filter unit (260) applied to the decoding device can apply the same filtering technology as that applied to the filter unit (180) applied to the encoding device described above.
[0103] The filter unit (260) can output a restored image. The restored block or restored image can be stored in the reference picture buffer (270) and used for inter prediction. The restored block that has passed through the filter unit (260) can be a part of the reference image. In other words, the reference image can be a restored image composed of restored blocks that have passed through the filter unit (260). The stored reference image can be used for inter-screen prediction or motion compensation thereafter.
[0104] FIG. 3 is a diagram schematically showing a video coding system to which the present invention can be applied.
[0105] A video coding system according to one embodiment may include an encoding device (10) and a decoding device (20). The encoding device (10) may transmit encoded video and / or image information or data to the decoding device (20) in the form of a file or streaming through a digital storage medium or a network.
[0106] An encoding device (10) according to one embodiment may include a video source generation unit (11), an encoding unit (12), and a transmission unit (13). A decoding device (20) according to one embodiment may include a reception unit (21), a decoding unit (22), and a rendering unit (23). The encoding unit (12) may be referred to as a video / image encoding unit, and the decoding unit (22) may be referred to as a video / image decoding unit. The transmission unit (13) may be included in the encoding unit (12). The reception unit (21) may be included in the decoding unit (22). The rendering unit (23) may include a display unit, and the display unit may be configured as a separate device or an external component.
[0107] The video source generation unit (11) can obtain video / images through a process of capturing, synthesizing, or generating video / images. The video source generation unit (11) can include a video / image capture device and / or a video / image generation device. The video / image capture device can include, for example, one or more cameras, a video / image archive including previously captured video / images, etc. The video / image generation device can include, for example, a computer, a tablet, a smartphone, etc., and can (electronically) generate video / images. For example, a virtual video / image can be generated through a computer, etc., in which case the video / image capture process can be replaced with a process of generating related data.
[0108] The encoding unit (12) can encode the input video / image. The encoding unit (12) can perform a series of procedures such as prediction, transformation, and quantization for compression and encoding efficiency. The encoding unit (12) can output encoded data (encoded video / image information) in the form of a bitstream. The detailed configuration of the encoding unit (12) can also be configured in the same manner as the encoding device (100) of FIG. 1 described above.
[0109] The transmission unit (13) can transmit encoded video / image information or data output in the form of a bitstream to the reception unit (21) of the decoding device (20) via a digital storage medium or a network in the form of a file or streaming. The digital storage medium can include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, etc. The transmission unit (13) can include an element for generating a media file through a predetermined file format and can include an element for transmission via a broadcasting / communication network. The reception unit (21) can extract / receive the bitstream from the storage medium or network and transmit it to the decoding unit (22).
[0110] The decoding unit (22) can decode video / image by performing a series of procedures such as inverse quantization, inverse transformation, and prediction corresponding to the operation of the encoding unit (12). The detailed configuration of the decoding unit (22) can also be configured in the same manner as the decoding device (200) of FIG. 2 described above.
[0111] The rendering unit (23) can render the decrypted video / image. The rendered video / image can be displayed through the display unit.
[0112]
[0113] Merge Mode is a prediction mode in which the predicted block of the current block is generated based on a merge candidate list. Here, the merge candidate list refers to a list containing motion information of the reference block of the current block.
[0114] Meanwhile, motion information may include information about a motion vector and a reference picture. Here, information about a reference picture may mean a reference picture index (refIdx).
[0115] In the regular merge mode, the decoder can generate a merge candidate list. Then, by performing motion compensation using the motion information of the merge candidates indicated by the signaled index among the merge candidates included in the generated merge candidate list, the decoder can generate a prediction block of the current block.
[0116] Meanwhile, in the default merge mode, the merge candidate list may include at least one of spatial merging candidates, temporal merging candidates, non-adjacent spatial merging candidates, history-based merging candidates, pairwise average merging candidates, and zero motion vector merging candidates. The merge candidates included in the merge candidate list may be arranged in the aforementioned order.
[0117] Meanwhile, if a merge candidate is at the Lth position in the merge candidate list, it may mean that its merge index is L-1. Therefore, a merge candidate located earlier in the merge candidate list may have a lower merge index. Here, L is an integer greater than or equal to 1.
[0118] Meanwhile, in this specification, the merge candidate at the Lth position of the merge candidate list,
[0119] The Lth merge candidate in the merge candidate list may have the same meaning. For example, the merge candidate at position 1 (merge index 0) in the merge candidate list may have the same meaning as the first merge candidate in the merge candidate list.
[0120] Meanwhile, the maximum number of merge candidates N that can be included in the merge candidate list of the current block can be derived from the sequence parameter set (SPS). In this case, the merge candidate list of the current block can include at most N merge candidates, where N is any positive integer.
[0121] Meanwhile, the maximum number of merge candidates can be determined by the promise of the encoder / decoder.
[0122] Alternatively, the maximum number of merge candidates may be determined at a higher level, such as a video parameter set (VPS), a sequence parameter set, a picture parameter set (PPS), a picture header (PH), a slice header (SH), or at a lower level, such as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, the maximum number of merge candidates may be determined explicitly or implicitly.
[0123] Meanwhile, MaxNumMergeCand can mean the maximum number of merge candidates.
[0124]
[0125] The pair-average merge candidate derivation method involves deriving a pair-average merge candidate and then adding it to the merge candidate list. Here, a pair-average merge candidate is a merge candidate derived using the motion information of two merge candidates included in the merge candidate list. The motion vector of the pair-average merge candidate is determined as the average of the motion vectors of the two merge candidates.
[0126] The pairwise average merge candidate derivation method can be performed when the merge candidate list has not yet reached the maximum number of merge candidates after being filled with spatial merge candidates, temporal merge candidates, non-adjacent spatial merge candidates, and history-based merge candidates. In other words, it is a method for providing additional candidates when the merge candidate list is not fully filled.
[0127] FIG. 4 is a diagram illustrating a method for deriving pair-average merge candidates according to one embodiment of the present invention. For example, FIG. 4 assumes that the maximum number of merge candidates is 6.
[0128] In Fig. 4, the merge index of the merge candidate list (400) indicates the index of each merge candidate. The maximum number of merge candidates is 6, and the merge index can have values from 0 to 5.
[0129] In Fig. 4, L0 represents a motion vector in the L0 direction, and L1 represents a motion vector in the L1 direction. Accordingly, the motion vector in the L0 direction of the first merge candidate (merge index 0) of the merge candidate list (400) of Fig. 4 is A0, and the motion vector in the L1 direction is A1. In addition, the motion vector in the L0 direction of the second merge candidate (merge index 1) of the merge candidate list (400) of Fig. 4 is B0, and the motion vector in the L1 direction is B1.
[0130] In Fig. 4, even after the merge candidate list is filled with spatial merge candidates, temporal merge candidates, non-adjacent spatial merge candidates, and history-based merge candidates, the number of merge candidates in the merge candidate list (400) is 4. In this case, since the number of merge candidates in the merge candidate list does not reach the maximum number of merge candidates, a pair-average merge candidate derivation method can be performed.
[0131] First, a pairwise average merge candidate can be derived by calculating the average of the first and second merge candidates (Averaging, 401).
[0132] The motion information of the pairwise averaging candidate (402) in FIG. 4 refers to the motion information of the pairwise averaging candidate derived through averaging calculation. Here, the motion information (402) may include a motion vector (MV) of the pairwise averaging candidate and a reference picture index (refIdx) of the pairwise averaging candidate.
[0133] Meanwhile, the motion vector of the pair average merge candidate can be determined as the average value of the motion vector of the first merge candidate and the motion vector of the second merge candidate. Referring to FIG. 4, the motion vector in the L0 direction of the pair average merge candidate is determined as the average value of the motion vector (A0) in the L0 direction of the first merge candidate in the merge candidate list (400) and the motion vector (B0) in the L0 direction of the second merge candidate. In addition, the motion vector in the L1 direction of the pair average merge candidate is determined as the average value of the motion vector (A1) in the L1 direction of the first merge candidate in the merge candidate list (400) and the motion vector (B1) in the L1 direction of the second merge candidate.
[0134] Meanwhile, the reference picture of the pair average merge candidate may be determined to be the same as the first merge candidate in the merge candidate list. Referring to FIG. 4, the reference picture indices (refIdx) in the L0 direction and L1 direction of the pair average merge candidate may be determined to be the reference picture index of the first merge candidate in the merge candidate list (400).
[0135] And, once a pair average merge candidate is derived, the derived pair average merge candidate can be added to the merge candidate list. Referring to Fig. 4, the pair average merge candidate can be added (Adding, 403) to the fifth position of the merge candidate list (400).
[0136] And, a prediction block of the current block can be generated based on the merge candidate list with the pair average merge candidate added.
[0137] Meanwhile, referring to FIG. 4, both the first merge candidate and the second merge candidate in the merge candidate list (400) have motion information in the L0 direction and the L1 direction, but this is just one example, and the merge candidates may only have motion information in a specific direction (L0 direction or L1 direction). In this case, if the two candidates have motion information in the same direction, the motion information of the pair average merge candidate can be derived through an average calculation. On the other hand, if only one candidate has motion information in a specific direction, the motion information of the pair average merge candidate in the corresponding direction can be derived by using that information as is.
[0138]
[0139] The aforementioned pairwise average merge candidate derivation method uses the first and second merge candidates in the merge candidate list to derive pairwise average merge candidates. Since this method uses merge candidates in fixed positions, it may not provide optimal merge candidates, resulting in lower prediction accuracy.
[0140] Hereinafter, an embodiment of an improved pair average merge candidate derivation method will be described.
[0141] In the improved pair-average merge candidate derivation method, the pair-average merge candidate can be derived based on merge candidates selected in a random manner, rather than being derived based on the first merge candidate and the second merge candidate in the merge candidate list.
[0142] According to one embodiment of the present invention, in an improved pair-average merge candidate derivation method, a pair-average merge candidate can be derived based on two specific merge candidates included in a merge candidate list. Specifically, the pair-average merge candidate can be derived based on two specific merge candidates from among a spatial merge candidate, a temporal merge candidate, a non-adjacent spatial merge candidate, and a history-based merge candidate included in the merge candidate list. For example, the pair-average merge candidate can be derived using only two spatial merge candidates included in the merge candidate list. As another example, the pair-average merge candidate can be derived using only two temporal merge candidates included in the merge candidate list.
[0143] According to another embodiment of the present invention, a pair-average merge candidate can be derived by selecting a merge candidate at any position in the merge candidate list without any constraints on a specific merge candidate. That is, rather than necessarily deriving a pair-average merge candidate based on the first and second merge candidates in the merge candidate list, a pair-average merge candidate can be derived based on two merge candidates having any merge index in the merge candidate list.
[0144] According to another embodiment of the present invention, the positions of pair-average merge candidates added to the merge candidate list may be restricted. Specifically, the improved pair-average merge candidate derivation method may restrict the positions where pair-average merge candidates can be added to the merge candidate list up to the Mth position. Here, M is an integer greater than or equal to 0, and may be less than or equal to a value obtained by subtracting 1 from the maximum number of merge candidates N. Through this, the number of pair-average merge candidates can be variably adjusted.
[0145] Specifically, if the merge candidate list is not fully filled with merge candidates, pairwise average merge candidates can be filled up to the limited Mth position. That is, pairwise average merge candidates can be added up to the merge index M-1 in the merge candidate list.
[0146] Meanwhile, pairwise average merge candidates can be added to the merge candidate list until the number of merge candidates in the merge candidate list becomes M.
[0147] Meanwhile, multiple pairwise average merge candidates may be added to the merge candidate list. In this case, pairwise average merge candidates may be added until the merge candidate list is completely filled, or until the Mth position described above is filled with merge candidates.
[0148] Meanwhile, if the merge candidate list is not fully filled even when the improved pairwise average merge candidate derivation method is used, or if the pairwise average merge candidate is not filled up to a limited position, the remaining merge candidate list can be filled using the zero-motion vector merge candidate.
[0149] Meanwhile, the aforementioned merge index M-1 can be determined by the promise of the encoder / decoder.
[0150] Alternatively, M-1 may be determined at a higher level, such as a video parameter set (VPS), a sequence parameter set, a picture parameter set (PPS), a picture header (PH), a slice header (SH), or at a lower level, such as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, M-1 may be explicitly determined or implicitly determined.
[0151] Meanwhile, M-1 can be determined independently from the maximum number of merge candidates.
[0152] According to another embodiment of the present invention, a pair average merge candidate may be added to a merge candidate list if it has motion information that does not overlap with other merge candidates in the merge candidate list.
[0153] Here, overlapping motion information may mean that the motion vector and reference picture of the pair average merge candidate are identical to the motion vector and reference picture of at least one other merge candidate in the merge candidate list.
[0154] Meanwhile, the encoder / decoder can perform a redundancy check on the pair average merge candidate and other merge candidates, and then add the pair average merge candidate with non-redundant motion information to the merge candidate list.
[0155] Meanwhile, the improved pairwise average merge candidate derivation method can also be applied to the fusion merge candidate derivation method described below. For example, a fusion merge candidate can be added to the merge candidate list if it has motion information that does not overlap with other merge candidates in the merge candidate list. In another example, multiple fusion merge candidates can be added to the merge candidate list. In another example, fusion merge candidates can be added to the merge candidate list until the number of merge candidates in the merge candidate list reaches M. In another example, multiple fusion merge candidates can be added to the merge candidate list.
[0156]
[0157] Conventional pairwise average merge candidate derivation methods derive pairwise average merge candidates by calculating the average of two merge candidates in a merge candidate list. This method uses a limited number of merge candidates and derives motion information solely through average calculations. Consequently, it fails to provide optimal merge candidates, potentially resulting in lower prediction accuracy.
[0158] Hereinafter, an embodiment of a method for deriving a fusion merge candidate will be described. The fusion merge candidate derivation method refers to a method of deriving a fusion merge candidate and then adding the candidate to a merge candidate list. Here, a fusion merge candidate may refer to a merge candidate derived by merging the motion information of at least two or more merge candidates included in an existing merge candidate list.
[0159] The method for deriving fusion merge candidates can be performed when the maximum number of merge candidates is not yet reached even after the merge candidate list is filled with spatial merge candidates, temporal merge candidates, non-adjacent spatial merge candidates, and history-based merge candidates. In other words, this is a method for providing additional candidates when the merge candidate list is not fully filled.
[0160] Meanwhile, in this specification, the fusion of motion information of merge candidates and the weighted sum of motion information of merge candidates may have the same meaning.
[0161] Meanwhile, in this specification, the terms "fused merge candidate" and "weighted merging candidate" may have the same meaning. Furthermore, if three or more existing merge candidates are derived through merging, the fused merge candidate and the multi-weighted merging candidate may have the same meaning.
[0162] FIG. 5 is a diagram illustrating a method for deriving fusion merge candidates according to one embodiment of the present invention. For example, FIG. 5 assumes that the maximum number of merge candidates is 6, and that 3 merge candidates are selected and fused.
[0163] In Fig. 5, the merge index of the merge candidate list (500) indicates the index of each merge candidate. The maximum number of merge candidates is 6, and the merge index can have values from 0 to 5.
[0164] In Fig. 5, even after the merge candidate list is filled with spatial merge candidates, temporal merge candidates, non-adjacent spatial merge candidates, and history-based merge candidates, the number of merge candidates in the merge candidate list (500) is 4. Since the number of merge candidates in the merge candidate list does not reach the maximum number of merge candidates, a method for deriving fused merge candidates can be performed.
[0165] Referring to Fig. 5, L0 in the merge candidate list (500) means a motion vector in the L0 direction, and L1 means a motion vector in the L1 direction. Therefore, the motion vector in the L0 direction of the first merge candidate (merge index 0) in the merge candidate list (500) is C 1,L0 , and the motion vector in the L1 direction is C 1,L1 Likewise, the motion vector in the L0 direction of the second merge candidate (merge index 1) of the merge candidate list (500) is C 2,L0 , and the motion vector in the L1 direction is C 2,L1 And, the motion vector in the L0 direction of the third merge candidate (merge index 2) of the merge candidate list (500) is C 3,L0 , and the motion vector in the L1 direction is C 3,L1 am.
[0166] First, to derive fusion merge candidates, K merge candidates can be selected from the merge candidate list. Here, K is a positive integer greater than or equal to 2.
[0167] Referring to FIG. 5, a first merge candidate, a second merge candidate, and a third merge candidate are selected, and a fusion merge candidate can be derived through fusion (Fusion, 501) of the merge candidates.
[0168] The motion information of the fusion merge candidate (502) in FIG. 5 refers to the motion information of the derived fusion merge candidate. Here, the motion information (502) may include a motion vector (MV) of the fusion merge candidate and a reference picture index (refIdx) of the fusion merge candidate.
[0169] Meanwhile, the motion vector of the fused merge candidate can be determined by weighting the motion vectors of the selected merge candidates. The weighted sum can be performed by applying weights to the motion vectors of the selected merge candidates. For example, the weighted sum can be performed in the manner shown in Equation 1.
[0170]
[0171]
[0172] Mathematical expression 1 assumes that three merge candidates are selected from the merge candidate list.
[0173] In mathematical expression 1, MV L0 refers to the motion vector in the L0 direction of the fusion merge candidate, and MV L1 refers to the motion vector in the L1 direction of the fusion merge candidate. In Fig. 5, the motion vectors in the L0 and L1 directions correspond to the motion information (502) of the fusion merge candidate.
[0174] And, C 1,L0 , C 2,L0 and C 3,L0 means the motion vector in the L0 direction of each of the merge candidates selected for deriving the fusion merge candidate. In Fig. 5, the merge candidate list
[0175] The motion vectors in the L0 direction of the merge candidates selected from (500) correspond to this.
[0176] And, W 1,L0 , W 2,L0 and W 3,L0represents the weight value of each motion vector in the L0 direction. Specifically, W 1,L0 is C 1,L0 It means the weighted value of W 2,L0 is C 2,L0 It means the weighted value of W 3,L0 is C 3,L0 It means the weighted value of W 1,L0 , W 2,L0 and W 3,L0 The sum is 1, and each weight value can have a value greater than or equal to 0 and less than or equal to 1.
[0177] And, C 1,L1 , C 2,L1 and C 3,L1 refers to the motion vector in the L1 direction of each of the merge candidates selected for deriving the fusion merge candidate. In Fig. 5, the motion vector in the L1 direction of the merge candidates selected from the merge candidate list (500) corresponds to this.
[0178] And, W 1,L1 , W 2,L1 and W 3,L1 means the weighted value of the motion vector in the L1 direction. Specifically, W 1,L1 is C 1,L1 It means the weighted value of W 2,L1 is C 2,L1 It means the weighted value of W 3,L1 is C 3,L1 It means the weighted value of W 1,L1 , W 2,L1 and W 3,L1 The sum is 1, and each weight value can have a value greater than or equal to 0 and less than or equal to 1.
[0179] Meanwhile, the weights of motion vectors can be determined by several methods.
[0180] According to one embodiment of the present invention, weights can be determined based on the magnitudes of motion vectors. For example, based on the sum of the motion vector magnitudes of all selected merge candidates, a lower weight can be assigned to a specific merge candidate with a larger motion vector, and conversely, a higher weight can be assigned to a smaller motion vector. This allows for the ratio of each motion vector to be reflected, thereby setting an appropriate weight.
[0181] Meanwhile, the magnitude of the motion vector can mean the absolute value of the motion vector.
[0182] According to another embodiment of the present invention, the weight value may be determined as a predetermined constant. For example, the weight value may be 1 / K. Here, K may represent the number of merge candidates selected for deriving fusion merge candidates.
[0183] According to another embodiment of the present invention, the weight value can be determined in the encoder and transmitted to the decoder.
[0184] According to another embodiment of the present invention, a specific weight value may be determined as 0. For example, W in Equation 1 1,L0 , W 2,L0 and W 3,L0 At least one of them can be determined as 0. In addition, W in mathematical expression 1 1,L1 , W 2,L1 and W 3,L1 At least one of them can be determined to be 0.
[0185] Meanwhile, the weight value of the motion vector may be determined through any one of the weight value determination methods described above, or may be determined through a combination of the methods described above.
[0186] Meanwhile, if the aforementioned weight determination method is used to determine the weight of the motion vector in the L0 direction, the weight of the motion vector in the L1 direction can also be determined to be the same as the weight of the motion vector in the L0 direction. Conversely, if the weight determination method is used to determine the weight of the motion vector in the L1 direction, the weight of the motion vector in the L0 direction can also be determined to be the same as the weight of the motion vector in the L1 direction. In this case, the weight of the motion vector in the L0 direction and the weight of the motion vector in the L1 direction can be determined through a single weight determination method.
[0187] Alternatively, the weights of the motion vectors in the L0 direction and the weights of the motion vectors in the L1 direction can be determined independently. In this case, the weights of the motion vectors in the L0 direction and the weights of the motion vectors in the L1 direction can be determined in different ways.
[0188] And, once a fusion merge candidate is derived, the derived fusion merge candidate can be added to the merge candidate list. Referring to Fig. 5, the fusion merge candidate can be added (Adding, 503) to the fifth position of the merge candidate list (500).
[0189] And, a prediction block of the current block can be generated based on the merge candidate list to which the fusion merge candidate has been added.
[0190] Meanwhile, in FIG. 5, the reference picture (refIdx) of the fusion merge candidate is determined to be the same as the first merge candidate in the merge candidate list, but this is only an example, and the reference picture of the fusion merge candidate may be determined to be the same as the reference picture of any one of the merge candidates selected to induce the fusion merge.
[0191] Meanwhile, a fused merge candidate can be derived by selecting merge candidates with the same reference picture from the merge candidate list. In this case, the reference picture of the fused merge candidate may be identical to the reference pictures of the selected merge candidates.
[0192]
[0193] Meanwhile, in FIG. 5, the fusion merge candidate is added (503) to the fifth position of the merge candidate list (500), but this is only an example, and the fusion merge candidate may be added to any position of the merge candidate list.
[0194] Meanwhile, in FIG. 5, one fusion merge candidate is added to the merge candidate list (500), but this is only one example, and multiple fusion merge candidates may be added to any position in the merge candidate list.
[0195] Meanwhile, in mathematical expression 1 and FIG. 5, the motion vectors of three merge candidates are selected to derive a fused merge candidate, but this is only one example, and K merge candidates may be selected from among the merge candidates included in the merge candidate list. Here, K is an integer greater than or equal to 2, and may be less than or equal to the number of merge candidates included in the merge candidate list.
[0196] Meanwhile, K can be determined by the promise of the encoder / decoder.
[0197] Alternatively, K may be determined at a higher level, such as a video parameter set (VPS), a sequence parameter set, a picture parameter set (PPS), a picture header (PH), a slice header (SH), or at a lower level, such as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, K may be explicitly determined or implicitly determined.
[0198] Meanwhile, the method for deriving fusion merge candidates may not be performed when the number of merge candidates included in the merge candidate list is 1.
[0199] Meanwhile, the merge candidate list before the fusion merge candidate is added can be called the first merge candidate list, and the merge candidate list after the fusion merge candidate is added can be called the second merge candidate list.
[0200]
[0201] Meanwhile, in mathematical expression 1 and FIG. 5, three consecutive merge candidates from the first merge candidate are selected among the merge candidates included in the merge candidate list, but this is only one example, and K merge candidates at any position among the merge candidates included in the merge candidate list can be selected by various methods.
[0202] According to one embodiment of the present invention, K merge candidates at arbitrary locations can be selected based on a cost value. Here, the cost value may refer to a template matching cost. Specifically, the merge candidates may be selected based on the similarity between the template of the reference block represented by each merge candidate included in the merge candidate list and the template of the current block. Here, a high similarity means that the distortion between the template of the current block and the template of the reference block represented by the merge candidate is small, which may mean that the template matching cost of the corresponding merge candidate is small.
[0203] For example, among the merge candidates included in the merge candidate list, K merge candidates at random positions can be selected in order of decreasing template matching cost.
[0204] As another example, K merge candidates can be selected based on the merge candidate with the smallest template matching cost among the merge candidates included in the merge candidate list. Here, the template matching cost of the merge candidate with the smallest template matching cost can be referred to as the minimum template matching cost.
[0205] Specifically, if the template matching cost of a specific merge candidate is less than the minimum template matching cost multiplied by R, the merge candidate can be selected as a merge candidate for deriving a fusion merge candidate, where R is an arbitrary positive real number.
[0206] According to another embodiment of the present invention, K merge candidates at arbitrary positions can be selected based on the distance between the current picture and the reference picture of the merge candidate.
[0207] For example, K merge candidates at arbitrary positions in the merge candidate list can be selected based on the picture order count (POC). Here, POC refers to a value that assigns a number by listing the order of pictures in chronological order. Therefore, a reference picture with a POC value smaller than the POC value of the current picture may indicate a picture that is relatively past the current picture, and a reference picture with a POC value larger than the POC value of the current picture may indicate a picture that is relatively future the current picture. Therefore, the POC difference may indicate the distance between pictures.
[0208] And, the merge candidate can be selected based on the absolute value of the POC difference between the reference picture of the merge candidate included in the merge candidate list and the current picture.
[0209] Alternatively, among the merge candidates included in the merge candidate list, K merge candidates may be selected in descending order of the absolute value of the POC difference between the reference picture and the current picture.
[0210] Alternatively, a merge candidate whose absolute value of the POC difference between the reference picture and the current picture is the smallest may be selected based on the merge candidate whose absolute value of the POC difference between the reference picture and the current picture is less than or equal to a predetermined threshold.
[0211] According to another embodiment of the present invention, K merge candidates at arbitrary positions may be selected based on the reference pictures of the merge candidates. Specifically, among the merge candidates in the merge candidate list, K merge candidates having the same reference pictures may be selected.
[0212] For example, four merge candidates having the same reference picture can be selected from among the merge candidates included in the merge candidate list. Then, a fused merge candidate can be derived based on the four selected merge candidates.
[0213]
[0214] According to one embodiment of the present invention, before adding a fusion merge candidate, the merge candidate list may be reordered. Specifically, the merge candidates in the merge candidate list may be reordered from highest to lowest priority, and merge candidates for deriving a fusion merge candidate may be selected from the reordered merge candidate list. In this case, a higher priority may mean matching a lower index in the merge candidate list.
[0215] Meanwhile, the method for reordering merge candidates in the merge candidate list may be an adaptive reordering of merge candidates with template matching (ARMC-TM). The adaptive reordering of merge candidates with template matching refers to a method for reordering merge candidates in the merge candidate list based on the similarity between the template of the reference block represented by each merge candidate and the template of the current block.
[0216] For example, merge candidates can be reordered so that a merge candidate corresponding to a template of a reference block with a high degree of similarity to the template of the current block has a high priority in the merge candidate list. In this case, a high degree of similarity may indicate a small distortion between the template of the current block and the template of the reference block indicated by the merge candidate. Meanwhile, the distortion measurement method may be any of various methods, such as the sum of absolute difference (SAD), the sum of square error (SSE), and the sum of absolute transformed differences (SATD).
[0217] Meanwhile, the reordering of merge candidates in the merge candidate list can be performed on a subgroup basis. This can reduce computational complexity.
[0218]
[0219] According to one embodiment of the present invention, a method for deriving a fusion merge candidate in a decoder may additionally include a step of obtaining information regarding whether the fusion merge candidate derivation method is performed.
[0220] At this time, the information regarding whether the fusion merge candidate derivation method is performed is information indicating whether the fusion merge candidate derivation method is performed to generate a merge candidate list. In addition, the decoder can perform the fusion merge candidate derivation method only if the information indicates whether the fusion merge candidate derivation method is performed.
[0221] Meanwhile, information regarding whether the fusion merge candidate derivation method is performed may be an activation flag of the fusion merge candidate derivation method. In this case, if the flag is 1, the decoder can perform the fusion merge candidate derivation method.
[0222] Meanwhile, information regarding whether the fusion merge candidate derivation method is performed may be a deactivation flag of the fusion merge candidate derivation method. In this case, if the flag is 0, the decoder can perform the fusion merge candidate derivation method.
[0223] Meanwhile, information on whether a method for deriving fusion merge candidates is performed can be conditionally signaled based on information on whether a merge mode is used.
[0224] Specifically, information about whether a fusion merge candidate derivation method is performed may be signaled only if information about whether merge mode is used indicates that merge mode is used.
[0225] For example, it can be implemented with a syntax transmission / parsing structure as in Table 1. In Table 1, information regarding whether merge mode is used can be general_merge_flag. In addition, the activation flag for the fusion merge candidate derivation method can be fusion_merge_flag.
[0226]
[0227]
[0228] Meanwhile, the syntax of information on whether the method for deriving fusion merge candidates in Table 1 is performed and information on whether the merge mode is used can be implemented through a specific syntax transmission / parsing structure, as an example, and the implementation method can be defined in various forms.
[0229] Meanwhile, information regarding whether the fusion merge candidate derivation method is performed may be influenced by the constraint flag of the fusion merge candidate derivation method. Specifically, if the constraint flag is enabled, the fusion merge candidate derivation method is not performed, and conversely, if the constraint flag is disabled, the fusion merge candidate derivation method may be performed.
[0230] For example, if the constraint flag is 1, the activation flag of the fusion merge candidate derivation method can be determined as 0. Then, the decoder does not perform the fusion merge candidate derivation method.
[0231] As another example, if the constraint flag is 1, the disable flag of the fusion merge candidate derivation method can be determined to be 1. Then, the decoder does not perform the fusion merge candidate derivation method.
[0232] Meanwhile, information on whether a method for deriving fusion merge candidates is performed can be signaled at a higher level, such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), and a slice header (SH), or can be signaled at a lower level, such as a coding unit (CU), a prediction unit (PU), and a transform unit (TU).
[0233] Meanwhile, the fusion merge candidate derivation method can be performed as a replacement for the pairwise average merge candidate derivation method. Alternatively, the fusion merge candidate derivation method can be performed independently from the pairwise average merge candidate derivation method to add the fusion merge candidate to the merge candidate list. Alternatively, even when pairwise average merge candidates are added to the merge candidate list, if the number of merge candidates included in the merge candidate list is less than the maximum number of merge candidates, the fusion merge candidate can be added to the merge candidate list.
[0234] Meanwhile, information about the positions of the K merge candidates selected for fusion merge induction within the merge candidate list can be signaled dependently on information about whether the fusion merge candidate induction method is performed.
[0235]
[0236] Conventional merge modes perform predictions based on a single merge candidate selected from a list of merge candidates. This method directly uses the motion information of the selected merge candidate to predict the current block, which can result in lower prediction accuracy.
[0237] Hereinafter, an embodiment of a prediction block fusion method will be described. Here, the prediction block fusion method may refer to a method in which multiple initial prediction blocks are weighted and combined to generate a prediction block for the current block. Furthermore, the initial prediction block may refer to a prediction block derived based on each of multiple merge candidates selected from a merge candidate list.
[0238] First, a list of merge candidates for the current block can be generated.
[0239] According to one embodiment of the present invention, the merge candidate list may mean a merge candidate list generated in a general merge mode.
[0240] According to another embodiment of the present invention, a merge candidate list can be generated based on the improved pairwise average merge candidate derivation method described above.
[0241] According to another embodiment of the present invention, a merge candidate list can be generated based on the aforementioned fusion merge candidate derivation method.
[0242] Meanwhile, the current block's merge candidate list can be reordered. Specifically, the merge candidates in the merge candidate list can be reordered from highest to lowest priority, and multiple merge candidates can be selected from the reordered merge candidate list.
[0243] In this case, a high priority could mean matching a lower numbered index in the merge candidate list.
[0244] Meanwhile, the method for reordering merge candidates in the merge candidate list may be an adaptive reordering of merge candidates with template matching (ARMC-TM). The adaptive reordering of merge candidates with template matching refers to a method for reordering merge candidates in the merge candidate list based on the similarity between the template of the reference block represented by each merge candidate and the template of the current block.
[0245] For example, merge candidates can be reordered so that a merge candidate corresponding to a template of a reference block with a high degree of similarity to the template of the current block has a higher priority in the merge candidate list. In this case, a high degree of similarity may indicate a small distortion between the template of the current block and the template of the reference block represented by the merge candidate.
[0246] Meanwhile, the method for measuring distortion can be any one of various methods, such as the SAD method, the SSE method, and the SATD method.
[0247] Meanwhile, the reordering of merge candidates can be performed on a subgroup basis. This can reduce computational complexity.
[0248] And, Q merge candidates can be selected from the merge candidate list. Here, Q is an integer greater than or equal to 2, and can be less than or equal to the number of merge candidates included in the merge candidate list. When Q is 1, one merge candidate is selected, as in the existing merge mode, and therefore is not considered in the prediction block fusion method.
[0249] Meanwhile, Q can be determined by the promise of the encoder / decoder.
[0250] Alternatively, Q may be determined at a higher level, such as a video parameter set (VPS), a sequence parameter set, a picture parameter set (PPS), a picture header (PH), a slice header (SH), or at a lower level, such as a coding unit (CU), a prediction unit (PU), or a transform unit (TU). In this case, Q may be explicitly determined or implicitly determined.
[0251] According to one embodiment of the present invention, Q merge candidates at any position in a merge candidate list may be selected based on a cost value. Here, the cost value may refer to a template matching cost. Specifically, the merge candidates may be selected based on the similarity between the template of a reference block represented by each merge candidate included in the merge candidate list and the template of a current block. Here, a high similarity means that the distortion between the template of the current block and the template of the reference block represented by the merge candidate is small, which may mean that the template matching cost of the corresponding merge candidate is small.
[0252] For example, among the merge candidates included in the merge candidate list, Q merge candidates can be selected in order of smallest template matching cost.
[0253] According to another embodiment of the present invention, Q merge candidates at arbitrary positions in the merge candidate list can be selected based on the distance from the current block. Here, the distance can mean a spatial distance.
[0254] For example, among the merge candidates included in the merge candidate list, Q merge candidates can be selected in order of distance from the current block.
[0255] According to another embodiment of the present invention, Q merge candidates at any position in the merge candidate list can be selected based on the magnitude of the motion vector.
[0256] For example, among the merge candidates included in the merge candidate list, Q merge candidates can be selected in order of decreasing motion vector size.
[0257] According to another embodiment of the present invention, Q merge candidates at arbitrary positions can be selected based on the distance between the current picture and the reference picture of the merge candidate and the current picture.
[0258] For example, Q merge candidates at any position in the merge candidate list can be selected based on the absolute value of the POC (picture order count) difference.
[0259] Meanwhile, among the merge candidates included in the merge candidate list, Q merge candidates can be selected in the order of the smallest absolute value of the POC difference between the reference picture and the current block.
[0260] According to another embodiment of the present invention, Q merge candidates at arbitrary positions may be selected based on the reference pictures of the merge candidates. Specifically, among the merge candidates in the merge candidate list, Q merge candidates having the same reference pictures may be selected.
[0261]
[0262] Then, a merge candidate list consisting of Q merge candidates selected from the merge candidate list can be generated. Here, the merge candidate list consisting of Q merge candidates can be referred to as a merge candidate list for prediction block fusion.
[0263] Additionally, a prediction block can be generated based on each merge candidate included in the merge candidate list for prediction block fusion. Here, the prediction block generated based on each merge candidate can be referred to as an initial prediction block.
[0264] Additionally, the prediction block of the current block can be generated through a weighted sum of the initial prediction blocks. Here, the prediction block of the current block generated through the weighted sum of the initial prediction blocks can be referred to as the final prediction block.
[0265] Meanwhile, the weighted sum can be performed in the same manner as in mathematical expression 2.
[0266]
[0267]
[0268] In Equation 2, Pred represents the predicted block of the current block generated through weighted sum. And, Pred_B i (i=1, …, Q) denote the initial prediction blocks generated based on each merge candidate. And, W i (i=1, …, Q) represents the weight value of each initial prediction block.
[0269] Meanwhile, the weights of the initial prediction blocks can be determined by several methods.
[0270] According to one embodiment of the present invention, weights may be determined based on the template matching cost of the merge candidate. For example, based on the sum of the template matching costs of all selected merge candidates, a lower weight may be assigned to the initial prediction block generated based on a particular merge candidate if the template matching cost of the candidate is higher. Conversely, a lower cost may be assigned a higher weight. This allows the prediction performance of each merge candidate to be reflected, thereby setting an appropriate weight.
[0271] According to another embodiment of the present invention, the weight value may be determined as a constant. For example, the weight value may be 1 / Q, where Q is the number of merge candidates included in the merge candidate list for predictive block fusion and is an integer greater than or equal to 2.
[0272] According to another embodiment of the present invention, the weight value can be determined in the encoder and transmitted to the decoder.
[0273] According to another embodiment of the present invention, a specific weight value may be determined as 0. For example, W in Equation 1 i At least one of (i=1, …, Q) can be determined to be 0.
[0274] Meanwhile, the weight value may be determined through any one of the weight value determination methods described above, or may be determined through a combination of the methods described above.
[0275]
[0276] According to one embodiment of the present invention, the prediction block fusion method in the decoder may additionally include a step of obtaining information regarding whether the prediction block fusion method is performed.
[0277] At this time, the information regarding whether the prediction block fusion method is performed is information indicating whether the prediction block fusion method is performed to generate a prediction block of the current block. Furthermore, the decoder can perform the prediction block fusion method only if the information indicates that the prediction block fusion method is performed.
[0278] Meanwhile, information regarding whether the prediction block fusion method is performed may be an activation flag of the prediction block fusion method. In this case, if the flag is 1, the decoder can perform the prediction block fusion method.
[0279] Meanwhile, information regarding whether the prediction block fusion method is performed may be a deactivation flag for the prediction block fusion method. In this case, if the flag is 0, the decoder can perform the prediction block fusion method.
[0280] Meanwhile, information on whether the prediction block fusion method is performed can be conditionally signaled based on information on whether the merge mode is used.
[0281] Specifically, information about whether a prediction block fusion method is performed may be signaled only if information about whether merge mode is used indicates that merge mode is used.
[0282] For example, it can be implemented with a syntax transmission / parsing structure as in Table 2. In Table 2, information regarding whether merge mode is used can be general_merge_flag. In addition, the activation flag for the prediction block fusion method can be prediction_fusion_flag.
[0283]
[0284]
[0285] Meanwhile, the syntax of information on whether the prediction block fusion method of Table 2 is performed and information on whether the merge mode is used can be implemented through a specific syntax transmission / parsing structure, as an example, and the implementation method can be defined in various forms.
[0286] Meanwhile, information regarding whether or not the prediction block fusion method is performed may be influenced by the constraint flag of the prediction block fusion method. Specifically, if the constraint flag is enabled, the prediction block fusion method is not performed, and conversely, if the constraint flag is disabled, the prediction block fusion method may be performed.
[0287] For example, if the constraint flag is 1, the activation flag of the prediction block fusion method can be determined as 0. Then, the decoder does not perform the prediction block fusion method.
[0288] As another example, if the constraint flag is 1, the disable flag of the prediction block fusion method can be determined to be 1. Then, the decoder does not perform the prediction block fusion method.
[0289] Meanwhile, information on whether or not a prediction block fusion method is performed can be signaled at a higher level, such as a sequence parameter set (SPS), a picture parameter set (PPS), a picture header (PH), and a slice header (SH), or can be signaled at a lower level, such as a coding unit (CU), a prediction unit (PU), and a transform unit (TU).
[0290] Meanwhile, information about the positions of the Q selected merge candidates within the merge candidate list for generating the merge candidate list for prediction block fusion can be signaled dependently on information about whether the prediction block fusion method is performed.
[0291]
[0292] Figure 6 is a flowchart illustrating an image decoding method according to one embodiment of the present invention. Specifically, Figure 6 is a flowchart illustrating a prediction method based on deriving fusion merge candidates in the image decoding method. The prediction method based on deriving fusion merge candidates of Figure 6 can be performed by a decoding device.
[0293] The video decoding device can generate a merge candidate list based on at least one of motion information of surrounding blocks of the current block and motion information of a block decoded before the current block (S600).
[0294] Meanwhile, the step of generating the merge candidate list can be performed by reordering the merge candidates in the merge candidate list.
[0295] Meanwhile, when the merge candidates in the merge candidate list are reordered, they may be reordered based on the similarity between the template of the reference block indicated by the merge candidate and the template of the current block.
[0296] And, if the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a fusion merge candidate can be added to the merge candidate list (S610).
[0297] Meanwhile, the above-mentioned fusion merge candidate may be added to the merge candidate list if it has motion information that does not overlap with the merge candidates in the merge candidate list.
[0298] And, based on the merge candidate list, a prediction block of the current block can be generated (S620).
[0299] Meanwhile, the above fusion merge candidate can be derived based on at least two merge candidates selected from the merge candidate list.
[0300] Meanwhile, the above fusion merge candidate can be added to the merge candidate list until the number of merge candidates in the merge candidate list reaches a predefined number.
[0301] Meanwhile, the position of the merge candidate selected from the merge candidate list within the merge candidate list may be any position.
[0302] Meanwhile, the motion vector of the above fusion merge candidate can be determined by weighting the motion vectors of the merge candidates selected from the merge candidate list.
[0303] Meanwhile, the weight of the motion vector can be determined based on the size of the motion vector.
[0304] Meanwhile, the merge candidate selected from the merge candidate list may be selected based on the similarity between the template of the reference block indicated by the merge candidate in the merge candidate list and the template of the current block.
[0305] Meanwhile, the merge candidate selected from the merge candidate list may be selected based on the distance between the reference picture of the merge candidate in the merge candidate list and the current picture.
[0306] Meanwhile, a merge candidate selected from the merge candidate list may be a merge candidate having the same reference picture among the merge candidates in the merge candidate list.
[0307] Meanwhile, the step of generating the prediction block may include a step of generating a merge candidate list for prediction block fusion by selecting at least two merge candidates from among the merge candidate list, a step of generating initial prediction blocks based on each of the merge candidates in the merge candidate list for prediction block fusion, and a step of generating a final prediction block of the current block by weighting the initial prediction blocks.
[0308] Meanwhile, the step of generating a merge candidate list for the prediction block fusion can be performed based on the similarity between the template of the reference block indicated by the merge candidate of the merge candidate list and the template of the current block.
[0309] Meanwhile, the step of generating a merge candidate list for the prediction block fusion may be performed based on the distance between the merge candidate of the merge candidate list and the current block.
[0310] Meanwhile, the step of generating a merge candidate list for the above prediction block fusion can be performed based on the distance between the reference picture of the merge candidate of the merge candidate list and the current picture.
[0311] Meanwhile, the step of generating a merge candidate list for the above prediction block fusion can be performed by selecting a merge candidate having the same reference picture among the merge candidates included in the merge candidate list.
[0312] Meanwhile, the weights of the initial prediction blocks can be determined based on the similarity between the template of the reference block indicated by the merge candidate in the merge candidate list for the prediction block fusion and the template of the current block.
[0313] Meanwhile, the step of generating a prediction block of the current block may include a step of rearranging the merge candidate list based on the similarity between the template of the reference block indicated by the merge candidate of the merge candidate list and the template of the current block.
[0314] Meanwhile, the steps described in FIG. 6 can be performed in the same manner in an image encoding method. Furthermore, a bitstream can be generated by an image encoding method including the steps described in FIG. 6. The bitstream can be stored on a non-transitory computer-readable recording medium and can also be transmitted (or streamed).
[0315]
[0316] FIG. 7 is a drawing exemplarily showing a content streaming system to which an embodiment according to the present invention can be applied.
[0317] As illustrated in FIG. 7, a content streaming system to which an embodiment of the present invention is applied may largely include an encoding server, a streaming server, a web server, a media storage, a user device, and a multimedia input device.
[0318] The encoding server compresses content input from multimedia input devices such as smartphones, cameras, and CCTVs into digital data, generates a bitstream, and transmits it to the streaming server. Alternatively, if multimedia input devices such as smartphones, cameras, and CCTVs directly generate bitstreams, the encoding server may be omitted.
[0319] The above bitstream can be generated by an image encoding method and / or an image encoding device to which an embodiment of the present invention is applied, and the streaming server can temporarily store the bitstream during the process of transmitting or receiving the bitstream.
[0320] The streaming server transmits multimedia data to a user device based on a user request via a web server, and the web server can act as an intermediary to inform the user of available services. When a user requests a desired service from the web server, the web server transmits the request to the streaming server, and the streaming server can transmit multimedia data to the user. At this time, the content streaming system may include a separate control server, and in this case, the control server may control commands / responses between each device within the content streaming system.
[0321] The streaming server can receive content from a media repository and / or encoding server. For example, when receiving content from the encoding server, the content can be received in real time. In this case, to provide a smooth streaming service, the streaming server can store the bitstream for a certain period of time.
[0322] Examples of the user devices may include mobile phones, smart phones, laptop computers, digital broadcasting terminals, personal digital assistants (PDAs), portable multimedia players (PMPs), navigation devices, slate PCs, tablet PCs, ultrabooks, wearable devices (e.g., smartwatches, smart glasses, HMDs), digital TVs, desktop computers, digital signage, etc.
[0323] Each server within the above content streaming system can be operated as a distributed server, in which case data received from each server can be processed in a distributed manner.
[0324]
[0325] The above embodiments can be performed in the same or corresponding manner in an encoding device and a decoding device. In addition, an image can be encoded / decoded using at least one or a combination of at least one of the above embodiments.
[0326] The order in which the above embodiments are applied may be different in the encoding device and the decoding device. Alternatively, the order in which the above embodiments are applied may be the same in the encoding device and the decoding device.
[0327] The above embodiments can be performed for each of the luminance and chrominance signals. Alternatively, the above embodiments can be performed identically for the luminance and chrominance signals.
[0328] In the above embodiments, the methods are described based on a flowchart as a series of steps or units. However, the present invention is not limited to the order of the steps, and some steps may occur in a different order or simultaneously with other steps described above. Furthermore, those skilled in the art will understand that the steps depicted in the flowchart are not exclusive, and that other steps may be included, or one or more steps in the flowchart may be deleted without affecting the scope of the present invention.
[0329] The above embodiments may be implemented in the form of program commands that can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program commands, data files, data structures, etc., either singly or in combination. The program commands recorded on the computer-readable recording medium may be those specifically designed and constructed for the present invention, or may be known and usable by those skilled in the art of computer software.
[0330] The bitstream generated by the encoding method according to the above embodiment can be stored in a non-transitory computer-readable recording medium. In addition, the bitstream stored in the non-transitory computer-readable recording medium can be decoded by the decoding method according to the above embodiment.
[0331] Here, examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specifically configured to store and execute program instructions such as ROMs, RAMs, and flash memories. Examples of program instructions include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices may be configured to operate as one or more software modules to perform processing according to the present invention, and vice versa.
[0332] Although the present invention has been described above with specific details such as specific components and limited examples and drawings, these are provided only to help a more general understanding of the present invention, and the present invention is not limited to the above examples, and those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and variations from this description.
[0333] Therefore, the idea of the present invention should not be limited to the embodiments described above, and all things that are modified equally or equivalently to the following claims as well as the claims are considered to fall within the scope of the idea of the present invention.
[0334] The present invention can be used in a device for encoding / decoding an image and a recording medium storing a bitstream.
Claims
1. In the video decryption method, A step of generating a merge candidate list based on at least one of motion information of surrounding blocks of the current block and motion information of a block decrypted before the current block; If the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a step of adding a fusion merge candidate to the merge candidate list; and A step of generating a prediction block of the current block based on the merge candidate list, The above fusion merge candidate is derived based on at least two merge candidates selected from the merge candidate list, A video decoding method, characterized in that the above fusion merge candidate is added to the merge candidate list until the number of merge candidates in the merge candidate list reaches a predefined number.
2. In paragraph 1, A video decoding method, characterized in that the above fusion merge candidate is added to the merge candidate list if it has motion information that does not overlap with merge candidates in the merge candidate list.
3. In paragraph 1, An image decoding method, characterized in that the position of a merge candidate selected from the merge candidate list is an arbitrary position within the merge candidate list.
4. In paragraph 1, An image decoding method, characterized in that the motion vector of the above fusion merge candidate is determined by weighting the motion vectors of the merge candidates selected from the merge candidate list.
5. In paragraph 4, An image decoding method, characterized in that the weight of the motion vector is determined based on the size of the motion vector.
6. In paragraph 1, A method for decoding an image, characterized in that the step of generating the above merge candidate list is performed by rearranging the merge candidates in the above merge candidate list.
7. In paragraph 6, A video decoding method characterized in that when the merge candidates in the merge candidate list are reordered, the reordering is based on the similarity between the template of the reference block indicated by the merge candidate and the template of the current block.
8. In paragraph 1, A video decoding method, characterized in that a merge candidate selected from the merge candidate list is selected based on the similarity between the template of the reference block indicated by the merge candidate in the merge candidate list and the template of the current block.
9. In paragraph 1, A video decoding method, characterized in that a merge candidate selected from the merge candidate list is selected based on the distance between a reference picture of the merge candidate in the merge candidate list and the current picture.
10. In paragraph 1, A video decoding method, characterized in that a merge candidate selected from the above merge candidate list is a merge candidate having the same reference picture among the merge candidates in the above merge candidate list.
11. In the video encoding method, A step of generating a merge candidate list based on at least one of motion information of surrounding blocks of a current block and motion information of a block encoded before the current block; If the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a step of adding a fusion merge candidate to the merge candidate list; and A step of generating a prediction block of the current block based on the merge candidate list, The above fusion merge candidate is derived based on at least two merge candidates selected from the merge candidate list, A video encoding method, characterized in that the above fusion merge candidate is added to the merge candidate list until the number of merge candidates in the merge candidate list reaches a predefined number.
12. In a non-transitory computer-readable recording medium storing a bitstream generated by an image encoding method, The above image encoding method is, A step of generating a merge candidate list based on at least one of motion information of surrounding blocks of a current block and motion information of a block encoded before the current block; If the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a step of adding a fusion merge candidate to the merge candidate list; and A step of generating a prediction block of the current block based on the merge candidate list, The above fusion merge candidate is derived based on at least two merge candidates selected from the merge candidate list, A recording medium, characterized in that the above fusion merge candidate is added to the merge candidate list until the number of merge candidates in the merge candidate list reaches a predefined number.
13. In a method for transmitting a bitstream generated by a video encoding method, The above transmission method includes a step of transmitting the bitstream, The above image encoding method is, A step of generating a merge candidate list based on at least one of motion information of surrounding blocks of a current block and motion information of a block encoded before the current block; If the number of merge candidates in the merge candidate list is less than the maximum number of merge candidates, a step of adding a fusion merge candidate to the merge candidate list; and A step of generating a prediction block of the current block based on the merge candidate list, The above fusion merge candidate is derived based on at least two merge candidates selected from the merge candidate list, A transmission method, characterized in that the above fusion merge candidate is added to the merge candidate list until the number of merge candidates in the merge candidate list reaches a predefined number.
Citation Information
Patent Citations
Reference Picture Interpolation Method and Apparatus and Video Coding Method and Apparatus Using Same
KR101678968B1
Ultraviolet ray irrdiation apparatus and method of manufacturing a semiconductor package using the apparatus
KR1020210028910A
Vehicle door latch apparatus
KR1020250024319A
Method and apparatus for encoding / decoding image and recording medium for storing bitstream
KR102438181B1
Video signal encoding / decoding method, and device therefor
WO2020130714A1