Image encoding / decoding method and recording medium for the method

By using a combined merge candidate list and multi-directional prediction in video encoding/decoding, the problems of low efficiency and hardware complexity in traditional merging modes are solved, achieving more efficient motion compensation and simplified hardware logic design.

CN116614638BActive Publication Date: 2026-08-25INTELLECTUAL DISCOVERY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310761221.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-07-12
Filing Date
2017-07-12
Publication Date
2026-08-25
Estimated Expiration
2037-07-12

AI Technical Summary

Technical Problem

Existing technologies suffer from low encoding efficiency, increased memory access bandwidth, and complex hardware logic during the encoding/decoding of high-resolution and high-quality images. In particular, in motion compensation using traditional merging modes, there are dependencies and non-parallelism between merging candidate derivation processes.

Method used

A combined list of merging candidates is adopted, including spatial merging candidates, temporal merging candidates, modified spatial merging candidates, and modified temporal merging candidates. Motion compensation is performed through unidirectional prediction, bidirectional prediction, tridirectional prediction, and quadridirectional prediction. The merging candidate derivation process is parallelized, the dependencies between merging candidate derivation processes are removed, and the hardware logic is simplified.

Benefits of technology

It improves video encoding/decoding efficiency, increases throughput in merged mode, simplifies hardware logic structure, and reduces memory access bandwidth requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116614638B_ABST
    Figure CN116614638B_ABST
Patent Text Reader

Abstract

The present invention relates to an image encoding / decoding method and a recording medium therefor. The image decoding method can include the steps of generating a merge candidate list of a current block, wherein the merge candidate list of the current block includes at least one merge candidate among merge candidates respectively corresponding to a plurality of reference picture lists; determining at least one piece of motion information by using the merge candidate; and generating a prediction block of the current block by using the determined at least one piece of motion information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of invention application No. 201780043622.4, filed on July 12, 2017, entitled "Image Encoding / Decoding Method and Recording Medium for the Method". Technical Field

[0002] This invention relates to a method and apparatus for encoding / decoding video. More specifically, this invention relates to a method and apparatus for performing motion compensation by using a merging pattern. Background Technology

[0003] Recently, the demand for high-resolution and high-quality images, such as high-definition (HD) and ultra-high-definition (UHD) images, has grown across various application areas. However, the data volume of higher-resolution and higher-quality image data increases compared to traditional image data. Therefore, the costs of transmission and storage increase when transmitting image data using media such as traditional wired and wireless broadband networks, or when storing image data using traditional storage media. To address these challenges arising from the increasing resolution and quality of image data, efficient image encoding / decoding technologies are needed for higher-resolution and higher-quality images.

[0004] Image compression techniques encompass a variety of methods, including: inter-frame prediction techniques that predict pixel values ​​included in the current frame from previous or subsequent frames; intra-frame prediction techniques that predict pixel values ​​included in the current frame using pixel information from the current frame; energy transformation and quantization techniques for compressing residual signals; entropy coding techniques that assign short codes to high-frequency values ​​and long codes to low-frequency values; and so on. By using such image compression techniques, image data can be effectively compressed and transmitted or stored.

[0005] In motion compensation using the traditional merging pattern, only spatial merging candidates, temporal merging candidates, bidirectional prediction merging candidates, and zero merging candidates are added to the list of merging candidates to be used. Therefore, only unidirectional and bidirectional prediction are used, which limits the improvement of coding efficiency.

[0006] In motion compensation using the traditional merging pattern, there are limitations in throughput due to the dependency between temporal merging candidate derivation and bidirectional prediction merging candidate derivation. Furthermore, merging candidate derivation cannot be performed in parallel.

[0007] In motion compensation using the traditional merging mode, bidirectional predictive merging candidates, generated through bidirectional predictive merging candidate derivation, are used as motion information. Therefore, compared to unidirectional predictive merging candidates, memory access bandwidth is increased during motion compensation.

[0008] In motion compensation using the traditional merging mode, zero-merging candidate derivation is performed differently depending on the stripe type, thus complicating the hardware logic. Furthermore, generating bidirectional predictive zero-merging candidates through bidirectional predictive zero-merging candidate derivation processing, which is used in motion compensation, increases memory access bandwidth. Summary of the Invention

[0009] Technical issues

[0010] The objective of this invention is to provide a method and apparatus for improving the encoding / decoding efficiency of video by using a combination of merged candidates to perform motion compensation.

[0011] Another objective of the present invention is to provide a method and apparatus for improving the encoding / decoding efficiency of video by performing motion compensation through one-way prediction, two-way prediction, three-way prediction and four-way prediction.

[0012] Another objective of the present invention is to provide a method and apparatus for determining motion information by parallelizing the merging candidate derivation process, removing the dependency between merging candidate derivation processes, bidirectionally predicting merging candidate segmentation, and unidirectionally predicting zero merging candidate derivation, thereby increasing the throughput of the merging mode and simplifying the hardware logic.

[0013] Technical solution

[0014] A method for decoding video according to the present invention includes: generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists; determining at least one piece of motion information by using the merge candidate list; and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0015] In a method for decoding video, the merge candidate list may include at least one of spatial merge candidates derived from spatially neighboring blocks of the current block, temporal merge candidates derived from co-located blocks of the current block, modified spatial merge candidates derived by modifying spatial merge candidates, modified temporal merge candidates derived by modifying temporal merge candidates, and merge candidates having predefined motion information values.

[0016] In a method for decoding video, the merge candidate list may also include merge candidates derived by using a combination of at least two merge candidates selected from a group including spatial merge candidates, temporal merge candidates, modified spatial merge candidates, and modified temporal merge candidates.

[0017] In methods for decoding video, spatial merge candidates can be derived from sub-blocks of neighboring blocks adjacent to the current block, and temporal merge candidates can be derived from sub-blocks of co-located blocks of the current block.

[0018] In a method for decoding video, the step of generating a prediction block for the current block using the determined at least one piece of motion information may include: generating a plurality of temporal prediction blocks based on an inter-frame prediction indicator for the current block; and generating a prediction block for the current block by applying at least one of a weighting factor and an offset to the generated plurality of temporal prediction blocks.

[0019] In a method for decoding video, at least one of a weighting factor and an offset may be shared in blocks smaller than a predetermined block or in blocks deeper than the predetermined block.

[0020] In methods for decoding video, a list of merge candidates may be shared in blocks smaller than a predetermined block or in blocks deeper than the predetermined block.

[0021] In a method for decoding video, when the size of the current block is smaller than a predetermined block or the depth of the current block is greater than the predetermined block, a list of merging candidates can be generated based on the higher-level blocks of the current block, wherein the size or depth of the higher-level blocks is equal to the size or depth of the predetermined block.

[0022] A method for encoding video according to the present invention includes: generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists; determining at least one piece of motion information by using the merge candidate list; and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0023] In a method for encoding video, the merge candidate list may include at least one of spatial merge candidates derived from spatially neighboring blocks of the current block, temporal merge candidates derived from co-located blocks of the current block, modified spatial merge candidates derived by modifying spatial merge candidates, modified temporal merge candidates derived by modifying temporal merge candidates, and merge candidates having predetermined motion information values.

[0024] In a method for encoding video, the merge candidate list may also include merge candidates derived by using a combination of at least two merge candidates selected from a group including spatial merge candidates, temporal merge candidates, modified spatial merge candidates, and modified temporal merge candidates.

[0025] In methods for encoding video, spatial merge candidates can be derived from sub-blocks of neighboring blocks adjacent to the current block, and temporal merge candidates can be derived from sub-blocks of co-located blocks of the current block.

[0026] In a method for encoding video, the step of generating a prediction block for the current block using the determined at least one piece of motion information may include: generating a plurality of temporal prediction blocks based on an inter-frame prediction indicator for the current block; and generating a prediction block for the current block by applying at least one of a weighting factor and an offset to the generated plurality of temporal prediction blocks.

[0027] In a method for encoding video, at least one of a weighting factor and an offset may be shared in blocks smaller than a predetermined block or in blocks deeper than the predetermined block.

[0028] In methods for encoding video, a list of merge candidates can be shared in blocks smaller than a predetermined block or in blocks deeper than the predetermined block.

[0029] In a method for encoding video, when the size of the current block is smaller than a predetermined block or the depth of the current block is greater than the predetermined block, a list of merge candidate blocks is generated based on the higher-level blocks of the current block, wherein the size or depth of the higher-level blocks is equal to the size or depth of the predetermined block.

[0030] An apparatus for decoding video according to the present invention, the apparatus comprising: an inter-frame prediction unit, generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists, determining at least one piece of motion information by using the merge candidate list, and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0031] An apparatus for encoding video according to the present invention, the apparatus comprising: an inter-frame prediction unit, generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists, determining at least one piece of motion information by using the merge candidate list, and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0032] A readable medium according to the invention stores a bitstream formed by a method for encoding video, the method comprising: generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists; determining at least one piece of motion information by using the merge candidate list; and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0033] Beneficial effects

[0034] In this invention, a method and apparatus are provided to improve the encoding / decoding efficiency of video by performing motion compensation through the combined merging of candidates.

[0035] This invention provides a method and apparatus for improving video encoding / decoding efficiency by performing motion compensation through one-way prediction, two-way prediction, three-way prediction, and four-way prediction.

[0036] In this invention, a method and apparatus are provided to perform motion compensation by parallelizing the merging candidate derivation process, removing the dependencies between merging candidate derivation processes, bidirectionally predicting merging candidate segmentation, and unidirectionally predicting zero merging candidate derivation, thereby increasing the throughput of the merging mode and simplifying the hardware logic. Attached Figure Description

[0037] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.

[0038] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.

[0039] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded.

[0040] Figure 4 This is a diagram showing the form of a prediction unit (PU) that can be included in a coding unit (CU).

[0041] Figure 5 This is a diagram showing the form of a transform unit (TU) that can be included in an encoding unit (CU).

[0042] Figure 6 This is a diagram illustrating an embodiment of the processing used to explain intra-frame prediction.

[0043] Figure 7 This is a diagram illustrating an embodiment of the processing used to explain inter-frame prediction.

[0044] Figure 8 It is a diagram used to interpret the transform set based on the intra-frame prediction mode.

[0045] Figure 9 This is a diagram used to explain the processing of transformations.

[0046] Figure 10 It is a diagram used to interpret the scanning of the transformation coefficients of quantization.

[0047] Figure 11It is a diagram used to explain block partitions.

[0048] Figure 12 This is a flowchart illustrating a method for encoding video using a merging mode according to the present invention.

[0049] Figure 13 This is a flowchart illustrating a method for decoding video using a merging mode according to the present invention.

[0050] Figure 14 This is a diagram illustrating an example of deriving spatial merge candidates for the current block.

[0051] Figure 15 This is a diagram illustrating an example of adding spatial merge candidates to the merge candidate list.

[0052] Figure 16 This is a diagram illustrating an embodiment of deriving and sharing space merging candidates in a CTU.

[0053] Figure 17 This is a diagram illustrating an example of deriving the time merge candidates for the current block.

[0054] Figure 18 This is a diagram illustrating an example of adding time merging candidates to the merge candidate list.

[0055] Figure 19 This is a diagram illustrating an example of scaling the motion vector of a co-located block to derive a temporal merging candidate for the current block.

[0056] Figure 20 This is a diagram showing the index of the combination.

[0057] Figure 21a and Figure 21b This is a diagram illustrating an embodiment of a method for deriving combined candidate combinations.

[0058] Figure 22 and Figure 23 This is a diagram illustrating an embodiment of deriving a combined merge candidate by using at least one of spatial merge candidates, temporal merge candidates, and zero merge candidates, and adding the combined merge candidate to the merge candidate list.

[0059] Figure 24 This is a diagram illustrating the advantages of deriving combined merging candidates by using only spatial merging candidates in motion compensation using merging modes.

[0060] Figure 25 This is a diagram illustrating an embodiment of a bidirectional predictive merge candidate method for split-and-merge merging.

[0061] Figure 26This is a diagram illustrating an embodiment of a method for deriving zero-merging candidates.

[0062] Figure 27 This is a diagram illustrating an example of adding derived zero merge candidates to the merge candidate list.

[0063] Figure 28 This is a diagram illustrating another embodiment of the method for deriving zero-merging candidates.

[0064] Figure 29 This is a diagram illustrating an embodiment of deriving and sharing a list of merge candidates in a CTU.

[0065] Figure 30 and Figure 31 This is a diagram illustrating an example of the syntax for information about motion compensation.

[0066] Figure 32 This is a diagram illustrating an embodiment of using a merge mode in a CTU where the size is smaller than a predetermined block.

[0067] Figure 33 This is a diagram illustrating a method for decoding video according to the present invention.

[0068] Figure 34 This is a diagram illustrating a method for encoding video according to the present invention. Detailed Implementation

[0069] Various modifications can be made to this invention, and various embodiments of the invention exist, wherein examples of the embodiments will now be provided with reference to the accompanying drawings, and examples of the embodiments will be described in detail. However, the invention is not limited thereto, although exemplary embodiments may be interpreted as including all modifications, equivalents, or substitutions within the technical concept and scope of the invention. Similar reference numerals refer to functions that are the same or similar in respect of each other. In the drawings, the shapes and sizes of elements may be exaggerated for clarity. In the following detailed description of the invention, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice this disclosure. It should be understood that the various embodiments of this disclosure, though different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics associated with one embodiment described herein may be implemented in other embodiments without departing from the spirit and scope of this disclosure. Furthermore, it should be understood that the positions or arrangements of the various elements within each disclosed embodiment may be modified without departing from the spirit and scope of this disclosure. Therefore, the following detailed description is not intended to be limiting, and the scope of this disclosure is defined only by the appended claims (and, where appropriate, the full scope of the equivalents claimed in the claims).

[0070] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components are not to be construed as limiting the terms. The terms are used only to distinguish one component from another. For example, without departing from the scope of the invention, a "first" component may be referred to as a "second" component, and a "second" component may similarly be referred to as a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.

[0071] It will be understood that in this specification, when an element is simply referred to as "connected to" or "joined to" another element rather than "directly connected to" or "directly joined to" another element, it can be "directly connected to" or "directly joined to" another element, or connected to or joined to another element with other elements inserted in between. Conversely, it should be understood that when an element is referred to as "directly joined" or "directly connected to" another element, there are no intermediate elements.

[0072] Furthermore, the components shown in the embodiments of the present invention are illustrated independently to present distinct functionalities. Therefore, this does not imply that each component is composed as a separate hardware or software unit. In other words, for convenience, each component includes every one of the enumerated components. Thus, at least two components in each component can be combined to form a single component, or a single component can be divided into multiple components to perform each function. Embodiments where each component is combined and embodiments where a component is divided are also included within the scope of the invention without departing from its spirit.

[0073] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include plural expressions unless they have a distinct meaning in the context. In this specification, it will be understood that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, quantities, steps, actions, elements, components, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, quantities, steps, actions, elements, components, or combinations thereof may be present or added. In other words, when a particular element is referred to as “comprising,” elements other than the corresponding element are not excluded; rather, additional elements may be included in embodiments of the invention or within the scope of the invention.

[0074] Furthermore, some components may not be essential for performing the necessary functions of the invention, but rather optional components that merely enhance its performance. The invention can be implemented by including only the essential components necessary for carrying out the invention, excluding components used to enhance performance. Structures that include only the essential components and exclude optional components used solely for enhancing performance are also included within the scope of the invention.

[0075] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing exemplary embodiments of the invention, well-known functions or structures will not be described in detail, as they would unnecessarily obscure the understanding of the invention. The same constituent elements in the drawings are denoted by the same reference numerals, and repeated descriptions of the same elements will be omitted.

[0076] Furthermore, in the following text, "image" may refer to a frame that constitutes a video, or it may refer to the video itself. For example, "encoding or decoding an image, or both" may mean "encoding or decoding a video, or both," and may also mean "encoding or decoding one of a plurality of images in a video, or both." Here, "frame" and "image" may have the same meaning.

[0077] Terminology Description

[0078] Encoder: can refer to a device that performs encoding.

[0079] Decoder: can refer to a device that performs decoding.

[0080] Explanation: This could refer to determining the value of a syntax element by performing entropy decoding, or it could refer to entropy decoding itself.

[0081] A block can refer to a sample of an M×N matrix. Here, M and N are positive integers, and a block can refer to a sample matrix in two-dimensional form.

[0082] Sample: A sample is the basic unit of a block and can indicate a value ranging from 0 to 2 Bd – 1 depending on the bit depth (Bd). In this invention, a sample can refer to a pixel.

[0083] A unit can refer to a unit used for encoding and decoding an image. During image encoding and decoding, a unit can be a region created by partitioning an image. Furthermore, a unit can refer to a sub-partition unit when an image is partitioned into multiple sub-partition units during encoding or decoding. During image encoding and decoding, predetermined processing can be performed for each unit. A unit can be partitioned into sub-units smaller than the unit's size. Depending on its function, a unit can refer to a block, macroblock, coding tree unit, coding tree block, coding unit, coding block, prediction unit, prediction block, transform unit, transform block, etc. Furthermore, to distinguish a unit from a block, a unit can include a luma component block, a chroma component block of the luma component block, and syntax elements for each chroma component block. Units can have various sizes and shapes; specifically, the shape of a unit can be a two-dimensional geometric shape, such as a rectangle, square, trapezoid, triangle, pentagon, etc. Additionally, unit information can include at least one of the following: unit type (indicating coding unit, prediction unit, transform unit, etc.), unit size, unit depth, and the order in which the unit is encoded and decoded.

[0084] Reconstructing neighboring units: This can refer to reconstructed units that have been previously encoded or decoded, and which are spatially / temporally adjacent to the target unit being encoded / decoded. Here, reconstructing neighboring units can refer to reconstructing neighboring blocks.

[0085] Neighboring block: Refers to a block adjacent to the target block being encoded / decoded. A block adjacent to the target block can be a block with a boundary that contacts the target block. A neighboring block can also be a block located at an adjacent vertex of the target block. A neighboring block can also refer to a reconstructed neighboring block.

[0086] Cell depth: This refers to the degree to which cells are partitioned. In a tree structure, the root node can be the highest node, and the leaf nodes can be the lowest nodes.

[0087] Symbols: can refer to the syntax elements, encoding parameters, transform coefficients, etc. of the encoding / decoding target unit.

[0088] Parameter set: This refers to the header information in the structure of a bitstream. A parameter set can include at least one parameter set from a video parameter set, sequence parameter set, frame parameter set, or adaptive parameter set. Additionally, a parameter set can refer to slice header information and tile header information, etc.

[0089] Bitstream: can refer to a string of bits that includes encoded image information.

[0090] Prediction unit: This refers to the basic unit used when performing inter-frame or intra-frame prediction and compensation for prediction. A prediction unit can be partitioned into multiple partitions. In this case, each of the multiple partitions can be a basic unit when performing prediction and compensation, and each partition obtained from the prediction unit partitioning can be a prediction unit. Furthermore, a prediction unit can be partitioned into multiple smaller prediction units. Prediction units can have various sizes and shapes, and specifically, the shape of a prediction unit can be a two-dimensional geometric figure, such as a rectangle, square, trapezoid, triangle, pentagon, etc.

[0091] Prediction cell partitioning: can refer to the shape of the prediction cells partitioned.

[0092] Reference frame list: This can refer to a list that includes at least one reference frame, wherein the at least one reference frame is used for inter-frame prediction or motion compensation. The reference frame list can be of the following types: List Combined (LC), List 0 (L0), List 1 (L1), List 2 (L2), List 3 (L3), etc. At least one reference frame list can be used for inter-frame prediction.

[0093] Inter-frame prediction indicator: can refer to one of the following: the inter-frame prediction direction (unidirectional prediction, bidirectional prediction, etc.) of the encoded / decoded target block in the case of inter-frame prediction, the number of reference frames used to generate prediction blocks through the encoded / decoded target block, and the number of reference blocks used to perform inter-frame prediction or motion compensation through the encoded / decoded target block.

[0094] Reference screen index: can refer to the index of a specific reference screen in the reference screen list.

[0095] Reference frame: This refers to the frame that a specific unit references for inter-frame prediction or motion compensation. A reference image can be called a reference frame.

[0096] Motion vector: A two-dimensional vector used for inter-frame prediction or motion compensation, and can refer to the offset between the encoded / decoded target frame and the reference frame. For example, (mvX, mvY) can indicate a motion vector, where mvX indicates the horizontal component and mvY indicates the vertical component.

[0097] Motion vector candidate: can refer to the cell that becomes a prediction candidate when predicting motion vectors, or it can refer to the motion vector of that cell.

[0098] Motion vector candidate list: can refer to a list configured by using motion vector candidates.

[0099] Motion vector candidate index: This refers to an indicator that points to a motion vector candidate in the motion vector candidate list. The motion vector candidate index can also be referred to as the index of the motion vector predictor.

[0100] Motion information: may refer to motion vectors, reference frame indexes and inter-frame prediction indicators, as well as information including at least one of the following: reference frame list information, reference frames, motion vector candidates, motion vector candidate indexes, etc.

[0101] Merge candidate list: can refer to a list configured by using merge candidates.

[0102] Merging candidates can include spatial merging candidates, temporal merging candidates, combined merging candidates, combined bidirectional prediction merging candidates, zero merging candidates, etc. Merging candidates can include motion information such as prediction type information, reference frame indexes for each list, motion vectors, etc.

[0103] Merge Index: This can indicate information about merge candidates in the merge candidate list. Furthermore, the merge index can indicate a deduced merge candidate among reconstructed blocks that are spatially / temporally adjacent to the current block. Additionally, the merge index can indicate at least one of multiple motion information entries for a merge candidate.

[0104] Transform unit: This refers to the basic unit used when performing transformations, inverse transformations, quantization, dequantization, and encoding / decoding of transform coefficients on a residual signal. A transform unit can be divided into multiple smaller transform units. Transform units can have various sizes and shapes. Specifically, the shape of a transform unit can be a two-dimensional geometric figure, such as a rectangle, square, trapezoid, triangle, pentagon, etc.

[0105] Scaling: This refers to the process of multiplying a factor by the levels of the transform coefficients, resulting in the transformation coefficients being generated. Scaling can also be called inverse quantization.

[0106] Quantization parameter: This refers to the value used during quantization and dequantization to scale the transform coefficient levels. Here, the quantization parameter can be a value mapped to the quantization step size.

[0107] Variable increment (Delta) quantization parameter: can refer to the difference between the quantization parameter of the encoding / decoding target unit and the predicted quantization parameter.

[0108] Scan: This can refer to a method of sorting the coefficients within a block or matrix. For example, the operation of sorting a two-dimensional matrix into a one-dimensional matrix can be called a scan, and the operation of sorting a one-dimensional matrix into a two-dimensional matrix can be called a scan or inverse scan.

[0109] Transformation coefficients: These refer to the coefficient values ​​generated after performing a transformation. In this invention, the quantized transformation coefficient levels (i.e., the transformation coefficients to which quantization has been applied) can be referred to as transformation coefficients.

[0110] Non-zero transform coefficients: can be a transform coefficient whose value is not 0, or can be a transform coefficient level whose value is not 0.

[0111] Quantization matrix: This refers to a matrix used in quantization and dequantization to improve the subject quality or object quality of an image. The quantization matrix can also be called a scaling list.

[0112] Quantization matrix coefficients: These refer to each element of the quantization matrix. Quantization matrix coefficients can also be called matrix coefficients.

[0113] Default matrix: can refer to a predefined quantization matrix that is defined in the encoder and decoder.

[0114] Non-default matrix: can refer to a quantization matrix sent / received by the user without being predefined in the encoder and decoder.

[0115] A coding tree unit can consist of one luminance component (Y) coding tree unit and two associated chrominance component (Cb, Cr) coding tree units. Each coding tree unit can be partitioned using at least one partitioning method (such as a quadtree, binary tree, etc.) to form sub-units such as coding units, prediction units, transform units, etc. The term "coding tree unit" can be used to refer to pixel blocks (i.e., processing units in the decoding / encoding process of an image, such as partitions of the input image).

[0116] Coding tree block: can be used as a term to indicate one of the Y coding tree unit, Cb coding tree unit, and Cr coding tree unit.

[0117] Figure 1 This is a block diagram illustrating the configuration of an encoding device according to an embodiment of the present invention.

[0118] Encoding device 100 can be a video encoding device or an image encoding device. Video may include one or more images. Encoding device 100 can encode one or more images of the video in chronological order.

[0119] Reference Figure 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switcher 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference frame buffer 190.

[0120] Encoding device 100 can encode the input frame in intra-frame mode, inter-frame mode, or both. Furthermore, encoding device 100 can generate a bitstream by encoding the input frame and can output the generated bitstream. When intra-frame mode is used as the prediction mode, switcher 115 can switch to intra-frame mode. When inter-frame mode is used as the prediction mode, switcher 115 can switch to inter-frame mode. Here, intra-frame mode can be referred to as intra-frame prediction mode, and inter-frame mode can be referred to as inter-frame prediction mode. Encoding device 100 can generate prediction blocks of input blocks of the input frame. Furthermore, after generating prediction blocks, encoding device 100 can encode the residual between the input block and the prediction block. The input frame can be referred to as the current image as the target of the current encoding. The input block can be referred to as the current block or as the encoding target block as the target of the current encoding.

[0121] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use the pixel values ​​of the previous coded blocks adjacent to the current block as reference pixels. The intra-frame prediction unit 120 can perform spatial prediction by using reference pixels and can generate prediction samples of the input block by using spatial prediction. Here, intra-frame prediction may refer to intra-frame prediction.

[0122] When the prediction mode is inter-frame mode, the motion prediction unit 111 can search for the region that best matches the input block from the reference frame during motion prediction processing, and can derive the motion vector by using the searched region. The reference frame can be stored in the reference frame buffer 190.

[0123] The motion compensation unit 112 can generate prediction blocks by performing motion compensation using motion vectors. Here, the motion vectors can be two-dimensional vectors used for inter-frame prediction. Furthermore, the motion vectors can indicate the offset between the current frame and the reference frame. Here, inter-frame prediction can refer to inter-frame prediction.

[0124] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate a prediction block by applying an interpolation filter to a portion of the reference frame. To perform inter-frame prediction or motion compensation based on the coding unit, the method used for motion prediction and compensation in the coding unit can be determined from among skip mode, merge mode, AMVP mode, and current frame reference mode. Inter-frame prediction or motion compensation can be performed according to each mode. Here, the current frame reference mode can refer to a prediction mode that uses a pre-constructed region of the current frame with the coding target block. To specify the pre-constructed region, a motion vector for the current frame reference mode can be defined. Whether the coding target block is encoded according to the current frame reference mode can be encoded using the reference frame index of the coding target block.

[0125] Subtractor 125 generates a residual block by using the difference between the input block and the prediction block. The residual block may be referred to as the residual signal.

[0126] Transformation unit 130 can generate transformation coefficients by transforming the residual block and can output the transformation coefficients. Here, the transformation coefficients can be coefficient values ​​generated by transforming the residual block. In transform skip mode, transformation unit 130 can skip the transformation of the residual block.

[0127] A quantized transformation coefficient level can be generated by applying quantization to the transformation coefficients. In the following, in embodiments of the invention, the quantized transformation coefficient level may be referred to as the transformation coefficient.

[0128] The quantization unit 140 can generate quantized transformation coefficient levels by quantizing the transformation coefficients according to quantization parameters, and can output the quantized transformation coefficient levels. Here, the quantization unit 140 can quantize the transformation coefficients using a quantization matrix.

[0129] The entropy coding unit 150 can generate a bitstream by performing entropy coding on values ​​calculated by the quantization unit 140 or on coding parameter values ​​calculated in the coding process according to a probability distribution, and can output the generated bitstream. The entropy coding unit 150 can perform entropy coding on information used for decoding the image, and can also perform entropy coding on information about the image's pixels. For example, the information used for decoding the image may include syntax elements, etc.

[0130] When entropy coding is applied, the size of the bitstream encoding the target symbol is reduced by allocating a small number of bits to symbols with high occurrence probabilities and a large number of bits to symbols with low occurrence probabilities. Therefore, the compression performance of image coding can be improved through entropy coding. For entropy coding, the entropy coding unit 150 can use coding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code / code (VLC) table. Furthermore, the entropy coding unit 150 can derive a binaryization method for the target symbol and a probability model for the target symbol / bits, and can subsequently perform arithmetic coding by using the derived binaryization method or the derived probability model.

[0131] To encode the transform coefficient levels, the entropy coding unit 150 can transform the coefficients from two-dimensional block form to one-dimensional vector form using a transform coefficient scanning method. For example, by scanning the coefficients of the block using an upper-right scan, the two-dimensional coefficients can be transformed into one-dimensional vectors. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan for scanning the coefficients in the two-dimensional block form along the column direction and a horizontal scan for scanning the coefficients in the two-dimensional block form along the row direction can be used instead of an upper-right scan. That is, based on the size of the transform unit and the intra-frame prediction mode, it can be determined which scanning method among the upper-right scan, vertical scan, and horizontal scan will be used.

[0132] Encoding parameters may include information such as syntax elements encoded by the encoder and sent to the decoder, and may include information that can be deduced during the encoding or decoding process. Encoding parameters may refer to information necessary for encoding or decoding an image. For example, encoding parameters may include at least one value or combination of the following: block size, block depth, block partitioning information, cell size, cell depth, cell partitioning information, quadtree partitioning flag, binary tree partitioning flag, binary tree partitioning direction, intra-frame prediction mode, intra-frame prediction direction, reference sample filtering method, prediction block boundary filtering method, filter taps, filter coefficients, inter-frame prediction mode, motion information, motion vectors, reference frame index, inter-frame prediction direction, inter-frame prediction indicator, reference frame list, motion vector prediction factor, motion vector candidate list, information on whether motion merging mode is used, motion merging candidates, motion merging candidate list, information on whether skip mode is used, and interpolation filter type. The information includes: motion vector magnitude, accuracy of motion vector representation, transform type, transform size, information on whether an additional (secondary) transform is used, information on the presence of residual signals, code block style, code block flags, quantization parameters, quantization matrix, filter information within the loop, information on whether filters are applied within the loop, filter coefficients within the loop, binaryization / debinding method, context model, context bits, bypass bits, transform coefficients, transform coefficient levels, transform coefficient level scanning method, image display / output order, stripe identification information, stripe type, stripe partition information, parallel block identification information, parallel block type, parallel block partition information, frame type, bit depth, and information on luminance or chrominance signals.

[0133] The residual signal can refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. A residual block can be the residual signal of a block unit.

[0134] When the encoding device 100 performs encoding using inter-frame prediction, the encoded current frame can be used as a reference frame for another image that will be processed subsequently. Therefore, the encoding device 100 can decode the encoded current frame and store the decoded image as a reference frame. To perform decoding, inverse quantization and inverse transform can be performed on the encoded current frame.

[0135] The quantized coefficients can be dequantized by the dequantization unit 160 and inverse transformed by the inverse transform unit 170. The dequantized and inverse transformed coefficients can be added to the prediction block by the adder 175, thereby generating the reconstructed block.

[0136] The reconstructed block can be processed by filter unit 180. Filter unit 180 can apply at least one of deblocking filter, sample adaptive offset (SAO), and adaptive loop filter (ALF) to the reconstructed block or reconstructed image. Filter unit 180 may be referred to as a loop filter.

[0137] Deblocking filters remove block distortion that occurs at the boundaries between blocks. To determine whether a deblocking filter is being applied, it can be determined based on the pixels included in several rows or columns within the block. When a deblocking filter is applied to a block, a strong or weak filter can be applied depending on the desired deblocking filter strength. Furthermore, horizontal and vertical filtering can be processed in parallel when applying a deblocking filter.

[0138] Sample-adaptive offset adds an optimal offset value to a pixel value to compensate for coding errors. Sample-adaptive offset corrects the offset between the deblocked image and the original image for each pixel. To perform offset correction on a specific image, one can use a method that considers the edge information of each pixel to apply the offset, or use the following method: divide the image's pixels into a predetermined number of regions, determine the regions to be offset corrected, and apply the offset correction to the determined regions.

[0139] An adaptive loop filter performs filtering based on values ​​obtained by comparing the reconstructed image with the original image. The pixels of the image can be partitioned into predetermined groups, a filter is determined for each group, and different filters can be performed for each group. Information regarding whether an adaptive loop filter is applied to the luminance signal can be sent for each coding unit (CU). The shape and filter coefficients of the adaptive loop filter applied to each block can vary. Furthermore, adaptive loop filters with the same form (fixed form) can be applied without considering the characteristics of the target block.

[0140] The reconstructed block after passing through the filter unit 180 can be stored in the reference frame buffer 190.

[0141] Figure 2 This is a block diagram illustrating the configuration of a decoding device according to an embodiment of the present invention.

[0142] Decoding device 200 can be a video decoding device or an image decoding device.

[0143] Reference Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference frame buffer 270.

[0144] Decoding device 200 can receive bitstreams output from encoding device 100. Decoding device 200 can decode the bitstreams in intra-frame mode or inter-frame mode. In addition, decoding device 100 can generate reconstructed images by performing decoding and can output the reconstructed images.

[0145] When the prediction mode used in decoding is intra-frame mode, the switcher can be switched to intra-frame mode. When the prediction mode used in decoding is inter-frame mode, the switcher can be switched to inter-frame mode.

[0146] Decoding device 200 can obtain reconstructed residual blocks from the input bitstream and can generate prediction blocks. When the reconstructed residual blocks and prediction blocks are obtained, decoding device 200 can generate a reconstructed block as the decoding target block by adding the reconstructed residual blocks and the prediction blocks. The decoding target block can be referred to as the current block.

[0147] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols with quantized transform coefficient levels. Here, the entropy decoding method may be similar to the entropy encoding method described above. For example, the entropy decoding method may be the inverse process of the entropy encoding method described above.

[0148] To decode the transform coefficient levels, the entropy decoding unit 210 can perform a transform coefficient scan, thereby transforming the coefficients from one-dimensional vector form to two-dimensional block form. For example, by scanning the coefficients of the block using an upper-right scan, the coefficients from one-dimensional vector form can be transformed into two-dimensional block form. Depending on the size of the transform unit and the intra-frame prediction mode, vertical and horizontal scans can be used instead of upper-right scans. That is, based on the size of the transform unit and the intra-frame prediction mode, it can be determined which scanning method among upper-right, vertical, and horizontal scans is used.

[0149] The quantized transform coefficient levels can be dequantized by dequantization unit 220 and inversely transformed by inverse transform unit 230. The quantized transform coefficient levels are dequantized and inversely transformed to generate a reconstruction residual block. Here, dequantization unit 220 can apply a quantization matrix to the quantized transform coefficient levels.

[0150] When the intra-frame mode is used, the intra-frame prediction unit 240 can generate a prediction block by performing spatial prediction, wherein the spatial prediction uses the pixel values ​​of the previous decoded block adjacent to the decoded target block.

[0151] When inter-frame mode is used, motion compensation unit 250 can generate prediction blocks by performing motion compensation, which uses both a reference frame stored in reference frame buffer 270 and motion vectors. When the value of the motion vector is not an integer, motion compensation unit 250 can generate prediction blocks by applying an interpolation filter to a portion of the reference frame. To perform motion compensation, based on the coding unit, it can be determined which method among skip mode, merge mode, AMVP mode, and current frame reference mode the motion compensation method of the prediction unit in the coding unit uses. Furthermore, motion compensation can be performed according to the mode. Here, current frame reference mode may refer to a prediction mode that uses a previously reconstructed region within the current frame having the decoding target block. The previously reconstructed region may not be adjacent to the decoding target block. To specify the previously reconstructed region, a fixed vector can be used for the current frame reference mode. In addition, a flag or index indicating whether the decoding target block is a block decoded according to the current frame reference mode can be transmitted by signaling and can be derived by using the reference frame index of the decoding target block. The current frame for the current frame reference mode can exist at a fixed position within the reference frame list for the decoded target block (e.g., the position with reference frame index 0 or the last position). Alternatively, the current frame can be variably located within the reference frame list; for this purpose, a reference frame index indicating the position of the current frame can be signaled. Here, signaling a flag or index can instruct the encoder to entropy-encode the corresponding flag or index and include it in the bitstream, and the decoder to entropy-decode the corresponding flag or index from the bitstream.

[0152] The reconstructed residual block and the prediction block can be added together by adder 255. The resulting block, obtained by adding the reconstructed residual block and the prediction block, can be passed through filter unit 260. Filter unit 260 can apply at least one of deblocking filter, sample adaptive offset, and adaptive loop filter to the reconstructed block or reconstructed frame. Filter unit 260 can output the reconstructed frame. The reconstructed frame can be stored in reference frame buffer 270 and can be used for inter-frame prediction.

[0153] Figure 3 It is a schematic diagram illustrating the partitioning structure of an image when it is encoded and decoded. Figure 3 An embodiment of dividing a cell into multiple sub-cells is illustrated schematically.

[0154] To effectively partition an image, coding units (CUs) can be used in encoding and decoding. Here, a coding unit can refer to a unit that is encoded. A unit can be a combination of 1) a syntax element and 2) a block that includes image samples. For example, "partitioning of a unit" can refer to "partitioning of the block associated with the unit." Block partitioning information can include information about the depth of the unit. Depth information can indicate the number of times the unit is partitioned or the degree to which the unit is partitioned, or both.

[0155] Reference Figure 3 Image 300 is sequentially partitioned for each maximum coding unit (LCU), and the partitioning structure is determined for each LCU. Here, LCU and coding tree unit (CTU) have the same meaning. A unit may have depth information based on a tree structure and may be hierarchically partitioned. Each sub-unit from a partition may have depth information. The depth information indicates the number of times the unit is partitioned or the degree to which the unit is partitioned, or both; therefore, the depth information may include information about the size of the sub-units.

[0156] The partitioning structure refers to the distribution of coding units (CUs) in the LCU 310. A CU can be a unit used for efficient encoding / decoding of an image. The distribution can be determined based on whether a CU will be partitioned multiple times (i.e., a positive integer equal to or greater than 2, including 2, 4, 8, 16, etc.). The width and height dimensions of the partitioned CUs can be half the width and half the height of the original CU, respectively. Alternatively, depending on the number of partitions, the width and height dimensions of the partitioned CUs can be smaller than the width and height dimensions of the original CU, respectively. The partitioned CUs can be recursively partitioned into multiple further partitioned CUs, wherein, following the same partitioning method, the further partitioned CUs have width and height dimensions smaller than those of the partitioned CUs.

[0157] Here, the partitioning of a CU can be performed recursively until a predetermined depth is reached. Depth information can be information indicating the size of the CU and can be stored for each CU. For example, the depth of an LCU can be 0, and the depth of a minimum coding unit (SCU) can be a predetermined maximum depth. Here, an LCU can be a coding unit with the aforementioned maximum size, and an SCU can be a coding unit with the minimum size.

[0158] Whenever LCU 310 begins to be partitioned, and the width and height dimensions of the CU are reduced through the partitioning operation, the depth of the CU increases by 1. In the case of a CU that cannot be partitioned, the CU can have a 2N×2N size for each depth. In the case of a CU that can be partitioned, a CU with a 2N×2N size can be partitioned into multiple CUs of N×N size. Each time the depth increases by 1, the size N is halved.

[0159] For example, when a coding unit is partitioned into four sub-coding units, the width and height of one of the four sub-coding units can be half the width and half the height of the original coding unit, respectively. For example, when a 32×32 coding unit is partitioned into four sub-coding units, each of the four sub-coding units can have a size of 16×16. When a coding unit is partitioned into four sub-coding units, the coding unit can be partitioned in a quadtree format.

[0160] For example, when a coding unit is partitioned into two sub-coding units, the width or height of one of the two sub-coding units can be half the width or half the height of the original coding unit, respectively. For example, when a 32×32 coding unit is vertically partitioned into two sub-coding units, each of the two sub-coding units can have a size of 16×32. For example, when a 32×32 coding unit is horizontally partitioned into two sub-coding units, each of the two sub-coding units can have a size of 32×16. When a coding unit is partitioned into two sub-coding units, the coding unit can be partitioned in a binary tree format.

[0161] Reference Figure 3 An LCU with a minimum depth of 0 can be 64×64 pixels, and an SCU with a maximum depth of 3 can be 8×8 pixels. Here, a CU with 64×64 pixels (i.e., LCU) can be represented by depth 0, a CU with 32×32 pixels can be represented by depth 1, a CU with 16×16 pixels can be represented by depth 2, and a CU with 8×8 pixels (i.e., SCU) can be represented by depth 3.

[0162] Furthermore, partition information of a CU can indicate whether or not a CU will be partitioned. Partition information can be 1 bit. Partition information can be included in all CUs except the SCU. For example, when the partition information value is 0, the CU may not be partitioned; when the partition information value is 1, the CU may be partitioned.

[0163] Figure 4 This is a diagram showing the form of a prediction unit (PU) that can be included in a coding unit (CU).

[0164] The CUs that are no longer to be partitioned from the LCU can be partitioned into at least one prediction unit (PU). This process can also be referred to as partitioning.

[0165] A PU can be the basic unit used for prediction. A PU can be encoded and decoded according to any of the skip mode, inter-frame mode, and intra-frame mode. A PU can be partitioned in various forms according to the said mode.

[0166] Furthermore, the coding unit may not be divided into multiple prediction units, and the coding unit and the prediction unit may have the same size.

[0167] like Figure 4 As shown, in skip mode, the CU may not be partitioned. In skip mode, a 2N×2N pattern 410 with the same size as the unpartitioned CU can be supported.

[0168] In inter-frame mode, the CU supports eight partition formats. For example, in inter-frame mode, it supports 2N×2N mode 410, 2N×N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440, and nR×2N mode 445. In intra-frame mode, it supports 2N×2N mode 410 and N×N mode 425.

[0169] A coding unit can be partitioned into one or more prediction units. A prediction unit can be partitioned into one or more sub-prediction units.

[0170] For example, when a prediction unit is partitioned into four sub-prediction units, the width and height of one of the four sub-prediction units can be half the width and half the height of the original prediction unit. For example, when a 32×32 prediction unit is partitioned into four sub-prediction units, each of the four sub-prediction units can have a size of 16×16. When a prediction unit is partitioned into four sub-prediction units, the prediction unit can be partitioned in a quadtree format.

[0171] For example, when a prediction unit is partitioned into two sub-prediction units, the width or height of one of the two sub-prediction units can be half the width or half the height of the original prediction unit. For example, when a 32×32 prediction unit is vertically partitioned into two sub-prediction units, each of the two sub-prediction units can have a size of 16×32. For example, when a 32×32 prediction unit is horizontally partitioned into two sub-prediction units, each of the two sub-prediction units can have a size of 32×16. When a prediction unit is partitioned into two sub-prediction units, the prediction unit can be partitioned in a binary tree format.

[0172] Figure 5 This is a diagram showing the form of a transform unit (TU) that can be included in an encoding unit (CU).

[0173] A transform unit (TU) can be a basic unit within a CU used for transforming, quantizing, inverse transforming, and dequantizing. A TU can have a square or rectangular shape, etc. A TU can be independently determined according to the size or form of the CU, or both.

[0174] The CUs that are no longer partitioned from the LCU can be partitioned into at least one TU. Here, the partitioning structure of the TU can be a quadtree structure. For example, as... Figure 5 As shown, a CU 510 can be partitioned once or more according to a quadtree structure. A CU being partitioned at least once is referred to as recursive partitioning. By partitioning, a CU 510 can be formed from TUs of different sizes. Optionally, a CU can be partitioned into at least one TU based on the number of vertical lines or horizontal lines used to partition the CU, or both. A CU can be partitioned into TUs that are symmetrical to each other, or it can be partitioned into TUs that are asymmetrical to each other. To partition a CU into symmetrical TUs, information about the size / shape of the TUs can be transmitted by signals and can be derived from the size / shape information of the CUs.

[0175] Furthermore, the coding unit may not be divided into transform units, and the coding unit and transform unit may have the same size.

[0176] A coding unit can be partitioned into at least one transform unit, and a transform unit can be partitioned into at least one sub-transform unit.

[0177] For example, when a transform unit is partitioned into four sub-transform units, the width and height of one of the four sub-transform units can be half the width and half the height of the original transform unit, respectively. For example, when a 32×32 transform unit is partitioned into four sub-transform units, each of the four sub-transform units can have a size of 16×16. When a transform unit is partitioned into four sub-transform units, the transform unit can be partitioned in a quadtree format.

[0178] For example, when a transform unit is partitioned into two sub-transform units, the width or height of one of the two sub-transform units can be half the width or half the height of the original transform unit, respectively. For example, when a 32×32 transform unit is vertically partitioned into two sub-transform units, each of the two sub-transform units can have a size of 16×32. For example, when a 32×32 transform unit is horizontally partitioned into two sub-transform units, each of the two sub-transform units can have a size of 32×16. When a transform unit is partitioned into two sub-transform units, the transform unit can be partitioned in a binary tree format.

[0179] When performing a transformation, the residual block can be transformed using at least one of a predetermined transformation method. For example, the predetermined transformation methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), KLT, etc. Which transformation method is applied to the residual block can be determined by using at least one of the following: inter-frame prediction mode information of the prediction unit, intra-frame prediction mode information of the prediction unit, and the size / shape of the transformed block. Information indicating the transformation method can be transmitted using a signal.

[0180] Figure 6 This is a diagram illustrating an embodiment of the processing used to explain intra-frame prediction.

[0181] Intra-frame prediction modes can be non-directional or directional. Non-directional modes can be DC or planar modes. Directional modes can be prediction modes with a specific direction or angle, and the number of directional modes can be M, equal to or greater than 1. A directional mode can be indicated by at least one of a mode number, a mode value, and a mode angle.

[0182] The number of intra-frame prediction modes can be equal to or greater than 1, N, including non-directional and directional modes.

[0183] The number of intra-prediction modes can vary depending on the block size. For example, when the block size is 4×4 or 8×8, the number of intra-prediction modes can be 67; when the block size is 16×16, the number of intra-prediction modes can be 35; when the block size is 32×32, the number of intra-prediction modes can be 19; and when the block size is 64×64, the number of intra-prediction modes can be 7.

[0184] The number of intra-prediction modes can be fixed at N, regardless of the block size. For example, the number of intra-prediction modes can be fixed at at least one of 35 or 67, regardless of the block size.

[0185] The number of intra-frame prediction modes can vary depending on the type of color component. For example, the number of prediction modes can vary depending on whether the color component is a luminance signal or a chrominance signal.

[0186] Intra-frame coding and / or decoding can be performed using sample values ​​or coding parameters included in the reconstructed neighboring blocks.

[0187] In order to encode / decode the current block according to intra-frame prediction, it is possible to identify whether samples included in the reconstructed neighboring blocks can be used as reference samples for encoding / decoding the target block. When there are samples that cannot be used as reference samples for encoding / decoding the target block, the sample values ​​are copied and / or interpolated to the samples that cannot be used as reference samples by using at least one of the samples included in the reconstructed neighboring blocks, thereby making the samples that cannot be used as reference samples usable as reference samples for encoding / decoding the target block.

[0188] In intra-frame prediction, filters can be applied to at least one of reference samples or prediction samples based on at least one of the intra-frame prediction mode and the size of the coded / decoded target block. Here, the coded / decoded target block can refer to the current block, and can refer to at least one of the coded block, prediction block, and transform block. The type of filter applied to the reference sample or prediction sample can vary depending on at least one of the intra-frame prediction mode or the size / shape of the current block. The type of filter can vary depending on at least one of the number of filter taps, filter coefficient values, or filter strength.

[0189] In the non-directional plane mode of intra-frame prediction mode, when generating a prediction block for an encoded / decoded target block, the sample value in the prediction block can be generated by using the weighted sum of the upper reference sample of the current sample, the left reference sample of the current sample, the upper right reference sample of the current block, and the lower left reference sample of the current block, based on the sample position.

[0190] In the non-directional DC mode of intra-frame prediction, when generating a prediction block for an encoded / decoded target block, the prediction block can be generated using the average of the upper reference sample and the left reference sample of the current block. Furthermore, filtering can be performed on one or more upper rows and one or more left columns of the encoded / decoded block adjacent to the reference sample using the reference sample values.

[0191] In the case of multiple orientation modes (angle modes) within intra-frame prediction modes, prediction blocks can be generated using upper-right reference samples and / or lower-left reference samples, and these multiple orientation modes can have different orientations. Real-valued interpolation can be performed to generate prediction sample values.

[0192] To perform intra-prediction, the intra-prediction mode of the current prediction block can be predicted from the intra-prediction modes of neighboring prediction blocks. When predicting the intra-prediction mode of the current prediction block using mode information predicted from neighboring intra-prediction modes, if the current prediction block and neighboring prediction blocks have the same intra-prediction mode, this information can be transmitted using predetermined flag information. If the intra-prediction mode of the current prediction block differs from that of neighboring prediction blocks, entropy coding can be performed to encode the intra-prediction mode information of the target block being encoded / decoded.

[0193] Figure 7 This is a diagram illustrating an embodiment of the processing used to explain inter-frame prediction.

[0194] Figure 7 The squares shown can indicate images (or screens). Furthermore, Figure 7 The arrows indicate the prediction direction. That is, an image can be encoded or decoded, or encoded and decoded, depending on the prediction direction. Based on the encoding type, each image can be classified as an I-frame (intra-frame), P-frame (one-way prediction frame), B-frame (two-way prediction frame), etc. Each frame can be encoded and decoded according to its own encoding type.

[0195] When the target image is an I-frame, the frame itself can be intra-coded without inter-frame prediction. When the target image is a P-frame, the image can be encoded using inter-frame prediction or motion compensation performed only on the forward reference frame. When the target image is a B-frame, the image can be encoded using inter-frame prediction or motion compensation performed on both the forward and backward reference frames. Alternatively, the image can be encoded using inter-frame prediction or motion compensation performed on either the forward or backward reference frame. Here, when inter-frame prediction mode is used, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform motion compensation in response to the encoder. Images of P-frames and B-frames that are encoded or decoded using reference frames, or encoded and decoded, can be considered as images used for inter-frame prediction.

[0196] The inter-frame prediction according to the embodiments will be described in detail below.

[0197] Inter-frame prediction or motion compensation can be performed using both reference frames and motion information. Furthermore, inter-frame prediction can utilize the skip mode described above.

[0198] The reference frame can be at least one of the previous and subsequent frames of the current frame. Here, inter-frame prediction can predict blocks of the current frame based on the reference frame. Here, the reference frame can refer to the image used when predicting the blocks. Here, the region within the reference frame can be indicated by using a reference frame index (refIdx) indicating the reference frame, motion vectors, etc.

[0199] Inter-frame prediction can select a reference frame and a reference block within that frame that is related to the current block. The predicted block for the current block can be generated using the selected reference block. The current block can be a block within the current frame that is the current encoding target or the current decoding target.

[0200] Motion information can be derived from inter-frame prediction processing by encoding device 100 and decoding device 200. Furthermore, the derived motion information can be used when performing inter-frame prediction. Here, encoding device 100 and decoding device 200 can improve encoding efficiency or decoding efficiency or both by using motion information of reconstructed neighboring blocks or motion information of col-blocks (col blocks). A col-block can be a block within a previously reconstructed col-frame that relates to the spatial location of the encoded / decoded target block. Reconstructed neighboring blocks can be blocks within the current frame, as well as blocks previously reconstructed through encoding or decoding, or both. Furthermore, a reconstructed block can be a block adjacent to the encoded / decoded target block, or a block located at the outer corner of the encoded / decoded target block, or both. Here, a block located at the outer corner of the encoded / decoded target block can be a block vertically adjacent to a horizontally adjacent neighboring block of the encoded / decoded target block. Alternatively, a block located at the outer corner of the encoded / decoded target block can be a block horizontally adjacent to a vertically adjacent neighboring block of the encoded / decoded target block.

[0201] Encoding device 100 and decoding device 200 can each determine a block existing within the col frame at a location related to the encoding / decoding target block space, and can determine a predefined relative position based on the determined block. The predefined relative position can be an internal or external position of the block existing at a location related to the encoding / decoding target block space, or both. Furthermore, encoding device 100 and decoding device 200 can respectively derive the col block based on the determined predefined relative position. Here, the col frame can be one of at least one reference frame included in a list of reference frames.

[0202] The method for deriving motion information can vary depending on the prediction mode of the encoded / decoded target block. For example, prediction modes applied to inter-frame prediction may include Advanced Motion Vector Prediction (AMVP), merging modes, etc. Here, the merging mode can be referred to as the motion merging mode.

[0203] For example, when AMVP is applied as a prediction mode, encoding device 100 and decoding device 200 can generate motion vector candidate lists by reconstructing motion vectors of neighboring blocks or motion vectors of col blocks, or both. The motion vectors of reconstructing neighboring blocks or motion vectors of col blocks, or both, can be used as motion vector candidates. Here, the motion vectors of col blocks can be referred to as temporal motion vector candidates, and the motion vectors of reconstructing neighboring blocks can be referred to as spatial motion vector candidates.

[0204] Encoding device 100 can generate a bitstream, which may include motion vector candidate indices. That is, encoding device 100 can generate a bitstream by entropy encoding the motion vector candidate indices. The motion vector candidate indices can indicate the optimal motion vector candidate selected from the motion vector candidates included in the motion vector candidate list. The motion vector candidate indices can be transmitted from encoding device 100 to decoding device 200 via the bitstream.

[0205] The decoding device 200 can entropy decode the motion vector candidate index from the bit stream, and can select the motion vector candidate of the target block from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index.

[0206] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the target block and the motion vector candidates, and entropy encode the MVD. The bitstream may include the entropy-encoded MVD. The MVD can be sent from encoding device 100 to decoding device 200 via the bitstream. Here, decoding device 200 can entropy decode the received MVD from the bitstream. Decoding device 200 can deduce the motion vector of the target block from the sum of the decoded MVD and the motion vector candidates.

[0207] The bitstream may include a reference frame index indicating a reference frame, and the reference frame index may be entropy encoded and transmitted from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 may predict the motion vector of the target block to be decoded using motion information from neighboring blocks, and may deduce the motion vector of the target block to be decoded using the predicted motion vector and the motion vector difference. The decoding device 200 may generate a predicted block of the target block to be decoded based on the deduced motion vector and the reference frame index information.

[0208] As another method for deriving motion information, a merging pattern is used. A merging pattern can refer to the merging of motion from multiple blocks. A merging pattern can refer to the application of motion information from one block to another. When a merging pattern is applied, the encoding device 100 and the decoding device 200 can generate a merging candidate list by reconstructing motion information from neighboring blocks or motion information from the col block, or both. Motion information may include at least one of the following: 1) a motion vector, 2) a reference frame index, and 3) an inter-frame prediction indicator. The prediction indicator may indicate unidirectional (L0 prediction, L1 prediction) or bidirectional prediction.

[0209] Here, the merging mode can be applied to each CU or each PU. When the merging mode is executed in each CU or each PU, the encoding device 100 can generate a bitstream by entropy decoding of predefined information and can send the bitstream to the decoding device 200. The bitstream may include the predefined information. The predefined information may include: 1) a merging flag indicating whether the merging mode is executed for each block partition; and 2) a merging index indicating which block among the neighboring blocks adjacent to the encoding target block is merged. For example, the neighboring blocks adjacent to the encoding target block may include the left neighboring block of the encoding target block, the upper neighboring block of the encoding target block, the time neighboring block of the encoding target block, etc.

[0210] The merge candidate list indicates a list storing motion information. Furthermore, the merge candidate list can be generated before executing the merge mode. The motion information stored in the merge candidate list can be at least one of the following: motion information of neighboring blocks adjacent to the encoding / decoding target block, motion information of co-occurring blocks in the reference frame related to the encoding / decoding target block, newly generated motion information through pre-combining of motion information existing in the motion candidate list, and zero merge candidates. Here, the motion information of neighboring blocks adjacent to the encoding / decoding target block can be referred to as spatial merge candidates. The motion information of co-occurring blocks in the reference frame related to the encoding / decoding target block can be referred to as temporal merge candidates.

[0211] A skip mode can be a mode that applies mode information from neighboring blocks to the encoded / decoded target block. A skip mode can be one of the modes used for inter-frame prediction. When a skip mode is used, encoding device 100 can entropy-encode information about which block's motion information is used as the motion information for the encoded target block and transmit this information to decoding device 200 via a bitstream. Encoding device 100 may not transmit other information (e.g., syntax element information) to decoding device 200. Syntax element information may include at least one of motion vector difference information, coded block flags, and transform coefficient levels.

[0212] The residual signal generated after intra-frame or inter-frame prediction can be transformed to the frequency domain through a transform process as part of the quantization process. Here, the initial transform can use DCT Type 2 (DCT-II) and various DCT and DST kernels. These transform kernels can perform separable transforms on the residual signal for 1D transforms along the horizontal and / or vertical directions, or they can perform 2D non-separable transforms on the residual signal.

[0213] For example, in the case of 1D transforms, the DCT and DST types used in the transform can be DCT-II, DCT-V, DCT-VIII, DST-I, and DST-VII as shown in the following tables. For example, as shown in Tables 1 and 2, the DCT or DST types used in the transform by synthesizing transform sets can be derived.

[0214] [Table 1]

[0215]

[0216] [Table 2]

[0217]

[0218] For example, such as Figure 8 As shown, different transform sets are defined for the horizontal and vertical directions according to the intra-prediction mode. Next, the encoder / decoder can perform transforms and / or inverse transforms using the intra-prediction mode of the current encoding / decoding target block and the transforms of the associated transform sets. In this case, entropy encoding / decoding is not performed on the transform sets, and the encoder / decoder can define the transform sets according to the same rules. In this case, information indicating which transform among the transforms of the transform set is used can be entropy encoded / decoded. For example, when the block size is equal to or less than 64×64, three transform sets are synthesized according to the intra-prediction mode, as shown in Table 2, and three transforms are used for each horizontal and vertical transform to combine and perform a total of nine multi-transform methods. Next, the residual signal is encoded / decoded using the optimal transform method, thereby improving coding efficiency. Here, truncated unary binaryization can be used to entropy encode / decode information about which transform method among the three transforms in a transform set is used. Here, to perform at least one of the vertical and horizontal transforms, entropy encoding / decoding can be performed on information indicating which transform among the transforms of the transform set is used.

[0219] After completing the first transformation described above, as follows Figure 9As shown, the encoder can perform a secondary transformation on the transform coefficients to improve energy concentration. The secondary transformation can perform a separable transformation for 1D transformation along the horizontal and / or vertical directions, or a non-separable 2D transformation. The transformation information used can be transmitted or derived by the encoder / decoder based on current and neighboring encoded information. For example, a transform set for the secondary transformation can be defined, such as for 1D transformation. Entropy encoding / decoding is not performed on this transform set, and the encoder / decoder can define the transform set according to the same rules. In this case, information indicating which transform among the transforms in the transform set is used can be transmitted, and this information can be applied to at least one residual signal via intra-frame prediction or inter-frame prediction.

[0220] At least one of the number or type of transform candidates varies for each transform set. At least one of the number or type of transform candidates may be determined differently based on at least one of the following: the location, size, partitioning form of the block (CU, PU, ​​TU, etc.), and the orientation / non-orientation of the prediction mode (intra-frame / inter-frame mode) or intra-frame prediction mode.

[0221] The decoder can perform a second inverse transform based on whether the second inverse transform has been performed, and can perform a first inverse transform based on whether the first inverse transform has been performed from the result of the second inverse transform.

[0222] The aforementioned first and second transformations can be applied to at least one signal component in the luminance / chrominance components, or can be applied according to the size / shape of any coded block. Entropy encoding / decoding can be performed on indices indicating whether the first / second transformation is used and both the first / second transformation used in any coded block. Alternatively, these indices can be derived by default by the encoder / decoder based on at least one current / nearby coded information.

[0223] The residual signal generated after intra-frame prediction or inter-frame prediction is quantized after the first and / or second transforms, and the quantized transform coefficients are entropy-coded. Here, as... Figure 10 As shown, the quantized transform coefficients can be scanned in the diagonal, vertical and horizontal directions based on at least one of the intra-frame prediction mode or the size / shape of the minimum block.

[0224] Furthermore, the quantized transform coefficients that have undergone entropy decoding can be arranged in blocks by inverse scanning, and at least one of dequantization or inverse transform can be performed on the relevant blocks. Here, as a method of inverse scanning, at least one of diagonal scanning, horizontal scanning, and vertical scanning can be performed.

[0225] For example, when the current coded block size is 8×8, the residual signal for the 8×8 block can be subjected to a first transform, a second transform, and quantization. Then, according to... Figure 10 At least one of the three scanning order methods shown performs scanning and entropy encoding on the quantized transform coefficients for each of the four 4×4 sub-blocks. Furthermore, an inverse scan can be performed on the quantized transform coefficients by performing entropy decoding. The quantized transform coefficients that have undergone inverse scanning become transform coefficients after dequantization, and at least one of a second inverse transform or a first inverse transform is performed, thereby generating the reconstructed residual signal.

[0226] In video encoding processing, a block can be like... Figure 11 The blocks are partitioned as shown, and indicators corresponding to the partition information can be transmitted using signals. Here, the partition information can be at least one of the following: a partition flag (split_flag), a quadtree / binary tree flag (QB_flag), a quadtree partition flag (quadtree_flag), a binary tree partition flag (binarytree_flag), and a binary tree partition type flag (Btype_flag). Here, split_flag indicates whether the block is partitioned, QB_flag indicates whether the block is partitioned in quadtree or binary tree form, quadtree_flag indicates whether the block is partitioned in quadtree form, binarytree_flag indicates whether the block is partitioned in binary tree form, and Btype_flag indicates whether the block is vertically or horizontally partitioned in the case of binary tree partitioning.

[0227] When the partition flag is 1, it indicates that the partition is executed; when the partition flag is 0, it indicates that the partition is not executed. In the case of the quadtree / binary tree flag, 0 indicates a quadtree partition, and 1 indicates a binary tree partition. Optionally, 0 can indicate a binary tree partition, and 1 can indicate a quadtree partition. In the case of the binary tree partition type flag, 0 can indicate a horizontal partition, and 1 can indicate a vertical partition. Optionally, 0 can indicate a vertical partition, and 1 can indicate a horizontal partition.

[0228] For example, it can be derived by sending at least one of the quadtree_flag, binarytree_flag, and Btype_flag as shown in Table 3 using a signal. Figure 11 Partition information.

[0229] [Table 3]

[0230]

[0231] For example, it can be derived by transmitting at least one of the split_flag, QB_flag, and Btype_flag as shown in Table 4 using a signal. Figure 11 Partition information.

[0232] [Table 4]

[0233]

[0234] The partitioning method can be performed either in quadtree or binary tree form only, depending on the size / shape of the block. In this case, `split_flag` can be interpreted as a flag indicating whether partitioning is performed in quadtree or binary tree form. The size / shape of the block can be derived from the block's depth information, which can be transmitted using signals.

[0235] When the block size is within a predetermined range, partitioning can be performed only in quadtree form. Here, the predetermined range can be defined as at least one of the largest block size or the smallest block size that can be partitioned only in quadtree form. Information indicating the largest / minimum block size that allows quadtree-form partitioning can be transmitted via a bitstream signal, and this information can be transmitted via signal in units of at least one of sequence, frame parameters, or stripes (segments). Alternatively, the largest / minimum block size can be a fixed size preset in the encoder / decoder. For example, when the block size ranges from 256×256 to 64×64, partitioning can be performed only in quadtree form. In this case, split_flag can be a flag indicating whether partitioning is performed in quadtree form.

[0236] When the block size is within a predetermined range, partitioning can be performed only in a binary tree format. Here, the predetermined range can be defined as at least one of the largest block size or the smallest block size that can be partitioned only in a binary tree format. Information indicating the largest / minimum block size that allows binary tree partitioning can be transmitted via a bitstream signal, and this information can be transmitted via a signal in units of at least one of sequence, frame parameters, or stripes (segments). Alternatively, the largest / minimum block size can be a fixed size preset in the encoder / decoder. For example, when the block size ranges from 16×16 to 8×8, partitioning can be performed only in a binary tree format. In this case, split_flag can be a flag indicating whether partitioning is performed in a binary tree format.

[0237] After partitioning a block according to a binary tree, when the partitioned block is further partitioned, partitioning can be performed only according to the binary tree.

[0238] When the width or length of a partitioned block cannot be further partitioned, at least one indicator may not be transmitted by signal.

[0239] In addition to binary tree partitioning based on quadtrees, quadtree-based partitioning can be performed after binary tree partitioning.

[0240] Based on the above description, the method for encoding / decoding images according to the present invention will be described in detail.

[0241] Figure 12 This is a flowchart illustrating a method for encoding video using a merging pattern according to the present invention. Figure 13 This is a flowchart illustrating a method for decoding video using a merging mode according to the present invention.

[0242] Reference Figure 12 In step S1201, the encoding device can deduce merging candidates and generate a merging candidate list based on the deduce merging candidates. When the merging candidate list is generated, in step S1202, motion information is determined by using the generated merging candidate list, and in step S1203, motion compensation for the current block can be performed using the determined motion information. Next, in step S1204, the encoding device can perform entropy encoding on the information regarding motion compensation.

[0243] Reference Figure 13 In step S1301, the decoding device performs entropy decoding on the motion compensation information received from the encoding device, and in step S1302, it derives merging candidates and generates a merging candidate list based on the derived merging candidates. When the merging candidate list is generated, in step S1303, the motion information of the current block can be determined by using the generated merging candidate list. Next, in step S1304, the decoding device performs motion compensation using the motion information.

[0244] The following text will describe it in detail. Figure 12 and Figure 13 The steps are shown in the diagram.

[0245] First, the derivation of the merging candidates in steps S1201 and S1302 will be described in detail.

[0246] The current block's merge candidates may include at least one of spatial merge candidates, temporal merge candidates, and additional merge candidates.

[0247] Spatial merge candidates for the current block can be derived from the reconstructed blocks adjacent to it. For example, motion information from the reconstructed blocks adjacent to the current block can be used to determine spatial merge candidates for the current block. Here, motion information may include at least one of motion vectors, reference frame indices, and prediction list usage flags.

[0248] In this case, the motion information for the spatial merging candidates may include motion information corresponding to L0 and L1, as well as motion information corresponding to L0, L1, ..., LX. Here, X can be a positive integer including zero. Therefore, the list of reference frames may include at least one of L0, L1, ..., LX.

[0249] Figure 14 This is a diagram illustrating an example of deriving spatial merge candidates for the current block. Here, deriving spatial merge candidates can refer to deriving spatial merge candidates and adding them to the merge candidate list.

[0250] Reference Figure 14 The spatial merging candidates for the current block can be derived from the neighboring blocks adjacent to the current block X. The neighboring blocks adjacent to the current block can include at least one of the following: the block above the current block (B1), the block to the left of the current block (A1), the block corresponding to the upper right corner of the current block (B0), the block to the upper left corner of the current block (B2), and the block to the lower left corner of the current block (A0). Furthermore, the neighboring blocks adjacent to the current block can have square or non-square shapes.

[0251] To derive spatial merge candidates for the current block, it can be determined whether neighboring blocks adjacent to the current block can be used to derive spatial merge candidates for the current block. Here, the determination of whether neighboring blocks adjacent to the current block can be used to derive spatial merge candidates for the current block can be based on a predetermined priority. For example, in Figure 14 In the example shown, the availability of spatial merge candidates can be determined based on the order of the block at position A1, the block at position B1, the block at position B0, the block at position A0, and the block at position B2. Spatial merge candidates determined based on the order used to determine availability can be added sequentially to the merge candidate list of the current block. Below are examples of neighboring blocks that cannot be used to derive spatial merge candidates for the current block.

[0252] 1) When the neighboring block is at position B2, derive the space merge candidate from the blocks at positions A0, A1, B0, and B1.

[0253] 2) Cases where neighboring blocks do not exist (cases where the current block exists at the screen boundary, stripe boundary, or parallel block boundary, etc.)

[0254] 3) Case where neighboring blocks are intra-coded

[0255] 4) The case where at least one of the motion vector, reference frame index, and reference frame of a neighboring block is the same as at least one of the motion vector, reference frame index, and reference frame of a previously derived spatial merging candidate.

[0256] 5) The motion vector of a neighboring block points to the outer region of the boundary of at least one of the current block's frame, strip, and parallel block.

[0257] Figure 15 This is a diagram illustrating an example of adding spatial merge candidates to the merge candidate list.

[0258] Reference Figure 15 When four spatial merge candidates are derived from the neighboring block at position A1, the neighboring block at position B0, the neighboring block at position A0, and the neighboring block at position B2, the derived spatial merge candidates can be added to the merge candidate list in sequence.

[0259] `maxNumSpatialMergeCand` can represent the maximum number of spatial merge candidates that can be included in the merge candidate list, and `numMergeCand` can represent the number of merge candidates included in the merge candidate list. `maxNumSpatialMergeCand` can be a positive integer including zero. `maxNumSpatialMergeCand` can be preset so that the encoding and decoding devices use the same value. Optionally, the encoding device can encode the maximum number of merge candidates that can be included in the merge candidate list of the current block, and this maximum number can be signaled to the decoding device via a bitstream.

[0260] As described above, when at least one spatial merge candidate is derived from neighboring blocks A1, B1, B0, A0, and B2, a spatial merge candidate flag (spatialCand) indicating whether it is a spatial merge candidate can be set for each derived merge candidate. For example, when a spatial merge candidate is derived, spatialCand can be set to a predetermined value of 1. Otherwise, spatialCand can be set to a predetermined value of 0. Furthermore, the spatial merge candidate count (spatialCandCnt) is incremented by 1 each time a spatial merge candidate is derived.

[0261] Spatial merging candidates can be derived based on at least one of the encoding parameters of the current block or neighboring blocks.

[0262] Space merging candidates can be shared among these blocks: these blocks are smaller in size than the blocks whose motion compensation information is entropy-encoded / decoded, or these blocks are deeper than the blocks whose motion compensation information is entropy-encoded / decoded. Here, the motion compensation information can be at least one of the following: information about whether to use a skip mode, information about whether to use a merge mode, or merge index information.

[0263] The block containing information about motion compensation that is entropy encoded / decoded can be a CTU or a sub-unit of a CTU, a CU, or a PU.

[0264] In the following text, for ease of explanation, the size of the block in which motion compensation information is entropy-encoded / decoded is referred to as the first block size. The depth of the block in which motion compensation information is entropy-encoded / decoded is referred to as the first block depth.

[0265] Specifically, when the size of the current block is smaller than the size of the first block, a spatial merge candidate for the current block can be derived from at least one of the reconstructed blocks adjacent to the higher-level block with the size of the first block. Furthermore, blocks included in the higher-level block can share the derived spatial merge candidates. Here, the block with the size of the first block can be referred to as the higher-level block of the current block.

[0266] Figure 16 This is a diagram illustrating an embodiment of deriving and sharing space merging candidates in a CTU. (Refer to...) Figure 16 When the first block size is 32×32, blocks 1601, 1602, 1603 and 1604 smaller than 32×32 can be derived from at least one of the neighboring blocks adjacent to the high-level block 1600 with the first block size, and can share the derived space merging candidates.

[0267] For example, when the first block size is 32×32 and the coded block size is 32×32, prediction blocks smaller than 32×32 can derive spatial merging candidates from at least one motion information from neighboring blocks of the current block. Prediction blocks within a coded block can share the derived spatial merging candidates. Here, coded block and prediction block can refer to blocks as a more generalized representation.

[0268] When the current block's depth is greater than the first block's depth, at least one deduced space merge candidate can be selected from the reconstructed blocks adjacent to the higher block with the first block's depth. Furthermore, blocks included in higher blocks can share deduced space merge candidates. Here, the block with the first block's depth can be referred to as the higher-level block of the current block.

[0269] For example, when the first block depth is 2 and the coded block depth is 2, prediction blocks deeper than block depth 2 can derive spatial merging candidates based on at least one motion information from neighboring blocks of the coded block. Prediction blocks within the coded block can share the derived spatial merging candidates.

[0270] Here, shared space merge candidates can represent separate lists of merge candidates that can generate shared blocks based on the same space merge candidates.

[0271] Furthermore, shared spatial merge candidates can indicate that shared blocks can perform motion compensation using a list of merge candidates. Here, the shared merge candidate list may include at least one of the spatial merge candidates derived from higher-level blocks that are entropy-encoded / decoded based on information about motion compensation.

[0272] The neighboring block adjacent to the current block, or the current block itself, can have a square shape or a non-square shape.

[0273] Furthermore, neighboring blocks adjacent to the current block can be partitioned into sub-blocks. In this case, the motion information of one sub-block among the neighboring blocks adjacent to the current block can be determined as a spatial merge candidate for the current block. Additionally, a spatial merge candidate for the current block can be determined based on at least one piece of motion information from the sub-blocks of the neighboring blocks adjacent to the current block. Here, it can be determined whether the sub-blocks of the neighboring blocks can be used to derive spatial merge candidates to determine the spatial merge candidate for the current block. The availability for deriving spatial merge candidates may include at least one of the following: the existence of motion information for the sub-blocks of the neighboring blocks, and whether the motion information of the sub-blocks of the neighboring blocks can be used as a spatial merge candidate for the current block.

[0274] In addition, the median, average, minimum, maximum, weighted average, or pattern of at least one motion information (i.e., motion vector) of the neighboring block's sub-blocks can be identified as a spatial merging candidate for the current block.

[0275] Next, we will describe the method for deriving the time merge candidates for the current block.

[0276] The temporal merging candidate for the current block can be derived from the reconstructed blocks included in the co-frames of the current frame. Here, a co-frame is a frame that has been encoded / decoded before the current frame. A co-frame can be a frame with a different temporal order than the current frame.

[0277] Figure 17 This is a diagram illustrating an example of deriving time merge candidates for the current block. Here, deriving time merge candidates can mean deriving time merge candidates and adding them to the merge candidate list.

[0278] Reference Figure 17 In the same frame as the current block, the time merging candidate can be derived from blocks located outside the block corresponding to the same spatial position as the current block X, or from blocks located inside the block corresponding to the same spatial position as the current block X. Here, the time merging candidate can refer to the motion information of the same frame block. For example, the time merging candidate of the current block X can be derived from block H, which is adjacent to the lower right corner of block C corresponding to the same spatial position as the current block, or from block C3, which includes the center point of block C. Block H or block C3 used to derive the time merging candidate of the current block can be called a "same frame block".

[0279] Additionally, the sibling block of the current block or the current block itself can have a square shape or a non-square shape.

[0280] When the time-merging candidate for the current block can be derived from block H, which includes a position outside block C, block H can be set as the co-position block of the current block. In this case, the time-merging candidate for the current block can be derived based on the motion information of block H. Conversely, when the time-merging candidate for the current block cannot be derived from block H, block C3, which includes a position inside block C, can be set as the co-position block of the current block. In this case, the time-merging candidate for the current block can be derived based on the motion information of block C3. When the time-merging candidate for the current block cannot be derived from both block H and block C (e.g., when both block H and block C3 are intra-coded), the time-merging candidate for the current block may not be derived, or it may be derived from a block located at a different position than block H and block C3.

[0281] As another example, time merge candidates for the current block can be derived from multiple blocks in the same frame. For instance, multiple time merge candidates for the current block can be derived from blocks H and C3.

[0282] Figure 18 This is a diagram illustrating an example of adding time merging candidates to the merge candidate list.

[0283] Reference Figure 18 When a time merge candidate is derived from the co-occurrence block at position H1, the derived time merge candidate can be added to the merge candidate list.

[0284] The co-located blocks of the current block can be partitioned into sub-blocks. In this case, the motion information of one sub-block within the co-located blocks of the current block can be determined as a time-merging candidate for the current block. Furthermore, a time-merging candidate for the current block can be determined based on at least one piece of motion information from the sub-blocks of the co-located blocks of the current block.

[0285] Here, it can be determined whether the motion information of the child blocks of the same block exists or whether the motion information of the child blocks of the same block can be used as a time merge candidate for the current block to determine the time merge candidate for the current block.

[0286] In addition, the median, average, minimum, maximum, weighted average, or pattern of at least one motion information (i.e., motion vector) of the co-located block can be determined as a time merging candidate for the current block.

[0287] exist Figure 17 In this context, the time merge candidate for the current block can be derived from the block adjacent to the lower right corner of the co-located block or from the block that includes the center point of the co-located block. However, the location of the block used to derive the time merge candidate for the current block is not limited to... Figure 17Examples are shown in the figure. For example, the time merge candidate of the current block can be derived from the block adjacent to the top / bottom boundary, left / right boundary or corner of the co-block, and the time merge candidate of the current block can be derived from the block at a specific position in the co-block (i.e., the block adjacent to the corner boundary of the co-block).

[0288] The temporal merging candidates for the current block can be determined by considering the reference frame list (or prediction direction) of the current block and its sibling blocks. Simultaneously, the motion information for the temporal merging candidates can include motion information corresponding to L0 and L1, as well as motion information corresponding to L0, L1, ..., LX. Here, X can be a positive integer including zero.

[0289] For example, when the available list of reference frames for the current block is L0 (i.e., when the inter-frame prediction indicator indicates PRED_L0), the motion information corresponding to L0 of the co-located block can be derived as a time-merging candidate for the current block. In other words, when the available list of reference frames for the current block is LX (where X is an integer such as 0, 1, 2, or 3, indicating the index of the reference frame list), the motion information corresponding to LX of the co-located block (hereinafter referred to as "LX motion information") can be derived as a time-merging candidate for the current block.

[0290] When the current block uses multiple reference screen lists, the time merge candidates for the current block can be determined by considering the reference screen lists of the current block and the co-located blocks.

[0291] For example, when performing bidirectional prediction on the current block (i.e., with the inter-frame prediction indicator PRED_BI), at least two pieces of information selected from the group including L0 motion information, L1 motion information, L2 motion information, ..., and LX motion information of the co-located block can be derived as time-merging candidates. When performing tridirectional prediction on the current block (i.e., with the inter-frame prediction indicator PRED_TRI), at least three pieces of information selected from the group including L0 motion information, L1 motion information, L2 motion information, ..., and LX motion information of the co-located block can be derived as time-merging candidates. When performing quadridirectional prediction on the current block (i.e., with the inter-frame prediction indicator PRED_QUAD), at least four pieces of information selected from the group including L0 motion information, L1 motion information, L2 motion information, ..., and LX motion information of the co-located block can be derived as time-merging candidates.

[0292] In addition, at least one of the following can be derived based on the encoding parameters of the current block, neighboring blocks, and co-occurring blocks: time merging candidate, co-occurring picture, co-occurring block, prediction list using flags, and reference picture index.

[0293] When the number of derived spatial merge candidates is less than the maximum number of merge candidates, temporal merge candidates can be prepared. Therefore, when the number of derived spatial merge candidates reaches the maximum number of merge candidates, the processing of temporal merge candidates can be omitted.

[0294] For example, when the maximum number of merge candidates is two, and the two derived spatial merge candidates have different values, the processing of the derived temporal merge candidates can be omitted.

[0295] As another example, the time-merging candidates for the current block can be derived based on the maximum number of time-merging candidates. Here, the maximum number of time-merging candidates can be preset so that the encoding and decoding devices use the same value. Alternatively, information indicating the maximum number of time-merging candidates for the current block can be encoded via a bitstream and transmitted to the decoding device by signaling. For example, the encoding device can encode maxNumTemporalMergeCand indicating the maximum number of time-merging candidates for the current block, and maxNumTemporalMergeCand can be transmitted to the decoding device by signaling via a bitstream. Here, maxNumTemporalMergeCand can be set to a positive integer including 0. For example, maxNumTemporalMergeCand can be set to 1. The value of maxNumTemporalMergeCand can be derived differently based on information about the number of time-merging candidates transmitted by signaling, and maxNumTemporalMergeCand can be a fixed value preset in the encoder / decoder.

[0296] When the distance between the current frame (including the current block) and its reference frame differs from the distance between the co-frame (including the co-block) and its reference frame, the motion vector of the current block's time-merging candidate can be obtained by scaling the motion vector of the co-frame. Here, scaling can be performed based on at least one of the distance between the reference frame (referenced by the current frame) and the current block, and the distance between the reference frame (referenced by the co-frame) and the co-block. For example, the motion vector of the co-frame is scaled according to the ratio of the distance between the reference frame (referenced by the current frame) and the current block to the distance between the reference frame (referenced by the co-frame) and the co-block, thereby deriving the motion vector of the current block's time-merging candidate.

[0297] Based on the size (first block size) or depth (first block depth) of the blocks whose motion compensation information is entropy encoded / decoded, temporal merging candidates can be shared among these blocks: these blocks are smaller in size than the blocks whose motion compensation information is entropy encoded / decoded, or these blocks are deeper than the blocks whose motion compensation information is entropy encoded / decoded. Here, the motion compensation information can be at least one of information about whether a skip mode is used, information about whether a merge mode is used, and merge index information.

[0298] The block containing information about motion compensation that is entropy encoded / decoded can be a CTU or a sub-unit of a CTU, a CU, or a PU.

[0299] Specifically, when the size of the current block is smaller than the size of the first block, the time merge candidate for the current block can be derived from the co-located blocks of the higher-level block with the size of the first block. Furthermore, blocks included in higher-level blocks can share the derived time merge candidate.

[0300] Furthermore, when the current block's depth is greater than the first block's depth, time merge candidates can be derived from the co-located blocks of the higher-level block with the first block's depth. Additionally, blocks included in higher-level blocks can share the derived time merge candidates.

[0301] Here, shared time merge candidates can refer to separate lists of merge candidates for shared blocks that can be generated based on the same time merge candidates.

[0302] Furthermore, shared time merge candidates can indicate that shared blocks can perform motion compensation using a list of merge candidates. Here, the shared list of merge candidates may include time merge candidates derived from higher-level blocks that are entropy-encoded / decoded based on information about motion compensation.

[0303] Figure 19 This is a diagram illustrating an example of scaling motion vectors of motion information from co-located blocks to derive temporal merging candidates for the current block.

[0304] The motion vector of a co-position block can be scaled based on at least one of the difference (td) between the POC (Plot Order Count) indicating the display order of the co-position block and the POC of the reference picture of the co-position block, and the difference (tb) between the POC of the current picture and the POC of the reference picture of the current block.

[0305] Before scaling, td or tb can be adjusted so that td or tb exists within a predetermined range. For example, when the predetermined range indicates -128 to 127 and td or tb is less than -128, td or tb can be adjusted to -128. When td or tb is greater than 127, td or tb can be adjusted to 127. When td or tb is within the range of -128 to 127, td or tb is not adjusted.

[0306] The scaling factor DistScaleFactor can be calculated based on td or tb. Here, the scaling factor can be calculated based on Formula 1 below.

[0307] [Formula 1]

[0308]

[0309]

[0310] In Formula 1, the absolute value function is specified as Abs(), which outputs the absolute value of the input value.

[0311] The value of the scaling factor DistScaleFactor calculated based on Formula 1 can be adjusted to a predetermined range. For example, DistScaleFactor can be adjusted to exist in the range of -1024 to 1023.

[0312] By scaling the motion vector of the co-location block using a scaling factor, the motion vector of the time-merging candidate for the current block can be determined. For example, the motion vector of the time-merging candidate for the current block can be determined using Equation 2 below.

[0313] [Formula 2]

[0314]

[0315] In Equation 2, Sign() is a function that outputs the sign information of the values ​​in (). For example, Sign(-1) outputs -. In Equation 2, the motion vector of the co-position block can be specified as mvCol.

[0316] Next, we will describe the method for deriving additional merge candidates for the current block.

[0317] Additional merge candidates can refer to at least one of modified spatial merge candidates, modified temporal merge candidates, combined merge candidates, and merge candidates with predetermined motion information values. Here, deriving additional merge candidates can mean deriving additional merge candidates and adding them to the merge candidate list.

[0318] A modified spatial merging candidate can refer to a merging candidate in which at least one motion information of the derived spatial merging candidate has been modified.

[0319] A modified time merging candidate can refer to a merging candidate in which at least one motion information of the derived time merging candidate has been modified.

[0320] A combined merging candidate can refer to a merging candidate derived by combining at least one piece of motion information from the spatial merging candidate, temporal merging candidate, modified spatial merging candidate, modified temporal merging candidate, combined merging candidate, and merging candidate with a predetermined motion information value that exist in the merging candidate list.

[0321] Optionally, the combined merging candidate can refer to a merging candidate derived by combining at least one piece of motion information from a spatial merging candidate, a temporal merging candidate, a modified spatial merging candidate, a modified temporal merging candidate, a combined merging candidate, and a merging candidate with a predetermined motion information value. The spatial merging candidate and the temporal merging candidate are derived from a block that does not exist in the merging candidate list but can be used to derive at least one of the spatial and temporal merging candidates. The modified spatial merging candidate and the modified temporal merging candidate are generated based on the spatial merging candidate and the temporal merging candidate.

[0322] Alternatively, the merging candidates for the combination can be derived using motion information decoded from the bitstream entropy by the decoder. Here, the motion information used in deriving the merging candidates for the combination can be encoded into the bitstream by the encoder entropy.

[0323] The combined merging candidate can refer to the combined bidirectional prediction merging candidate. The combined bidirectional prediction merging candidate is a merging candidate that uses bidirectional prediction and can refer to a merging candidate with both L0 motion information and L1 motion information.

[0324] Furthermore, a merging candidate for a combination can refer to a merging candidate having at least N motion information selected from a group including L0 motion information, L1 motion information, L2 motion information, and L3 motion information. Here, N can refer to a positive integer equal to or greater than 2.

[0325] Merging candidates with predetermined motion information values ​​can refer to zero merging candidates with motion vectors of (0,0). Furthermore, merging candidates with predetermined motion information values ​​can be preset so that the encoding and decoding devices use the same values.

[0326] At least one of the following can be derived or generated based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks: a modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate with a predetermined motion information value. Furthermore, at least one of the following can be added to the merge candidate list based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks: a modified spatial merge candidate, a modified temporal merge candidate, a combined merge candidate, and a merge candidate with a predetermined motion information value.

[0327] Additional merge candidates can be derived for each sub-block of the current block, neighboring blocks, or co-located blocks. The merge candidates derived for each sub-block can be added to the merge candidate list of the current block.

[0328] Additional merge candidates can be derived only in the case of B strips / B frames or only in the case of strips / frames using a list of M reference frames. Here, M can be 3 or 4, and can refer to a positive integer equal to or greater than 3.

[0329] Up to N additional merge candidates can be derived. Here, N is a positive integer including 0. N can be a variable value derived based on information about the maximum number of merge candidates included in the merge candidate list. Alternatively, N can be a fixed value preset in the encoder / decoder. Here, N can vary depending on the size, shape, depth, or position of the block being encoded / decoded in merge mode.

[0330] The merge candidate list has a preset size and can be increased by the number of additional merge candidates generated after adding spatial or temporal merge candidates. In this case, all additional merge candidates generated can be included in the merge candidate list. Conversely, the size of the merge candidate list can be increased by a smaller amount than the number of additional merge candidates (e.g., the number of additional merge candidates). N (where N is a positive integer). In this case, only a portion of the resulting additional merge candidates can be included in the merge candidate list.

[0331] Furthermore, the size of the merge candidate list can be determined based on the encoding parameters of the current block, neighboring blocks, or co-occurring blocks, and the size of the merge candidate list can be changed based on the encoding parameters.

[0332] To increase the throughput of merging modes in the encoder and decoder, motion compensation using merging modes can be performed solely through spatial merging candidate derivation, temporal merging candidate derivation, and zero merging candidate derivation without deriving combined merging candidates. In cases where combined merging candidate derivation is performed after temporal merging candidate derivation, which requires relatively long cycle times, the worst-case hardware complexity of the merging mode is that of temporal merging candidate derivation, rather than combined merging candidate derivation following temporal merging candidate derivation, when combined merging candidate derivation is not performed. Therefore, the cycle time required to derive each merging candidate in the merging mode can be reduced. Furthermore, in merging modes without deriving combined merging candidates, there is no dependency between merging candidate derivation processes. Therefore, it is advantageous that spatial merging candidate derivation, temporal merging candidate derivation, and zero merging mode candidate derivation can be performed in parallel.

[0333] Figure 21a and Figure 21bThis is a diagram illustrating an embodiment of a method for deriving merge candidates for combinations. Before deriving merge candidates for combinations, the following steps can be performed: when at least one merge candidate exists in the merge candidate list, or when the number of merge candidates (numOrigMergeCand) in the merge candidate list is less than the maximum number of merge candidates (MaxNumMergeCand). Figure 21a and Figure 21b The derivation of the method for merging candidate combinations in the text.

[0334] Reference Figure 21a and Figure 21b The encoder / decoder can set the number of merge candidates in the input (numInputMergeCand) to the number of merge candidates in the current merge candidate list (numMergeCand), and can set the index of the combination (combIdx) to 0. The k-th (numMergeCand) can be derived. Merge candidates for the combination of numInputMergeCand.

[0335] In step S2101, the encoder / decoder can be configured using, for example... Figure 20 The combined indexes shown are used to derive at least one of the L0 candidate index (l0CandIdx), L1 candidate index (l1CandIdx), L2 candidate index (l2CandIdx), and L3 candidate index (l3CandIdx).

[0336] Each candidate index can indicate a merge candidate in the merge candidate list. Motion information based on candidate indices L0, L1, L2, and L3 can be motion information about L0, L1, L2, and L3 of the combined merge candidates.

[0337] In step S2102, the encoder / decoder can deduce the L0 candidate (l0Cand) as the merged candidate corresponding to the L0 candidate index in the merged candidate list (mergeCandList[l0CandIdx]), the L1 candidate (l1Cand) as the merged candidate corresponding to the L1 candidate index in the merged candidate list (mergeCandList[l1CandIdx]), the L2 candidate (l2Cand) as the merged candidate corresponding to the L2 candidate index in the merged candidate list (mergeCandList[l2CandIdx]), and the L3 candidate (l3Cand) as the merged candidate corresponding to the L3 candidate index in the merged candidate list (mergeCandList[l3CandIdx]).

[0338] The encoder / decoder may execute step S2104 if at least one of the following conditions is met in step S2103; otherwise, step S2105 may be executed.

[0339] 1) The condition for using L0 candidates for unidirectional prediction (predFlagL0l0Cand == 1)

[0340] 2) Conditions for using L1 candidates for one-way L1 prediction (predFlagL1l1Cand == 1)

[0341] 3) Conditions for using L2 unidirectional prediction for L2 candidates (predFlagL2l2Cand == 1)

[0342] 4) Conditions for using L3 unidirectional prediction for L3 candidates (predFlagL3l3Cand == 1)

[0343] 5) The condition that at least one reference frame of candidate L0, candidate L1, candidate L2, and candidate L3 is different from the reference frame of another candidate, and at least one motion vector of candidate L0, candidate L1, candidate L2, and candidate L3 is different from the motion vector of another candidate.

[0344] When at least one of the five conditions mentioned above is satisfied in step S2103-Yes, in step S2104, the encoder / decoder can determine the L0 motion information of the L0 candidate as the combined candidate L0 motion information, can determine the L1 motion information of the L1 candidate as the combined candidate L1 motion information, can determine the L2 motion information of the L2 candidate as the combined candidate L2 motion information, can determine the L3 motion information of the L3 candidate as the combined candidate L3 motion information, and can add the combined merged candidate (combCandk) to the merged candidate list.

[0345] For example, the information regarding merger candidates for a combination can be the same as the following.

[0346] The L0 reference screen index of the Kth combination candidate (refIdxL0combCandk) = the L0 reference screen index of the L0 candidate (refIdxL0l0Cand).

[0347] The L1 reference screen index of the Kth combination candidate (refIdxL1combCandk) = the L1 reference screen index of the L1 candidate (refIdxL1l1Cand).

[0348] The L2 reference screen index of the Kth combination candidate (refIdxL2combCandk) = the L2 reference screen index of the L2 candidate (refIdxL2l2Cand).

[0349] The L3 reference screen index of the Kth combination candidate (refIdxL3combCandk) = the L3 reference screen index of the L3 candidate (refIdxL3l3Cand).

[0350] The L0 prediction list of the merge candidates for the Kth combination uses the flag (predFlagL0combCandk = 1)

[0351] The L1 prediction list of the merge candidates for the Kth combination uses the flag (predFlagL1combCandk=1)

[0352] The L2 prediction list of the merge candidates for the Kth combination uses the flag (predFlagL2combCandk=1)

[0353] The L3 prediction list of the merge candidates for the Kth combination uses the flag (predFlagL3combCandk=1)

[0354] The x-component of the merged candidate L0 motion vector of the Kth combination (mvL0combCandk[0]) = the x-component of the candidate L0 motion vector (mvL0l0Cand[0])

[0355] The y-component of the merged candidate L0 motion vector of the Kth combination (mvL0combCandk[1]) = the y-component of the candidate L0 motion vector (mvL0l0Cand[1])

[0356] The x-component of the merged candidate L1 motion vector of the Kth combination (mvL1combCandk[0]) = the x-component of the candidate L1 motion vector (mvL1l1Cand[0])

[0357] The y-component of the merged candidate L1 motion vector of the Kth combination (mvL1combCandk[1]) = the y-component of the candidate L1 motion vector (mvL1l1Cand[1])

[0358] The x-component of the merged candidate L2 motion vector of the Kth combination (mvL2combCandk[0]) = the x-component of the candidate L2 motion vector (mvL2l2Cand[0])

[0359] The y-component of the merged candidate L2 motion vector of the Kth combination (mvL2combCandk[1]) = the y-component of the candidate L2 motion vector (mvL2l2Cand[1])

[0360] The x-component of the merged candidate L3 motion vector of the Kth combination (mvL3combCandk[0]) = the x-component of the candidate L3 motion vector (mvL3l3Cand[0])

[0361] The y-component of the merged candidate L3 motion vector of the Kth combination (mvL3combCandk[1]) = the y-component of the L3 motion vector of the L3 candidate (mvL3l3Cand[1])

[0362] numMergeCand = numMergeCand + 1

[0363] In addition, in step S2105, the encoder / decoder can increment the combined index by 1.

[0364] Furthermore, in step S2106, when the index of the combination is equal to (numOrigMergeCand×(numOrigMergeCand-1)) or when the number of merge candidates in the current merge candidate list (numMergeCand) is equal to the maximum number of merge candidates (MaxNumMergeCand), the encoder / decoder may terminate the merge candidate derivation step of the combination; otherwise, step S2101 may be executed.

[0365] When execution Figure 21a and Figure 21b When deriving the merge candidate of a combination in the method, the derived merge candidate of the combination can be added to, for example... Figure 22 The list of candidate mergers is shown below.

[0366] Simultaneously, before deriving the merge candidates for the combination, if there are at least two spatial merge candidates in the merge candidate list, or if the number of merge candidates in the merge candidate list (numOrigMergeCand) is less than the maximum number of merge candidates (MaxNumMergeCand), a method can be executed to derive the merge candidates for the combination using only spatial merge candidates. In this case, the following can be used: Figure 21a and Figure 21b The derivation of the method for merging candidate combinations in the text.

[0367] However, in Figure 21a The L0, L1, L2, and L3 candidate indices derived in step S2101 can indicate only the merge candidates whose spatial merge candidate flag (spatialCand) is 1. Therefore, the L0, L1, L2, and L3 candidates derived in step S2102 can be derived by using only the merge candidates whose spatial merge candidate flag (spatialCand) is 1 from the merge candidate list.

[0368] In addition, Figure 21b In step S2106, the value of (spatialCandCnt × (spatialCandCnt-1)) is compared with the index of the combination, but the value of (numOrigMergeCand × (numOrigMergeCand-1)) is not compared with the index of the combination. The combination candidate derivation step can be terminated when the index of the combination is equal to (spatialCandCnt × (spatialCandCnt-1)), or when the number of merge candidates (numMergeCand) in the current merge candidate list is equal to 'MaxNumMergeCand'; otherwise, step S2101 can be executed.

[0369] When executing a method that derives merge candidates by using only spatial merge candidates, merge candidates that combine with spatial merge candidates can be added to, for example... Figure 23 The list of candidate mergers is shown below.

[0370] Figure 22 and Figure 23 This is a diagram illustrating an example of deriving combined merge candidates by using at least one of spatial merge candidates, temporal merge candidates, and zero merge candidates, and adding the combined merge candidates to the merge candidate list.

[0371] Here, merging candidates having at least one of the motion information types L0, L1, L2, and L3 can be included in the merging candidate list. Meanwhile, the L0, L1, L2, and L3 reference frame lists have been described as examples, but are not limited thereto. Merging candidates having motion information related to L0~LX reference frame lists (where X is a positive integer) can be included in the merging candidate list.

[0372] Each piece of motion information may include at least one of the following: motion vector, reference frame index, and prediction list using flags.

[0373] like Figure 22 and Figure 23As shown, at least one of the merging candidates can be determined as the final merging candidate. The determined final merging candidate can be used as motion information for the current block. This motion information can be used in inter-frame prediction or motion compensation for the current block. Furthermore, by changing at least one value of the information corresponding to the motion information of the current block, the motion information can be used in inter-frame prediction or motion compensation for the current block. Here, the value to be changed in the information corresponding to the motion information can be at least one of the x-component of the motion vector, the y-component of the motion vector, and the reference frame index. Furthermore, when changing at least one value of the information corresponding to the motion information, the at least one value of the information corresponding to the motion information can be changed such that minimum distortion is indicated by using a distortion calculation method (SAD, SSE, MSE, etc.).

[0374] A predicted block for the current block can be generated by using at least one of the following motion information: L0 motion information, L1 motion information, L2 motion information, and L3 motion information, based on the motion information of the merged candidates. The generated predicted block can be used in inter-frame prediction or motion compensation of the current block.

[0375] When at least one of the motion information from L0, L1, L2, and L3 is used when generating a prediction block, the inter-frame prediction indicator can be specified as PRED_LX for unidirectional prediction of PRED_L0 or PRED_L1, and as PRED_BI_LX for bidirectional prediction, for a reference frame list X. Here, X can refer to a positive integer including 0, such as 0, 1, 2, 3, etc.

[0376] Furthermore, when at least three pieces of information selected from the group including L0 motion information, L1 motion information, L2 motion information, and L3 motion information are used, the inter-frame prediction indicator can be indicated as PRED_TRI for three-way prediction. Furthermore, when at least four pieces of information selected from the group including L0 motion information, L1 motion information, L2 motion information, and L3 motion information are used, the inter-frame prediction indicator can be indicated as PRED_QUAD for four-way prediction.

[0377] For example, when the inter-frame prediction indicator for reference frame list L0 is PRED_L0 and the inter-frame prediction indicator for reference frame list L1 is PRED_BI_L1, the inter-frame prediction indicator for the current block can be PRED_TRI. That is, the sum of the number of prediction blocks indicated by the inter-frame prediction indicators for each reference frame list can be the inter-frame prediction indicator for the current block.

[0378] In addition, at least one list of reference screens may exist, such as L0, L1, L2, L3, etc. For each list of reference screens, a list of reference screens can be generated, such as... Figure 22 and Figure 23The candidate list for merging is shown in the diagram. Therefore, when generating a prediction block for the current block, at least one to at most N prediction blocks can be generated to be used for inter-frame prediction or motion compensation of the current block. Here, N can refer to a positive integer equal to or greater than 1, such as 1, 2, 3, 4, etc.

[0379] To reduce memory bandwidth and improve processing speed, when at least one of the reference screen index and motion vector of a merged candidate is the same as at least one of the reference screen index and motion vector of another merged candidate, or when at least one of the reference screen index and motion vector of a merged candidate is within the prediction range, at least one of the reference screen index and motion vector of the merged candidate may be used when deriving the combined merged candidate.

[0380] For example, among the merge candidates included in the merge candidate list, a merge candidate with a reference screen index equal to a predetermined value can be used when deriving the merge candidate for the combination. Here, the predetermined value can be a positive integer including 0.

[0381] As another example, among the merge candidates included in the merge candidate list, merge candidates within a predetermined range can be used when deriving the combined merge candidates. Here, the predetermined range can be a range of positive integers including 0.

[0382] As another example, among the merge candidates included in the merge candidate list, merge candidates with motion vectors within a predetermined range can be used when deriving the merge candidates for the combination. Here, the predetermined range can be a range of positive integers including 0.

[0383] As another example, among the merge candidates included in the merge candidate list, merge candidates whose motion vector differences between merge candidates fall within a predetermined range can be used when deriving the combined merge candidates. Here, the predetermined range can be a range of positive integers including 0.

[0384] Here, at least one of a predetermined value and a predetermined range can be determined based on a value commonly set in the encoder / decoder. Alternatively, at least one of a predetermined value and a predetermined range can be determined based on a value that is entropy encoded / decoded.

[0385] Furthermore, when deriving modified spatial merging candidates, modified temporal merging candidates, and merging candidates with predetermined motion information values, if at least one of the reference frame index and motion vector of a merging candidate is the same as at least one of the reference frame index and motion vector of another merging candidate, or if at least one of the reference frame index and motion vector of a merging candidate is within the prediction range, at least one of the reference frame index and motion vector of the merging candidate can be used when deriving the combined merging candidate, wherein the combined merging candidate can be derived and added to the merging candidate list.

[0386] Figure 24 This is a diagram illustrating the advantages of deriving combined merging candidates by using only spatial merging candidates in motion compensation using merging modes.

[0387] Reference Figure 24 To increase the throughput of merging modes in the encoder and decoder, combined merging candidates can be derived using only spatial merging candidates without temporal merging candidates. Compared to spatial merging candidate derivation, temporal merging candidate derivation requires a relatively long cycle time due to motion vector scaling. Therefore, when combined merging candidate derivation is performed after temporal merging candidate derivation, a longer cycle time is required when determining motion information using merging modes.

[0388] However, when deriving combined merging candidates using only spatial merging candidates without temporal merging candidates, the combined merging candidate derivation process can be performed immediately after the spatial merging candidate derivation process, which requires a relatively shorter cycle time compared to the temporal merging candidate derivation process. Therefore, compared to methods that include temporal merging candidate derivation, the required cycle time can be reduced when determining motion information using merging patterns.

[0389] In other words, by eliminating the dependency between the temporal merge candidate derivation process and the combined merge candidate derivation process, the throughput of the merging mode can be improved. Furthermore, when errors occur in the reference frame due to transmission errors, the combined merge candidate can be derived using only the spatial merge candidate without using the temporal merge candidate, thereby improving the decoder's error recovery.

[0390] Furthermore, when using a method to derive combined merge candidates by using only spatial merge candidates and not temporal merge candidates, the same approach can be used to derive combined merge candidates by using temporal merge candidates and to derive combined merge candidates without using temporal merge candidates. These methods can be implemented in the same way, thus enabling hardware logic integration.

[0391] Figure 25 This is a diagram illustrating an embodiment of a method for combining bidirectional prediction merging candidates. Here, the combined bidirectional prediction merging candidate can be a merging candidate that includes a combination of two motion information from L0 motion information, ..., LX motion information. In the following description, we will assume that the combined bidirectional prediction merging candidate includes L0 motion information and L1 motion information. Figure 25 .

[0392] Reference Figure 25The encoder / decoder can add L0 motion information and L1 motion information, which are segmented from the combined bidirectional prediction merge candidate information in the merge candidate list, as new merge candidates to the merge candidate list.

[0393] Specifically, in step S2501, the encoder / decoder can determine bidirectional predictive merge candidates from the merge candidate list that are combinations that have been split using a split index (splitIdx). Here, the split index (splitIdx) can be index information indicating bidirectional predictive merge candidates of combinations that have been split.

[0394] In step S2502, the encoder / decoder can set the L0 motion information of the combined bidirectional prediction merging candidate as the motion information of the L0 segmentation candidate, add the motion information to the merging candidate list, and increment numMergeCand by 1.

[0395] The encoder / decoder determines whether the number of merge candidates (numMergeCand) in the current merge candidate list is the same as the maximum number of merge candidates (MaxNumMergeCand). If, in step S2503-Yes, they are the same, the segmentation process can be terminated. Conversely, if, in step S2503-No, they are different, in step S2504, the encoder / decoder can set the L1 motion information of the combined bidirectional predicted merge candidates as the motion information of the L1 segmentation candidates, add this motion information to the merge candidate list, and increment numMergeCand by 1. Next, in step S2505, the segmentation index (splitIdx) can be incremented by 1.

[0396] Furthermore, in step S2506, when the number of merge candidates (numMergeCand) in the current merge candidate list is equal to the maximum number of merge candidates (MaxNumMergeCand), the encoder / decoder may terminate the combined merge candidate segmentation process; otherwise, step S2501 may be executed.

[0397] It can be performed only in the case of B strip / B screen or only in the case of using a strip / screen with at least M reference screen lists. Figure 25 The method shown here is for segmenting bidirectional predictive merge candidates of a combination, wherein the method segments the bidirectional merge candidates. Here, M can be three or four, and can refer to a positive integer equal to or greater than three.

[0398] The segmentation of bidirectional prediction merge candidates of a combination can be performed using at least one of the following methods: 1) segmenting the bidirectional prediction merge candidates of a combination into unidirectional prediction merge candidates when bidirectional prediction merge candidates of a combination exist; 2) segmenting the bidirectional prediction merge candidates of a combination into unidirectional prediction merge candidates when bidirectional prediction merge candidates of a combination exist and the L0 reference screen and L1 reference screen are different from each other in the bidirectional prediction merge candidates of a combination; 3) segmenting the bidirectional prediction merge candidates of a combination into unidirectional prediction merge candidates when bidirectional prediction merge candidates of a combination exist and the L0 reference screen and L1 reference screen are the same in the bidirectional prediction merge candidates of a combination.

[0399] Compared to unidirectional prediction using reconstructed pixel data from a single reference frame, combined bidirectional prediction merging candidates utilize bidirectional prediction and perform motion compensation using reconstructed pixel data from up to two different reference frames, thus resulting in greater memory access bandwidth during motion compensation. Therefore, when segmenting combined bidirectional prediction merging candidates, they are divided into unidirectional prediction merging candidates. Consequently, when the segmented unidirectional prediction merging candidates are determined to be the motion information for the current block, memory access bandwidth during motion compensation can be reduced.

[0400] The encoder / decoder can derive zero-merging candidates with zero motion vectors having motion vectors of (0,0).

[0401] A zero-merging candidate can refer to a merging candidate whose motion vector is (0,0) for at least one of the motion information of L0, L1, L2 and L3.

[0402] Furthermore, a zero-merge candidate can be at least one of two types. A first zero-merge candidate can refer to a merge candidate whose motion vector is (0,0) and whose reference frame index is equal to or greater than zero. A second zero-merge candidate can refer to a merge candidate whose motion vector is (0,0) and whose reference frame index is only zero.

[0403] When the number of merge candidates in the current merge candidate list (numMergeCand) is different from the maximum number of merge candidates (MaxNumMergeCand) (i.e., when the merge candidate list is not full), at least one of the first zero merge candidate and the second zero merge candidate can be repeatedly added to the merge candidate list until the number of merge candidates (numMergeCand) equals the maximum number of merge candidates (MaxNumMergeCand).

[0404] In addition, a first zero merge candidate can be derived and added to the merge candidate list. When the merge candidate list is not full, a second zero merge candidate can be derived and added to the merge candidate list.

[0405] Figure 26 This is a diagram illustrating an embodiment of a method for deriving zero merge candidates. When the number of merge candidates (numMergeCand) in the current merge candidate list is less than the maximum number of merge candidates (MaxNumMergeCand), it can be done as follows: Figure 26 The same order of operations is used to derive the zero-merge candidate.

[0406] First, the encoder / decoder can set the number of merge candidates in the input (numInputMergeCand) to the number of merge candidates in the current merge candidate list (numMergeCand). Additionally, the reference screen index (zeroIdx) for the zero merge candidate can be set to zero. Here, the m-th (numMergeCand) The zero-merge candidate (numInputMergeCand) can be derived.

[0407] In step S2601, the encoder / decoder can determine whether the slice type (slice_type) is a P-slice.

[0408] When the slice type (slice_type) is P-slice in step S2601-, the encoder / decoder can set the number of reference frames (numRefIdx) to the number of available reference frames in the L0 list (num_ref_idx_l0_active_minus1 + 1).

[0409] Furthermore, in step S2602, the encoder / decoder can derive zero-merge candidates as follows and can increment numMergeCand by 1.

[0410] The L0 reference screen index (refIdxL0zeroCandm) of the m-th zero-merge candidate is equal to the reference screen index (zeroIdx) of the zero-merge candidate.

[0411] The L1 reference screen index (refIdxL1zeroCandm) of the m-th zero-merge candidate is -1

[0412] The L0 prediction list for the m-th zero-merge candidate uses the flag (predFlagL0zeroCandm) = 1.

[0413] The L1 prediction list of the m-th zero-merge candidate uses the flag (predFlagL1zeroCandm) = 0.

[0414] The x-component of the L0 motion vector of the mth zero-merging candidate (mvL0zeroCandm[0]) = 0

[0415] The y-component of the L0 motion vector of the mth zero-merging candidate (mvL0zeroCandm[1]) = 0

[0416] The x-component of the L1 motion vector of the mth zero-merging candidate (mvL1zeroCandm[0]) = 0

[0417] The y-component of the m-th zero-merging candidate L1 motion vector (mvL1zeroCandm[1]) = 0

[0418] Conversely, when in step S2601-No, the strip type is not a P strip (it is a B strip or another strip), the number of reference frames (numRefIdx) can be set to a value less than at least one of the following: the number of available reference frames in the L0 list (num_ref_idx_l0_active_minus1 + 1), the number of available reference frames in the L1 list (num_ref_idx_l1_active_minus1 + 1), the number of available reference frames in the L2 list (num_ref_idx_l2_active_minus1 + 1), and the number of available reference frames in the L3 list (num_ref_idx_l3_active_minus1 + 1).

[0419] Next, in step S2603, the encoder / decoder can derive the zero-merge candidate as follows and can increment numMergeCand by 1.

[0420] refIdxL0zeroCandm = zeroIdx

[0421] refIdxL1zeroCandm = zeroIdx

[0422] refIdxL2zeroCandm = zeroIdx

[0423] refIdxL3zeroCandm = zeroIdx

[0424] predFlagL0zeroCandm = 1

[0425] predFlagL1zeroCandm = 1

[0426] predFlagL2zeroCandm = 1

[0427] predFlagL3zeroCandm = 1

[0428] mvL0zeroCandm[0] = 0

[0429] mvL0zeroCandm[1] = 0

[0430] mvL1zeroCandm[0] = 0

[0431] mvL1zeroCandm[1] = 0

[0432] mvL2zeroCandm[0] = 0

[0433] mvL2zeroCandm[1] = 0

[0434] mvL3zeroCandm[0] = 0

[0435] mvL3zeroCandm[1] = 0

[0436] After executing step S2602 or step S2603, when the reference frame count (refCnt) is equal to the number of reference frames (numRefIdx) - 1, in step S2604, the encoder / decoder can set the reference frame index (zeroIdx) of the zero-merging candidate to 0; otherwise, it can increment refCnt and zeroIdx by 1.

[0437] Next, in step S2605, when numMergeCand equals MaxNumMergeCand, the encoder / decoder can terminate the zero-merge candidate derivation process; otherwise, step S2601 can be executed.

[0438] when Figure 26 When the method for deriving zero merge candidates is executed, the derived zero merge candidates can be added to, for example, Figure 27 The list of candidate mergers is shown below.

[0439] Figure 28 This is a diagram illustrating another embodiment of the method for deriving zero merge candidates. When the number of merge candidates (numMergeCand) in the current merge candidate list is less than the maximum number of merge candidates (MaxNumMergeCand), it can be done as follows: Figure 28 The same order is used to perform the derivation of the L0 one-way prediction zero-merging candidate.

[0440] First, the encoder / decoder can set the number of merge candidates in the input (numInputMergeCand) to the number of merge candidates in the current merge candidate list (numMergeCand). Additionally, the reference screen index (zeroIdx) for the zero merge candidate can be set to 0. Here, the m-th (numMergeCand) The zero-merge candidate (numInputMergeCand) can be derived. Furthermore, the number of reference frames (numRefIdx) can be set to the number of available reference frames in the L0 list (num_ref_idx_l0_active_minus1 + 1).

[0441] In step S2801, the encoder / decoder can derive zero-merge candidates as follows and can increment numMergeCand by 1.

[0442] The L0 reference screen index (refIdxL0zeroCandm) of the m-th zero-merge candidate is equal to the reference screen index (zeroIdx) of the zero-merge candidate.

[0443] The L1 reference screen index (refIdxL1zeroCandm) of the m-th zero-merge candidate is -1

[0444] The L0 prediction list for the m-th zero-merge candidate uses the flag (predFlagL0zeroCandm) = 1.

[0445] The L1 prediction list of the m-th zero-merge candidate uses the flag (predFlagL1zeroCandm) = 0.

[0446] The x-component of the L0 motion vector of the mth zero-merging candidate (mvL0zeroCandm[0]) = 0

[0447] The y-component of the L0 motion vector of the mth zero-merging candidate (mvL0zeroCandm[1]) = 0

[0448] The x-component of the L1 motion vector of the mth zero-merging candidate (mvL1zeroCandm[0]) = 0

[0449] The y-component of the m-th zero-merging candidate L1 motion vector (mvL1zeroCandm[1]) = 0

[0450] When the reference frame count (refCnt) equals numRefIdx – 1, in step S2802, the encoder / decoder can set zeroIdx to 0; otherwise, refCnt and zeroIdx can be incremented by 1.

[0451] Next, when numMergeCand equals MaxNumMergeCand in step S2803-Yes, the encoder / decoder can terminate the zero-merge candidate derivation step; otherwise, step S2801 can be executed in step S2803-No.

[0452] exist Figure 26 In the method for deriving zero-merge candidates, either bidirectional predictive zero-merge candidate derivation or L0 unidirectional predictive zero-merge candidate derivation is performed based on the stripe type. Therefore, two implementation methods are required depending on the stripe type.

[0453] exist Figure 28 In the method for deriving zero-merging candidates, neither bidirectional nor L0 unidirectional predictive zero-merging candidate derivation is performed based on the stripe type. Regardless of the stripe type, L0 unidirectional predictive zero-merging candidates are derived, thereby simplifying the hardware logic. Therefore, the cycle time required for deriving zero-merging candidates can be reduced. Furthermore, when L0 unidirectional predictive zero-merging candidates, rather than bidirectional predictive zero-merging candidates, are determined to be the motion information for the current block, unidirectional predictive motion compensation is performed instead of bidirectional predictive motion compensation. Therefore, memory access bandwidth can be reduced during motion compensation.

[0454] For example, except in the case of P-stripes, L0 unidirectional prediction zero merge candidates can be derived and added to the merge candidate list.

[0455] The encoder / decoder can add merge candidates other than zero merge candidates to the merge candidate list, and can subsequently add L0 unidirectional prediction zero merge candidates. Furthermore, the encoder / decoder can initialize the merge candidate list using L0 unidirectional prediction zero merge candidates, and can subsequently add spatial merge candidates, temporal merge candidates, combined merge candidates, zero merge candidates, additional merge candidates, etc., to the initialized merge candidate list.

[0456] Figure 29 This diagram illustrates an embodiment of deriving and sharing a merge candidate list in a CTU. The merge candidate list can be shared among blocks smaller than a predetermined block or blocks deeper than a predetermined block. Here, the size or depth of the predetermined block can be the size or depth of a block whose motion compensation information is entropy-encoded / decoded. Furthermore, the size or depth of the predetermined block can be information entropy-encoded in the encoder and entropy-decoded in the decoder, and can be a value commonly preset in both the encoder and decoder.

[0457] Reference Figure 29 When the predetermined block size is 128×128, blocks smaller than 128×128 ( Figure 29 Blocks in the slashed area can share the merge candidate list.

[0458] Next, the motion information of the current block will be determined in detail by using the merge candidate list generated in steps S1202 and S1303.

[0459] The encoder can determine the merge candidate to be used in motion compensation from the merge candidate list through motion estimation, and can encode the merge candidate index (merge_idx) indicating the determined merge candidate in the bitstream.

[0460] Simultaneously, to generate a prediction block, the encoder can select a merging candidate from the merging candidate list based on the merging candidate index to determine the motion information of the current block. Here, the prediction block of the current block can be generated by performing motion compensation based on the determined motion information.

[0461] For example, when the merge candidate index is three, the merge candidate indicated by merge candidate index 3 in the merge candidate list can be identified as motion information and can be used in motion compensation of the encoded target block.

[0462] The decoder decodes the merge candidate indices in the bitstream to determine the merge candidate indicated by the merge candidate index in the merge candidate list. The determined merge candidate can be used as motion information for the current block. The determined motion information can be used in motion compensation for the current block. Here, motion compensation can represent inter-frame prediction.

[0463] For example, when the merge candidate index is two, the merge candidate indicated by merge candidate index 2 in the merge candidate list can be identified as motion information and can be used in motion compensation of the encoded target block.

[0464] Furthermore, by changing at least one value of the information corresponding to the motion information of the current block, the motion information can be used in inter-frame prediction or motion compensation of the current block. Here, the changed value of the information corresponding to the motion information can be at least one of the x-component of the motion vector, the y-component of the motion vector, and the reference frame index. Moreover, when changing at least one value of the information corresponding to the motion information, the at least one value of the information corresponding to the motion information can be changed such that minimum distortion is indicated by using a distortion calculation method (SAD, SSE, MSE, etc.).

[0465] Next, the motion compensation of the current block will be performed in steps S1203 and S1304 by using determined motion information.

[0466] The encoder and decoder can perform inter-frame prediction or motion compensation by using motion information from determined merging candidates. Here, the current block (the target block for encoding / decoding) may have motion information from determined merging candidates.

[0467] The current block may have at least one to at most N motion information depending on the prediction direction. The final prediction block of the current block can be derived by using at least one to at most N prediction blocks generated from the motion information.

[0468] For example, when the current block has a motion information, the predicted block generated by using the motion information can be determined as the final predicted block of the current block.

[0469] Conversely, when the current block has several motion information pieces, several prediction blocks can be generated using these motion information pieces, and the final prediction block of the current block can be determined based on the weighted sum of these prediction blocks. Reference frames that each include several prediction blocks indicated by several motion information pieces can be included in different reference frame lists, or they can be included in the same reference frame list. Furthermore, when the current block has several motion information pieces, the motion information pieces of several reference frames can indicate the same reference frame.

[0470] For example, multiple prediction blocks can be generated based on at least one of spatial merging candidates, temporal merging candidates, modified spatial merging candidates, modified temporal merging candidates, merging candidates with predetermined motion information values ​​or combinations thereof, and additional merging candidates. The final prediction block of the current block can be determined based on the weighted sum of the multiple prediction blocks.

[0471] As another example, multiple prediction blocks can be generated based on merge candidates indicated by a preset merge candidate index. The final prediction block of the current block can be determined based on the weighted sum of the multiple prediction blocks. Furthermore, multiple prediction blocks can be generated based on merge candidates existing within a preset merge candidate index range. The final prediction block of the current block can be determined based on the weighted sum of the multiple prediction blocks.

[0472] The weighting factor applied to each prediction block can have the same value, 1 / N (where N is the number of prediction blocks generated). For example, when two prediction blocks are generated, the weighting factor applied to each prediction block can be 1 / 2. When three prediction blocks are generated, the weighting factor applied to each prediction block can be 1 / 3. When four prediction blocks are generated, the weighting factor applied to each prediction block can be 1 / 4. Optionally, the final prediction block for the current block can be determined by applying different weighting factors to the prediction blocks.

[0473] The weighting factor does not need to have a fixed value for each prediction block and can have a variable value for each prediction block. Here, the weighting factors applied to prediction blocks can be the same or different from each other. For example, when two prediction blocks are generated, the weighting factors applied to both prediction blocks for each block can be variable values, such as (1 / 2, 1 / 2), (1 / 3, 2 / 3), (1 / 4, 3 / 4), (2 / 5, 3 / 5), (3 / 8, 5 / 8), etc. At the same time, the weighting factor can be a positive real number or a negative real number. For example, the weighting factor can be a negative real number, such as (-1 / 2, 3 / 2), (-1 / 3, 4 / 3), (-1 / 4, 5 / 4), etc.

[0474] Simultaneously, to apply the variable weighting factor, at least one weighting factor information for the current block can be transmitted via signaling through the bitstream. Weighting factor information can be transmitted via signaling for each prediction block, and weighting factor information can be transmitted via signaling for each reference frame. Multiple prediction blocks can share one weighting factor information.

[0475] The encoder and decoder can determine whether to use motion information from merging candidates based on flags in the prediction block list. For example, when the prediction block list flag indicates a first value of 1 for each list of reference frames, it indicates that the encoder and decoder can use motion information from the merging candidates of the current block to perform inter-frame prediction or motion compensation. When the prediction block list flag indicates a second value of 0, it indicates that the encoder and decoder should not perform inter-frame prediction or motion compensation by using motion information from the merging candidates of the current block. Simultaneously, the first value of the prediction block list flag can be set to 0, and its second value can be set to 1.

[0476] Equations 3 through 5 below show examples of generating the final predicted block for the current block when the inter-frame prediction indicator for each current block is PRED_BI (or when the current block can use two motion information), PRED_TRI (or when the current block can use three motion information), and PRED_QUAD (or when the current block can use four motion information), and the prediction direction for each reference frame list is unidirectional.

[0477] [Formula 3]

[0478]

[0479] [Formula 4]

[0480]

[0481] [Formula 5]

[0482]

[0483] In Formulas 3 to 5, the final prediction block of the current block can be specified as P_BI, P_TRI, and P_QUAD, and the reference frame list can be specified as LX (X=0, 1, 2, 3). The weighting factor value of the prediction block generated using LX can be specified as WF_LX. The offset value of the prediction block generated using LX can be specified as OFFSET_LX. The prediction block generated using motion information from LX for the current block can be specified as P_LX. The rounding factor can be specified as RF and can be set to 0, a positive number, or a negative number. The LX reference frame list can include at least one of the following: long-term reference frame, reference frame without deblocking filtering, reference frame without sample adaptive offset, reference frame without adaptive loop filtering, reference frame with deblocking filtering and adaptive offset, reference frame with deblocking filtering and adaptive loop filtering, reference frame with sample adaptive offset and adaptive loop filtering, and reference frame with deblocking filtering, sample adaptive offset, and adaptive loop filtering. In this case, the LX reference screen list can be at least one of the L0 reference screen list, L1 reference screen list, L2 reference screen list, and L3 reference screen list.

[0484] When there are multiple prediction directions for a predetermined list of reference screens, the final prediction block for the current block can be obtained based on the weighted sum of the prediction blocks. Here, the weighting factors applied to prediction blocks derived from the same list of reference screens can have the same value or different values.

[0485] At least one of the weighting factor (WF_LX) and offset (OFFSET_LX) for multiple prediction blocks can be an encoding parameter that is entropy encoded / decoded.

[0486] As another example, weighting factors and offsets can be derived from neighboring encoded / decoded blocks adjacent to the current block. Here, neighboring blocks of the current block may include at least one of the blocks used when deriving spatial merge candidates for the current block and the blocks used when deriving temporal merge candidates for the current block.

[0487] As another example, the weighting factor and offset can be determined based on the display order (POC) between the current frame and the reference frame. In this case, the weighting factor or offset can be set to a smaller value when the current frame is far from the reference frame, and a larger value when the current frame is close to the reference frame. For example, when the POC difference between the current frame and the L0 reference frame is 2, the weighting factor applied to the prediction block generated by referring to the L0 reference frame can be set to 1 / 3. Conversely, when the POC difference between the current frame and the L0 reference frame is 1, the weighting factor applied to the prediction block generated by referring to the L0 reference frame can be set to 2 / 3. As mentioned above, the weighting factor or offset value can be inversely proportional to the display order difference between the current frame and the reference frame. As another example, the weighting factor or offset value can be directly proportional to the display order difference between the current frame and the reference frame.

[0488] As another example, at least one of the weighting factor and offset can be entropy-encoded / decoded based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks. Furthermore, a weighted sum of the predicted blocks can be calculated based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks.

[0489] The weighted sum of multiple prediction blocks can be applied only to a portion of the prediction block. Here, the portion of the prediction block can be the region corresponding to the boundary. As mentioned above, in order to apply the weighted sum only to a portion of the prediction block, the weighted sum can be applied to each sub-block of the prediction block.

[0490] Next, the entropy encoding / decoding of information about motion compensation in steps S1204 and S1301 will be described in detail.

[0491] Figure 30 and Figure 31 This is a diagram illustrating an example of the syntax for information about motion compensation. Figure 30 This shows an example of the syntax for information about motion compensation in a coding unit. Figure 31 This shows an example of the syntax for information about motion compensation in a prediction unit.

[0492] The encoding device can entropy encode motion compensation information in the bitstream, and the decoding device can entropy decode the motion compensation information included in the bitstream. Here, the motion compensation information entropy-encoded / decoded may include at least one of the following: information on whether a skip mode is used (cu_skip_flag), information on whether a merge mode is used (merge_flag), merge index information (merge_index), inter-frame prediction indicator (inter_pred_idc), weighting factor values ​​(wf_l0, wf_l1, wf_l2, wf_l3), and offset values ​​(offset_l0, offset_l1, offset_l2, offset_l3). Entropy encoding / decoding of motion compensation information can be performed in at least one of the CTU, the coded block, and the prediction block.

[0493] When the information regarding whether to use skip mode (cu_skip_flag) has a first value of 1, it indicates that skip mode should be used. When the information regarding whether to use skip mode (cu_skip_flag) has a second value of 2, it indicates that skip mode should not be used. Motion compensation for the current block can be performed by using skip mode based on the information regarding whether to use skip mode.

[0494] When the information regarding whether to use the merge mode (merge_flag) has a first value of 1, it indicates that the merge mode should be used. When the information regarding whether to use the merge mode (merge_flag) has a second value of 2, it indicates that the merge mode should not be used. Motion compensation for the current block can be performed by using the merge mode based on the information regarding whether to use the merge mode.

[0495] The merge index information (merge_index) can indicate information about the merge candidates in the merge candidate list.

[0496] In addition, merged index information can refer to information about merged indexes.

[0497] In addition, the merge index information can indicate the deduced merge candidate blocks among the reconstructed blocks that are spatially / temporally adjacent to the current block.

[0498] Furthermore, the merge index information can indicate at least one motion information of a merge candidate. For example, when the merge index information has a first value of 0, it can indicate the first merge candidate in the merge candidate list. When the merge index information has a second value of 1, it can indicate the second merge candidate in the merge candidate list. When the merge index information has a third value of 2, it can indicate the third merge candidate in the merge candidate list. Similarly, when the merge index information has a fourth to Nth value, it can indicate the merge candidate corresponding to that value according to the order in the merge candidate list. Here, N can refer to a positive integer including 0.

[0499] Motion compensation for the current block can be performed by using the merge mode, based on the merge mode index information.

[0500] When encoding / decoding the current block under inter-frame prediction, the inter-frame prediction indicator can indicate at least one of the inter-frame prediction direction and the number of prediction directions for the current block. For example, the inter-frame prediction indicator can indicate unidirectional or multi-directional prediction (such as bidirectional, tridirectional, quadridirectional, etc.). The inter-frame prediction indicator can also indicate the number of reference frames used when the current block generates a prediction block. Optionally, a reference frame can be used for multi-directional prediction. In this case, M reference frames are used to perform N (N>M) directional prediction. The inter-frame prediction indicator can also indicate the number of prediction blocks used when performing motion compensation or the inter-frame prediction of the current block.

[0501] As described above, based on the inter-frame prediction indicator, the number of reference frames used when generating the prediction block for the current block, the number of prediction blocks used when performing inter-frame prediction or motion compensation for the current block, or the number of reference frames available for the current block can be determined. Here, the number of reference frames is a positive integer N, such as 1, 2, 3, 4, or greater. For example, the reference frame list may include L0, L1, L2, L3, etc. Motion compensation can be performed on the current block using at least one reference frame list.

[0502] For example, by using at least one list of reference frames, the current block can generate at least one prediction block to perform motion compensation for the current block. For example, motion compensation can be performed by generating one or more prediction blocks using reference frame list L0. Optionally, motion compensation can be performed by generating one or more prediction blocks using reference frame lists L0 and L1. Optionally, motion compensation can be performed by generating one or more prediction blocks or up to N prediction blocks (where N is a positive integer equal to or greater than 3 or 2) using reference frame lists L0, L1, L2, and L3. Optionally, motion compensation can be performed by generating one or more prediction blocks or up to N prediction blocks (where N is a positive integer equal to or greater than 4 or 2) using reference frame lists L0, L1, L2, and L3.

[0503] The reference screen indicator can indicate one-way (PRED_LX), two-way (PRED_BI), three-way (PRED_TRI), four-way (PRED_QUAD) or more directions depending on the number of predicted directions for the current block.

[0504] For example, when performing unidirectional prediction for each reference frame list, the inter-frame prediction indicator PRED_LX can indicate that a prediction block is generated using the reference frame list LX (where X is an integer such as 0, 1, 2, or 3), and inter-frame prediction or motion compensation is performed using the generated prediction block. The inter-frame prediction indicator PRED_BI can indicate that two prediction blocks are generated using at least one of the reference frame lists L0, L1, L2, and L3, and inter-frame prediction or motion compensation is performed using the generated two prediction blocks. The inter-frame prediction indicator PRED_TRI can indicate that three prediction blocks are generated using at least one of the reference frame lists L0, L1, L2, and L3, and inter-frame prediction or motion compensation is performed using the generated three prediction blocks. The inter-frame prediction indicator PRED_QUAD can indicate that four prediction blocks are generated using at least one of the reference frame lists L0, L1, L2, and L3, and inter-frame prediction or motion compensation is performed using the generated four prediction blocks. In other words, the sum of the number of prediction blocks used when performing inter-frame prediction for the current block can be set as the inter-frame prediction indicator.

[0505] When performing multi-directional prediction for a reference frame list, the inter-frame prediction indicator PRED_BI can indicate that bi-directional prediction is performed for the L0 reference frame list. The inter-frame prediction indicator PRED_TRI can indicate: performing tri-directional prediction for the L0 reference frame list; performing uni-directional prediction for the L0 reference frame list and performing bi-directional prediction for the L1 reference frame list; or performing bi-directional prediction for the L0 reference frame list and performing uni-directional prediction for the L1 reference frame list.

[0506] As described above, the inter-frame prediction indicator may indicate the generation of at least one to at most N (here, N is the number of prediction directions indicated by the inter-frame prediction indicator) prediction blocks from at least one list of reference frames to perform motion compensation. Optionally, the inter-frame prediction indicator may indicate the generation of at least one to at most N prediction blocks from N reference frames and the performance of motion compensation for the current block using the generated prediction blocks.

[0507] For example, the inter-frame prediction indicator PRED_TRI can indicate that inter-frame prediction or motion compensation for the current block is performed by generating three prediction blocks using at least one of the L0, L1, L2, and L3 reference frames. Optionally, the inter-frame prediction indicator PRED_TRI can indicate that inter-frame prediction or motion compensation for the current block is performed by generating three prediction blocks using at least three selected from the group including the L0, L1, L2, and L3 reference frames. Furthermore, the inter-frame prediction indicator PRED_QUAD can indicate that inter-frame prediction or motion compensation for the current block is performed by generating four prediction blocks using at least one of the L0, L1, L2, and L3 reference frames. Optionally, the inter-frame prediction indicator PRED_QUAD can indicate that inter-frame prediction or motion compensation for the current block is performed by generating four prediction blocks using at least four selected from the group including the L0, L1, L2, and L3 reference frames.

[0508] Available inter-frame prediction directions can be determined based on the inter-frame prediction indicator, and all or some of the available inter-frame prediction directions can be used optionally based on the size and / or shape of the current block.

[0509] The prediction list uses flags to indicate whether prediction blocks are generated by using a reference screen list.

[0510] For example, when the prediction list uses a flag indicating a first value of 1, it indicates that prediction blocks are generated using a reference screen list. When the prediction list uses a flag indicating a second value of 0, it indicates that prediction blocks are not generated using a reference screen list. Here, the first value of the prediction list using the flag can be set to 0, and the second value of the prediction list using the flag can be set to 1.

[0511] In other words, when the prediction list uses a flag to indicate the first value, the prediction block for the current block can be generated by using motion information corresponding to the reference screen list.

[0512] Additionally, the prediction list usage flag can be set based on the inter-frame prediction indicator. For example, when the inter-frame prediction indicator indicates PRED_LX, PRED_BI, PRED_TRI, or PRED_QUAD, the prediction list usage flag predFlagLX can be set to a first value of 1. When the inter-frame prediction indicator is PRED_LN (where N is a positive integer other than X), the prediction list usage flag predFlagLX can be set to a second value of 0.

[0513] Furthermore, the inter-frame prediction indicator can be set based on flags used in the prediction list. For example, when the prediction list uses flags predFlagL0 and predFlagL1 to indicate a first value of 1, the inter-frame prediction indicator can be set to PRED_BI. For example, when only the prediction list uses flag predFlagL0 to indicate a first value of 1, the inter-frame prediction indicator can be set to PRED_L0.

[0514] When two or more prediction blocks are generated during motion compensation for the current block, the final prediction block for the current block can be generated by weighted summation for each prediction block. When calculating the weighted sum, at least one of a weighting factor and an offset can be applied to each prediction block. The weighting factor (such as a weighting factor or offset) used in calculating the weighted sum can be entropy-encoded / decoded for at least one of the following: a list of reference frames, reference frames, motion vector candidate indices, motion vector differences, motion vectors, information about whether a skip mode is used, information about whether a merge mode is used, and merge index information. Furthermore, the weighting factor for each prediction block can be entropy-encoded / decoded based on an inter-frame prediction indicator. Here, the weighting factor can include at least one of a weighting factor and an offset.

[0515] The weighting factor can be derived from index information of one of the predetermined settings in the specified encoding and decoding devices. In this case, entropy encoding / decoding can be performed on the index information used to specify at least one of the weighting factor and offset. Predetermined settings in the encoder and decoder can be defined for the weighting factor and offset, respectively. The predetermined settings may include at least one weighting factor candidate or at least one offset candidate. Optionally, a table defining the mapping relationship between the weighting factor and the offset can be used. In this case, the weighting factor value and offset value for the prediction block can be obtained from the table using a single index information. Entropy encoding / decoding can be performed on the index information regarding the offset mapped to each index information for the entropy-encoded / decoded weighting factor.

[0516] Information related to weighting and factors can be entropy encoded / decoded at block units, and can also be entropy encoded / decoded at higher levels. For example, weighting factors or offsets can be entropy encoded / decoded at block units such as CTU, CU, or PU, or at higher levels such as video parameter sets, sequence parameter sets, picture parameter sets, adaptive parameter sets, or strip headers.

[0517] Entropy encoding / decoding of weighted sum factors can be performed based on the weighted sum factor and the weighted sum factor difference between the weighted sum factor and the weighted sum factor predicted values. For example, the weighted sum factor predicted value and the weighted sum factor difference can be entropy encoded / decoded, or the offset predicted value and the offset difference can be entropy encoded / decoded. Here, the weighted sum factor difference can indicate the difference between the weighted sum factor and the weighted sum factor predicted value, and the offset difference can indicate the difference between the offset and the offset predicted value.

[0518] Here, the weighted sum factor difference is entropy-encoded / decoded by block unit, and the weighted sum factor prediction can be entropy-encoded / decoded at a higher level. When the weighted sum factor prediction (such as weighted factor prediction or offset prediction) is entropy-encoded / decoded by frame or strip unit, blocks included in a frame or strip can use a common weighted sum factor prediction.

[0519] Weighted sum factor predictions can be derived from specific regions within an image, strip, or parallel block, or from specific regions within a CTU or CU. For example, weighted factor values ​​or offset values ​​from specific regions within an image, strip, parallel block, CTU, or CU can be used as weighted factor predictions or offset predictions. In this case, entropy encoding / decoding of the weighted sum factor predictions can be omitted, and only entropy encoding / decoding of the weighted sum factor differences can be performed.

[0520] Optionally, the weighted sum factor prediction value can be derived from neighboring encoded / decoded blocks adjacent to the current block. For example, the weighted factor value or offset value of the neighboring encoded / decoded blocks adjacent to the current block can be set as the weighted factor prediction value or offset prediction value of the current block. Here, the neighboring blocks of the current block may include at least one of the blocks used when deriving spatial merge candidates and the blocks used when deriving temporal merge candidates.

[0521] When using weighted factor predictions and weighted factor differences, the decoding device can calculate the weighted factor value for the prediction block by adding the weighted factor predictions and weighted factor differences. Similarly, when using offset predictions and offset differences, the decoding device can calculate the offset value for the prediction block by adding the offset predictions and offset differences.

[0522] Entropy encoding / decoding of the weighted sum factor or the weighted sum factor difference can be performed based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks.

[0523] Based on at least one of the coding parameters of the current block, neighboring blocks, and co-occurring blocks, the weighted sum factor, weighted sum factor prediction, or weighted sum factor difference can be derived as the weighted sum factor, weighted sum factor prediction, or weighted sum factor difference of the current block.

[0524] The weighted sum factor of the encoded / decoded blocks adjacent to the current block can be used as the weighted sum factor of the current block without entropy encoding / decoding information about the weighted sum factor of the current block. For example, the weighted sum factor or offset of the current block can be set to have the same value as the weighted sum factor or offset of the encoded / decoded neighboring blocks adjacent to the current block.

[0525] Motion compensation can be performed on the current block by using at least one of the weighted sum factors, or by using at least one of the derived weighted sum factors.

[0526] Weighted sums and factors can be included in information about motion compensation.

[0527] At least one of the aforementioned pieces of information regarding motion compensation can be entropy encoded / decoded in at least one of the CTU and its sub-units (sub-CTUs). Here, the sub-units of the CTU may include at least one of the CU and PU. The blocks of the sub-units of the CTU may have a square shape or a non-square shape. For convenience, the information regarding motion compensation may refer to at least one piece of information regarding motion compensation.

[0528] When entropy encoding / decoding information about motion compensation is performed in a CTU, motion compensation can be performed on all or some blocks present in that CTU based on the value of the information about motion compensation.

[0529] When entropy encoding / decoding information about motion compensation is performed in a CTU or a sub-unit of a CTU, the information about motion compensation can be entropy encoded / decoded based on at least one of the size and depth of a predetermined block.

[0530] Here, entropy encoding / decoding can be performed on information about the size or depth of the predicted block. Optionally, information about the size or depth of the predetermined block can be determined based on at least one of preset values ​​and encoding parameters in the encoder and decoder, or based on at least one of other syntax element values.

[0531] Entropy encoding / decoding of motion compensation information can be performed only in blocks with a size greater than or equal to a predetermined block, and can be omitted in blocks with a size smaller than the predetermined block. In this case, motion compensation can be performed on sub-blocks within the block with a size greater than or equal to the predetermined block based on the motion compensation information entropy-encoded / decoded in the block. That is, sub-blocks within the block with a size greater than or equal to the predetermined block can share motion compensation information, which includes motion vector candidates, motion vector candidate lists, merge candidates, and merge candidate lists, etc.

[0532] Information regarding motion compensation can be entropy-encoded / decoded only in blocks with a depth shallower than or equal to the predetermined block, and not in blocks deeper than the predetermined block. In this case, motion compensation can be performed on sub-blocks within the block with a depth shallower than or equal to the predetermined block based on the motion compensation information entropy-encoded / decoded in the block. That is, sub-blocks within the block with a depth shallower than or equal to the predetermined block can share information regarding motion compensation, which includes motion vector candidates, motion vector candidate lists, merge candidates, merge candidate lists, etc.

[0533] For example, when entropy encoding / decoding information about motion compensation is performed in a 32×32 sub-cell within a 64×64 block size CTU, motion compensation can be performed on blocks smaller than 32×32 included in the 32×32 block based on the motion compensation information entropy encoded / decoded in the 32×32 block.

[0534] As another example, when entropy encoding / decoding information about motion compensation is performed in a 16×16 sub-cell within a 128×128 block size CTU, motion compensation can be performed on blocks whose size is less than or equal to 16×16 within the 16×16 block based on the motion compensation information entropy encoded / decoded in the 16×16 block.

[0535] As another example, when entropy encoding / decoding information about motion compensation is performed in a sub-cell of block depth 1 within a CTU of block depth 0, motion compensation can be performed on blocks included in the block of depth 1 that are deeper than the block of depth 1, based on the motion compensation information entropy encoded / decoded in the block of depth 1.

[0536] For example, when entropy encoding / decoding of at least one piece of information about motion compensation is performed in a sub-cell of block depth 2 in a CTU of block depth 0, motion compensation can be performed on blocks whose depth is equal to or deeper than that of the block of block depth 2 based on the motion compensation information entropy encoded / decoded in the block of block depth 2.

[0537] Here, the value of block depth can be a positive integer including 0. A larger block depth value indicates a deeper depth. A smaller block depth value indicates a shallower depth. Therefore, a larger block depth value allows for a smaller block size, and a smaller block depth value allows for a larger block size. Furthermore, the depth of a sub-block of a predetermined block can be greater than the depth of the predetermined block, and within the block corresponding to the predetermined block, the depth of a sub-block of the predetermined block can also be greater than the depth of the predetermined block.

[0538] Information about motion compensation can be entropy encoded / decoded for each block, and this can also be done at higher levels. For example, information about motion compensation can be entropy encoded / decoded for each block such as CTU, CU, or PU, or at higher levels such as video parameter sets, sequence parameter sets, frame parameter sets, adaptive parameter sets, or strip headers.

[0539] Information about motion compensation can be entropy-encoded / decoded based on information differences related to motion compensation, where the information difference indicates the difference between the information about motion compensation and the predicted value of the information about motion compensation. Considering an inter-frame prediction indicator (i.e., a piece of information about motion compensation) as an example, the predicted value of the inter-frame prediction indicator and the inter-frame prediction indicator difference can be entropy-encoded / decoded. Here, the inter-frame prediction indicator difference can be entropy-encoded / decoded for each block, and the predicted value of the inter-frame prediction indicator can be entropy-encoded / decoded at a higher level. The predicted value of information about motion compensation (such as the predicted value of the inter-frame prediction indicator, etc.) can be entropy-encoded / decoded for each frame or strip, and blocks in a frame or strip can use a common predicted value of information about motion compensation.

[0540] Motion compensation information predictions can be derived from specific regions within an image, strip, or parallel block, or from specific regions within a CTU or CU. For example, inter-frame prediction indicators for specific regions within an image, strip, parallel block, CTU, or CU can be used as inter-frame prediction indicator predictions. In this case, entropy encoding / decoding of the motion compensation information predictions can be omitted, and only the information difference regarding motion compensation can be entropy encoded / decoded.

[0541] Optionally, prediction values ​​for motion compensation information can be derived from neighboring encoded / decoded blocks adjacent to the previous block. For example, the inter-frame prediction indicator of a neighboring encoded / decoded block adjacent to the current block can be set as the prediction value of the current block's inter-frame prediction indicator. Here, the neighboring blocks of the current block may include at least one of the blocks used when deriving spatial merging candidates and the blocks used when deriving temporal merging candidates. Furthermore, neighboring blocks may have the same depth as the current block or a smaller depth than the current block. When multiple neighboring blocks exist, one neighboring block may be selectively used according to a predetermined priority. The neighboring block used to predict information about motion compensation may have a fixed position based on the current block and may have a variable position depending on the position of the current block. Here, the position of the current block may be based on the position of the frame or strip that includes the current block, or it may be based on the position of the CTU, CU, or PU that includes the current block.

[0542] Merged index information can be calculated using index information in the predefined settings of the encoder and decoder.

[0543] When using the predicted value of motion compensation and the difference of motion compensation information, the decoding device can calculate the motion compensation information value for the prediction block by adding the predicted value of motion compensation information and the difference of motion compensation information.

[0544] Entropy encoding / decoding of motion compensation information or information differences related to motion compensation can be performed based on at least one of the encoding parameters of the current block, neighboring blocks, and co-occurring blocks.

[0545] Based on at least one of the coding parameters of the current block, neighboring blocks, and co-occurring blocks, information about motion compensation, a predicted value of information about motion compensation, or a difference in information about motion compensation can be derived as information about motion compensation, a predicted value of information about motion compensation, or a difference in information about motion compensation for the current block.

[0546] Motion compensation information from adjacent encoded / decoded blocks can be used as motion compensation information for the current block without entropy encoding / decoding the motion compensation information of the current block. For example, the inter-frame prediction indicator of the current block can be set to the same value as the inter-frame prediction indicator of the adjacent encoded / decoded blocks.

[0547] Furthermore, at least one piece of information regarding motion compensation may have a preset fixed value in the encoder and decoder. The preset fixed value may be determined as the value of at least one piece of information regarding motion compensation. Blocks within a specific block whose size is smaller than that specific block may share at least one piece of information regarding motion compensation with a preset fixed value. Similarly, blocks with a depth greater than that specific block and which are sub-blocks of that specific block may share at least one piece of information regarding motion compensation with a preset fixed value. Here, the fixed value may be a positive integer including 0, or it may be an integer vector value including (0,0).

[0548] Here, sharing at least one piece of information about motion compensation can mean that multiple blocks have the same value for at least one piece of information about motion compensation, or it can mean that motion compensation is performed on multiple blocks by using the same value for at least one piece of information about motion compensation.

[0549] Information regarding motion compensation may also include at least one of the following: motion vector, motion vector candidate, motion vector candidate index, motion vector difference, motion vector prediction value, information on whether to use skip mode (skip_flag), information on whether to use merge mode (merge_flag), merge index information (merge_index), motion vector resolution information, overlapping block motion compensation information, local illumination compensation information, affine motion compensation information, decoder-side motion vector derivation information, and bidirectional optical flow information.

[0550] Motion vector resolution information can be information indicating whether a specific resolution is used for at least one of the motion vector and the motion vector difference. Here, resolution can represent precision. Furthermore, the specific resolution can be set to at least one of integer pixel (integer-pel) units, 1 / 2 pixel (1 / 2-pel) units, 1 / 4 pixel (1 / 4-pel) units, 1 / 8 pixel (1 / 8-pel) units, 1 / 16 pixel (1 / 16-pel) units, 1 / 32 pixel (1 / 32-pel) units, and 1 / 64 pixel (1 / 64-pel) units.

[0551] Overlapping block motion compensation information can be information indicating whether the weighted sum of the predicted blocks of the current block is calculated by using the motion vectors of the spatially adjacent neighboring blocks during the motion compensation of the current block.

[0552] Local illumination compensation information can be information indicating whether at least one of a weighting factor value and an offset value is applied when generating a prediction block for the current block. Here, at least one of the weighting factor value and the offset value can be a value calculated based on a reference block.

[0553] Affine motion compensation information can be information indicating whether an affine motion model is used during motion compensation of the current block. Here, the affine motion model can be a model used to partition a block into several sub-blocks using multiple parameters and to calculate the motion vectors of the partitioned sub-blocks using typical motion vectors.

[0554] Decoder-side motion vector derivation information can indicate whether to use the motion vectors required for motion compensation through derivation by the decoder. Based on the decoder-side motion vector derivation information, information about motion vectors may not be entropy encoded / decoded. Furthermore, when the decoder-side motion vector derivation information instructs the decoder to derive and use motion vectors, information about the merging mode can be entropy encoded / decoded. In other words, decoder-side motion vector derivation information can indicate whether a merging mode is used in the decoder.

[0555] Bidirectional optical flow information can be information indicating whether motion compensation is performed by correcting the motion vector for each pixel or sub-block. Based on bidirectional optical flow information, the motion vector for each pixel or sub-block may not be entropy encoded / decoded. Here, motion vector correction can be modifying the motion vector value from a block-per-motion vector to a pixel-per-sub-block motion vector.

[0556] Motion compensation is performed on the current block by using at least one piece of information about motion compensation, and entropy encoding / decoding can be performed on at least one piece of information about motion compensation.

[0557] Figure 32 This is a diagram illustrating an embodiment of using a merge mode in blocks smaller than a predetermined block in a CTU.

[0558] Reference Figure 32 When the size of the predefined block is 8×8, blocks smaller than 8×8 (diagonal blocks) can use the merge mode.

[0559] Furthermore, when comparing sizes between blocks, a size smaller than a predetermined block can indicate that the sum of the sample points in the block is smaller. For example, a 32×16 block may have 512 sample points, therefore a 32×16 block is smaller than a 32×32 block with 1024 sample points. A 4×16 block may have 64 sample points, therefore the size of a 4×16 block can be equal to that of an 8×8 block.

[0560] When entropy encoding / decoding information about motion compensation, binary methods can be used, such as truncated Rice binary method, K-order exponential Columbus binary method, finite K-order exponential Columbus binary method, fixed-length binary method, unary binary method, or truncated unary binary method, etc.

[0561] When entropy encoding / decoding information about motion compensation, the context model can be determined by using at least one of the following: information about motion compensation from neighboring blocks adjacent to the current block or region information of neighboring blocks; information about previously encoded / decoded motion compensation or region information of previously encoded / decoded blocks; information about the depth of the current block; and information about the size of the current block.

[0562] Furthermore, when entropy encoding / decoding information about motion compensation, entropy encoding / decoding can be performed by using at least one of the following: information about motion compensation from neighboring blocks, information about previously encoded / decoded motion compensation, information about the depth of the current block, and information about the size of the current block, as a prediction value for the motion compensation information of the current block.

[0563] Encoding / decoding processing can be performed on each of the luma and chroma signals. For example, in the encoding / decoding processing, at least one of the following methods can be applied differently to the luma and chroma signals: obtaining inter-frame prediction indicators, generating a merge candidate list, deriving motion information, and performing motion compensation.

[0564] Encoding / decoding processing can be performed identically for both luma and chroma signals. For example, in encoding / decoding processing applied to luma signals, at least one of the following can be applied identically to chroma signals: inter-frame prediction indicator, merge candidate list, merge candidate, reference frame, and reference frame list.

[0565] These methods can be performed in the encoder and decoder in the same manner. For example, in encoding / decoding processing, at least one of the following methods can be applied in the encoder and decoder: obtaining inter-frame prediction indicators, generating a merge candidate list, deriving motion information, and performing motion compensation. Furthermore, the order in which these methods are applied may differ between the encoder and decoder.

[0566] Embodiments of the invention can be applied based on the size of at least one of the coding block, prediction block, block, and unit. Here, the size can be defined as a minimum and / or maximum size for applying the embodiment, and can be defined as a fixed size to which the embodiment is applied. Furthermore, a first embodiment can be applied according to a first size, and a second embodiment can be applied according to a second size. That is, embodiments can be applied in various ways depending on the size. Moreover, embodiments of the invention can be applied only when the size is equal to or greater than the minimum size and equal to or less than the maximum size. That is, embodiments can be applied only when the block size is within a predetermined range.

[0567] For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 8×8. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 16×16. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 32×32. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 64×64. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 128×128. For example, embodiments can be applied only when the size of the encoded / decoded target block is 4×4. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or less than 8×8. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or less than 16×16. For example, embodiments can be applied only when the size of the encoded / decoded target block is equal to or greater than 8×8 and equal to or less than 16×16. For example, the implementation example can be applied only when the size of the encoded / decoded target block is equal to or greater than 16×16 and equal to or less than 64×64.

[0568] Embodiments of the invention can be applied according to time layers. Identifiers for identifying the time layers to which embodiments can be applied can be transmitted using signals, and embodiments can be applied to the time layers specified by the identifiers. Here, the identifiers can be defined as indicating the lowest and / or highest layers to which embodiments can be applied, and can be defined as indicating a specific layer to which embodiments can be applied.

[0569] For example, the embodiment can be applied only when the time layer of the current frame is the lowest layer. For example, the embodiment can be applied only when the time layer identifier of the current frame is 0. For example, the embodiment can be applied only when the time layer identifier of the current frame is equal to or greater than 1. For example, the embodiment can be applied only when the time layer of the current frame is the highest layer.

[0570] As described in the embodiments of the present invention, the reference screen set used in the process of reference screen list creation and reference screen list modification may use at least one of reference screen lists L0, L1, L2 and L3.

[0571] According to an embodiment of the present invention, when the deblocking filter calculates the boundary strength, at least one to at most N motion vectors of the encoded / decoded target block can be used. Here, N indicates a positive integer equal to or greater than 1, such as 2, 3, 4, etc.

[0572] In motion vector prediction, embodiments of the invention can be applied when the motion vector has at least one of the following: 16-pixel (16-pel) units, 8-pixel (8-pel) units, 4-pixel (4-pel) units, integer-pixel (integer-pel) units, 1 / 2-pixel (1 / 2-pel) units, 1 / 4-pixel (1 / 4-pel) units, 1 / 8-pixel (1 / 8-pel) units, 1 / 16-pixel (1 / 16-pel) units, 1 / 32-pixel (1 / 32-pel) units, and 1 / 64-pixel (1 / 64-pel) units. Furthermore, in the merge mode, the motion vector can be optionally used for each pixel unit.

[0573] The strip type of the embodiments of the present invention can be defined, and the embodiments of the present invention can be applied according to the strip type.

[0574] For example, when the stripe type is T (three-way prediction), prediction blocks can be generated using at least three motion vectors, and the final prediction block for the encoding / decoding target block can be obtained by calculating the weighted sum of at least three prediction blocks. Similarly, when the stripe type is Q (four-way prediction), prediction blocks can be generated using at least four motion vectors, and the final prediction block for the encoding / decoding target block can be obtained by calculating the weighted sum of at least four prediction blocks.

[0575] The embodiments of the present invention can be applied to inter-frame prediction and motion compensation methods using merging mode, inter-frame prediction and motion compensation methods using motion vector prediction, inter-frame prediction and motion compensation methods using skip mode, etc.

[0576] The shape of the block to which the embodiments of the present invention are applied may be a regular shape or a non-square shape.

[0577] The above is 12 to Figure 32 A method for encoding and decoding video using a merging pattern according to the present invention is described. In the following text, reference will be made to... Figure 33 and Figure 34 The present invention describes in detail a method for decoding video, a method for encoding video, a device for decoding video, a device for encoding video, and a bitstream.

[0578] Figure 33 This is a diagram illustrating a method for decoding video according to the present invention.

[0579] Reference Figure 33 In step S3301, a merge candidate list for the current block can be generated, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of the plurality of reference screen lists.

[0580] Here, the merge candidate corresponding to each of the multiple reference frame lists can represent a merge candidate with LX motion information corresponding to the LX of the reference frame list. For example, there are L0 merge candidates with L0 motion information, L1 merge candidates with L1 motion information, L2 merge candidates with L2 motion information, L3 merge candidates with L3 motion information, etc.

[0581] Simultaneously, the merge candidate list includes at least one of the following: spatial merge candidates derived from spatially neighboring blocks of the current block, temporal merge candidates derived from co-located blocks of the current block, modified spatial merge candidates derived by modifying spatial merge candidates, modified temporal merge candidates derived by modifying temporal merge candidates, and merge candidates with predefined motion information values. Here, a merge candidate with a predefined motion information value can be a zero merge candidate.

[0582] In this scenario, spatial merge candidates can be derived from the sub-blocks of neighboring blocks adjacent to the current block. Furthermore, temporal merge candidates can be derived from the sub-blocks of co-located blocks of the current block.

[0583] Additionally, the merge candidate list may include merge candidates derived by using a combination of at least two merge candidates selected from a group that includes spatial merge candidates, temporal merge candidates, modified spatial merge candidates, and modified temporal merge candidates.

[0584] Furthermore, in step S3302, at least one piece of motion information can be determined by using the generated merge candidate list.

[0585] Furthermore, in step S3303, a predicted block for the current block can be generated using at least one piece of determined motion information.

[0586] Here, the step of generating the prediction block of the current block in step S3303 may include: generating multiple temporal prediction blocks according to the inter-frame prediction indicator of the current block; and generating the prediction block of the current block by applying at least one of a weighting factor and an offset to the generated multiple temporal prediction blocks.

[0587] In this case, at least one of the weighting factor and offset can be shared in blocks smaller than the predetermined block or in blocks deeper than the predetermined block.

[0588] Simultaneously, the list of merge candidates can be shared in blocks smaller than the predetermined block or in blocks deeper than the predetermined block.

[0589] Furthermore, when the size of the current block is smaller than the predetermined block or the depth of the current block is greater than the predetermined block, a list of merge candidates can be generated based on the higher-level blocks of the current block. The size or depth of the higher-level blocks is equal to that of the predetermined block.

[0590] The prediction block for the current block can be generated by applying information about the weighted sum to multiple prediction blocks generated based on multiple merge candidates or multiple lists of merge candidates.

[0591] Figure 34 This is a diagram illustrating a method for encoding video according to the present invention.

[0592] Reference Figure 34 In step S3401, a merge candidate list for the current block can be generated, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of the plurality of reference screen lists.

[0593] In step S3402, at least one piece of motion information can be determined by using the generated merge candidate list.

[0594] Furthermore, in step S3403, a predicted block for the current block can be generated using at least one piece of determined motion information.

[0595] The apparatus for decoding video according to the present invention may include an inter-frame prediction unit, wherein the inter-frame prediction unit generates a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists, the inter-frame prediction unit determines at least one piece of motion information by using the merge candidate list, and generates a predicted block for the current block by using the determined at least one piece of motion information.

[0596] The apparatus for encoding video according to the present invention may include an inter-frame prediction unit, wherein the inter-frame prediction unit generates a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists, the inter-frame prediction unit determines at least one piece of motion information by using the merge candidate list, and generates a predicted block for the current block by using the determined at least one piece of motion information.

[0597] The bitstream according to the present invention can be a bitstream generated by a method for encoding video, the method comprising: generating a merge candidate list for a current block, wherein the merge candidate list for the current block includes at least one merge candidate corresponding to each of a plurality of reference frame lists; determining at least one piece of motion information by using the merge candidate list; and generating a predicted block for the current block by using the determined at least one piece of motion information.

[0598] In the above embodiments, the method is described based on a flowchart having a series of steps or units. However, the present invention is not limited to the order of the steps; rather, some steps may be performed simultaneously with other steps, or may be performed with other steps in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart are not mutually exclusive, and other steps may be added to the flowchart, or some steps may be deleted from the flowchart, without affecting the scope of the present invention.

[0599] The embodiments include various aspects of the examples. All possible combinations of these aspects may not be described, but those skilled in the art will recognize the different combinations. Therefore, the invention may include all substitutions, modifications, and alterations within the scope of the claims.

[0600] Embodiments of the present invention can be implemented in the form of program instructions, which can be executed by various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include individual program instructions, data files, data structures, etc., or combinations thereof. The program instructions recorded in the computer-readable recording medium may be specially designed and constructed for the present invention, or may be known to those skilled in the art of computer software. Examples of computer-readable recording media include: magnetic recording media (such as hard disks, floppy disks, and magnetic tapes); optical data storage media (such as CD-ROMs or DVD-ROMs); magneto-optical media (such as floppy disks); and hardware devices (such as read-only memory (ROM), random access memory (RAM), flash memory, etc.) specially constructed for storing and implementing program instructions. Examples of program instructions include not only machine language code generated by a compiler, but also high-level language code that can be implemented by a computer using an interpreter. The hardware device may be configured to operate by one or more software modules to perform the processing according to the present invention, or vice versa.

[0601] Although the invention has been described with reference to specific terminology (such as detailed elements) and limited embodiments and drawings, these are provided only to aid in a more general understanding of the invention, and the invention is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made from the above description.

[0602] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the full scope of the appended claims and their equivalents shall fall within the scope and spirit of the present invention.

[0603] Industrial availability

[0604] This invention can be used in devices for encoding / decoding images.

Claims

1. A method for decoding video, the method comprising: Generate a candidate list for the current block, wherein the candidate list for the current block includes at least one of spatial candidates derived from spatial neighboring blocks of the current block and temporal candidates derived from co-located blocks of the current block; Motion information is determined by using a candidate list; Multiple prediction blocks are generated based on deterministic motion information; and The final prediction block for the current block is generated based on the weighted sum of the generated multiple prediction blocks. The weighting factors used for the weighted sum are derived from the index information of multiple weighting factors included in the weighting factor set, specifying the weighting factors applied to the current block. Specifically, the index information is obtained from the bitstream only when the size of the current block is greater than or equal to a predefined value. The current block is obtained by dividing the coded tree into blocks based on partitioning information transmitted from the bit stream via signals, and The partitioning information includes partitioning flags that indicate whether to partition the coded tree into blocks.

2. The method as described in claim 1, wherein, The set of weighting factors includes positive weighting factors and negative weighting factors.

3. A method for encoding video, the method comprising: Generate a candidate list for the current block, wherein the candidate list for the current block includes at least one of spatial candidates derived from spatial neighboring blocks of the current block and temporal candidates derived from co-located blocks of the current block; Motion information is determined by using a candidate list; Multiple prediction blocks are generated based on deterministic motion information; and The final prediction block for the current block is generated based on the weighted sum of the generated multiple prediction blocks. Specifically, the index information specifying the weighting factors applied to the current block for weighted summation from multiple weighting factors included in the weighting factor set is encoded. Specifically, the index information is encoded into the bitstream only when the size of the current block is greater than or equal to a predefined value. The current block is obtained by dividing the coded tree into blocks based on partitioning information transmitted from the bit stream via signals, and The partitioning information includes partitioning flags that indicate whether to partition the coded tree into blocks.

4. A method for transmitting a bit stream, the method comprising: Perform the method for encoding video as described in claim 3 to generate the bitstream; as well as Send the bit stream.

Citation Information

Patent Citations

  • Method and apparatus for predicting motion vector for coding video or decoding video

    CN104488272A

  • A method and an apparatus for processing a multi-view video signal

    KR1020160001647A