Methods and apparatus for encoding / decoding images

By generating geometrically modified reference images and performing inter-frame prediction, the problem of reduced similarity between the reference image and the current image is solved, achieving more efficient image encoding and decoding.

CN116489350BActive Publication Date: 2026-01-06ELECTRONICS & TELECOMM RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310357623.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2015-11-20
Filing Date
2016-11-18
Publication Date
2026-01-06
Estimated Expiration
2036-11-18

AI Technical Summary

Technical Problem

In existing technologies, the reduced similarity between the reference image and the current image leads to a decrease in inter-frame prediction efficiency, necessitating improvements in image encoding and decoding methods to enhance prediction efficiency.

Method used

A geometrically modified reference image is generated by generating a geometrically modified reference image, and inter-frame prediction and intra-frame prediction are performed based on the geometric modification information. The optimal prediction block is selected for encoding, and the geometric modification information is generated for encoding and decoding.

Benefits of technology

It effectively encodes and decodes images, improves the efficiency of inter-frame prediction, and enhances the effect of image compression and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116489350B_ABST
    Figure CN116489350B_ABST
Patent Text Reader

Abstract

This invention relates to a method and apparatus for encoding / decoding images. The method for decoding an image includes: determining whether a merging mode is applied to a current block; when the merging mode is applied to the current block, determining whether geometric modification information is used for inter-frame prediction of the current block based on geometric modification usage information of the current block; when the geometric modification usage information indicates that the geometric modification information is used for inter-frame prediction of the current block, obtaining a predicted block of the current block within a decoded target image by performing inter-frame prediction based on a reference image and the geometric modification information; and obtaining a reconstructed block of the current block by summing the predicted block and the residual block, wherein the geometric modification used for the geometric modification information includes affine modification, and wherein the geometric modification information and the reference image index of the specified reference image are derived from merging candidates in a merging candidate list of the current block, and the merging candidate list includes at least one spatial merging candidate derived from at least one adjacent block of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application filed on November 18, 2016, with application number 201680067705.2 and entitled "Method for Encoding / Decoding Images and Apparatus for Encoding / Decoding Images". Technical Field

[0002] This invention generally relates to methods and apparatus for encoding / decoding images by using geometrically modified images generated by geometrically modifying a reference image. Background Technology

[0003] With the expansion of high-definition (HD) broadcasting nationwide and worldwide, many users have become accustomed to images with high resolution and high picture quality. Therefore, many organizations are driving the development of the next generation of imaging equipment. Furthermore, with the continued growth in interest in ultra-high definition (UHD), which has resolutions several times higher than HDTV, there is a need for technologies capable of compressing and processing images with even higher resolution and higher picture quality.

[0004] Various image compression techniques exist, such as inter-frame prediction, which predicts the pixel values ​​included in the current image based on images preceding or following it; intra-frame prediction, which predicts the pixel values ​​included in the current image using pixel information from the current image; energy transformation and quantization techniques for compressing residual signals; and entropy coding, where short codes are assigned to values ​​with high frequency of occurrence, and long codes are assigned to values ​​with low frequency of occurrence. Using these image compression techniques, image data can be transmitted and stored while being effectively compressed.

[0005] When global motion is included in a reference image referenced during inter-frame prediction, the similarity between the reference image and the current image decreases. This reduced similarity can lead to decreased prediction efficiency. Therefore, improvements are needed to address this issue. Summary of the Invention

[0006] Technical issues

[0007] The present invention aims to provide a method and apparatus for efficiently encoding / decoding images.

[0008] In addition, the present invention provides a method and apparatus for performing intra-frame prediction and / or inter-frame prediction by referring to a reference image and / or a geometrically modified reference image.

[0009] In addition, the present invention provides a method and apparatus for encoding information related to geometrically modified images.

[0010] The technical objectives to be achieved by this invention are not limited to those described above, and those skilled in the art will understand from the following description other technical objectives not described.

[0011] Technical solution

[0012] According to one aspect of the present invention, a method for encoding an image is provided. The method may include: generating a geometrically modified reference image by geometrically modifying a reference image; generating a prediction block of a current block within a target image by performing inter-frame prediction with reference to the reference image or the geometrically modified reference image; and encoding inter-frame prediction information of the current block.

[0013] According to the encoding method of the present invention, the method may further include: generating geometric modification information based on the relationship between the encoding target image and the reference image, and the geometric modification reference image may be generated based on the geometric modification information.

[0014] According to the encoding method of the present invention, the prediction block can be selected from either a first prediction block generated by inter-frame prediction with reference to the reference image or a second prediction block generated by inter-frame prediction with reference to the geometrically modified reference image.

[0015] According to the encoding method of the present invention, the selected prediction block can be selected for each of the first prediction block and the second prediction block based on the encoding efficiency of the current block.

[0016] According to the encoding method of the present invention, the encoding efficiency of the current block can be determined based on the rate distortion cost.

[0017] According to the encoding method of the present invention, the encoding of inter-frame prediction information can be performed by making predictions based on the inter-frame prediction information of the blocks adjacent to the current block.

[0018] According to the encoding method of the present invention, encoding inter-frame prediction information may include: generating a candidate list constructed using one or more candidate blocks adjacent to the current block; selecting a candidate block included in the candidate list; and encoding information for identifying the selected candidate block among the candidate blocks included in the candidate list.

[0019] According to the encoding method of the present invention, the inter-frame prediction information may include geometric modification usage information, and the geometric modification usage information may be information used to indicate whether the reference image or the geometrically modified reference image is used to generate the prediction block of the current block.

[0020] According to the encoding method of the present invention, when a prediction block for the current block is generated using bidirectional or more directional prediction, the geometric modification usage information is encoded in a symbol value corresponding to a combination of information about each prediction direction and geometric modification usage information about each prediction direction.

[0021] According to the encoding method of the present invention, the encoding of the symbol value can be performed by encoding the difference between the symbol value and the previously used symbol value corresponding to the geometric modification usage information.

[0022] According to another aspect of the present invention, a method for decoding an image is provided. The method may include: decoding inter-frame prediction information of a current block; and generating a predicted block of the current block within a target image by performing inter-frame prediction based on the inter-frame prediction information.

[0023] According to the decoding method of the present invention, the inter-frame prediction information may include geometric modification usage information, and the geometric modification usage information may be information used to indicate whether a reference image or a geometrically modified reference image is used to generate the prediction block of the current block.

[0024] According to the decoding method of the present invention, when the geometric modification usage information indicates the use of a geometric modification reference image, the decoding method may further include: decoding the geometric modification information; and generating a geometric modification reference image by geometrically modifying the reference image based on the geometric modification usage information, and the prediction block of the current block may be generated by inter-frame prediction with reference to the geometric modification reference image.

[0025] According to the decoding method of the present invention, the inter-frame prediction information can be decoded by predicting the inter-frame prediction information of the blocks adjacent to the current block.

[0026] According to the decoding method of the present invention, decoding inter-frame prediction information may include: generating a candidate list constructed using one or more candidate blocks adjacent to the current block; decoding information for identifying a candidate block among the candidate blocks included in the candidate list; selecting a candidate block from the candidate blocks included in the candidate list based on the identification information; and deriving inter-frame prediction information of the current block by using the inter-frame prediction information of the selected candidate block.

[0027] According to the decoding method of the present invention, the method may further include determining the inter-frame prediction information of the selected candidate block as the inter-frame prediction information of the current block.

[0028] According to the decoding method of the present invention, when bidirectional or more directional predictions are used to generate a prediction block for the current block, the geometric modification usage information can be obtained by decoding in a symbol value corresponding to a combination of information for each prediction direction and geometric modification usage information for each prediction direction.

[0029] According to the decoding method of the present invention, decoding a symbol value may include: decoding the difference between the symbol value and a previously used symbol value corresponding to the geometric modification usage information; and adding the previously used symbol value corresponding to the geometric modification usage information to the decoded difference.

[0030] According to the decoding method of the present invention, geometric modification information can be generated based on the relationship between the decoded target image and the reference image, and can have various forms, such as global motion information (global motion vector), transfer geometric modification matrix, size geometric modification matrix, rotation geometric modification matrix, affine geometric modification matrix and projection geometric modification matrix.

[0031] Still according to another aspect of the invention, an apparatus for encoding an image is provided. The apparatus may include: a geometrically modified reference image generator for generating a geometrically modified reference image by geometrically modifying a reference image; an inter-frame prediction unit for generating a prediction block of the current block within an encoded target image by performing inter-frame prediction with reference to the reference image or the geometrically modified reference image; and an encoder for encoding the inter-frame prediction information of the current block.

[0032] According to another aspect of the present invention, an apparatus for decoding an image is provided. The apparatus may include: a decoder for decoding inter-frame prediction information of a current block; and an inter-frame prediction unit for generating a predicted block of the current block within a decoded target image by performing inter-frame prediction based on the inter-frame prediction information, wherein the inter-frame prediction information includes geometric modification usage information, and wherein the geometric modification usage information may be information indicating whether a reference image or a geometrically modified reference image is used to generate the predicted block of the current block.

[0033] Beneficial effects

[0034] According to the present invention, images can be efficiently encoded / decoded.

[0035] Furthermore, according to the present invention, inter-frame prediction and / or intra-frame prediction can be performed by referring to a reference image and / or a geometrically modified image.

[0036] Furthermore, according to the present invention, information related to geometrically modified images can be efficiently encoded.

[0037] The effects obtained by the present invention are not limited to those described above, and other unmentioned effects can be clearly understood by those skilled in the art based on the following description. Attached Figure Description

[0038] Figure 1 This is a block diagram illustrating the configuration of an image encoding apparatus to which an embodiment of the present invention is applied.

[0039] Figure 2 This is a block diagram illustrating the configuration of an image decoding apparatus to which an embodiment of the present invention is applied.

[0040] Figure 3It is a diagram that schematically illustrates the partition structure of an image when the image is encoded.

[0041] Figure 4 This is a diagram showing the form of a prediction unit (PU) that can be included in a coding unit (CU).

[0042] Figure 5 This is a diagram showing the form of a transform unit (TU) that can be included in a coding unit (CU).

[0043] Figure 6 This is a diagram illustrating an example of intra-frame prediction processing.

[0044] Figure 7 This is a diagram illustrating an example of inter-frame prediction processing.

[0045] Figure 8 This is a diagram illustrating a transfer modification of an embodiment of geometric modification of an image according to the present invention.

[0046] Figure 9 This is a diagram illustrating the size modification of an embodiment of geometric modification of an image according to the present invention.

[0047] Figure 10 This is a diagram illustrating a rotational modification of an embodiment of geometric modification of an image according to the present invention.

[0048] Figure 11 This is a diagram illustrating an embodiment of geometric modification of an image according to the present invention, specifically an affine modification.

[0049] Figure 12 This is a diagram illustrating a projection modification of an embodiment of geometric modification of an image according to the present invention.

[0050] Figure 13 This is a diagram illustrating an example of a method for implementing homography according to the present invention.

[0051] Figure 14 This is an example method according to the present invention for deriving the relationship between two corresponding points between two images.

[0052] Figure 15 This is a diagram illustrating a method for generating a geometrically modified image based on a geometric modification matrix and an original image according to the present invention.

[0053] Figure 16 This is a diagram illustrating a method for generating a geometrically modified image using inverse mapping according to the present invention.

[0054] Figure 17The figure illustrates a method for generating a geometrically modified image based on a geometric modification matrix and an original image according to the present invention, wherein the geometric modification matrix may correspond to geometric modification information.

[0055] Figure 18 Reference numerals illustrating embodiments of the present invention are shown. Figure 17 A graph showing bilinear interpolation among various interpolation methods.

[0056] Figure 19 This is a diagram illustrating motion prediction using reference images and / or geometrically modified images according to an embodiment of the present invention.

[0057] Figure 20 This is a block diagram illustrating the configuration of an image encoding apparatus to which another embodiment of the present invention is applied.

[0058] Figure 21 This is a block diagram illustrating the configuration of an image decoding apparatus to which another embodiment of the present invention is applied.

[0059] Figure 22 This is a diagram illustrating the operation of an encoder including a prediction unit that uses geometric modification information according to an embodiment of the present invention.

[0060] Figure 23 This is a diagram illustrating the operation of an encoder that includes geometric modification using an information predictor according to an embodiment of the present invention.

[0061] Figure 24 This is a diagram illustrating a method for generating geometric modification information according to an embodiment of the present invention.

[0062] Figure 25 This is a configuration diagram of an inter-frame prediction unit including a geometric modification information prediction portion according to an embodiment of the present invention.

[0063] Figure 26 This is a diagram illustrating an example encoding method for using information to explain geometric modifications.

[0064] Figure 27 This is a diagram illustrating another example of an encoding method that uses information to explain geometric modifications.

[0065] Figure 28 This is a diagram illustrating an example of encoding an image using a merging pattern.

[0066] Figure 29 This is a diagram showing adjacent blocks included in the merge candidate list.

[0067] Figure 30 This is a diagram explaining the steps involved in generating a geometrically modified reference image.

[0068] Figure 31This is an example diagram illustrating the configuration of inter-frame prediction information included in blocks within the encoded target image.

[0069] Figure 32 This is a diagram illustrating an embodiment of predicting inter-frame prediction information for the current block X within the current image.

[0070] Figure 33 This is a diagram illustrating an embodiment of predicting inter-frame prediction information for the current block based on merging candidates.

[0071] Figure 34 This is a diagram illustrating an embodiment of a method for predicting inter-frame prediction information for the current block based on a merged candidate list. Detailed Implementation

[0072] Because various modifications can be made to the invention and various embodiments of the invention exist, examples will now be provided with reference to the accompanying drawings, and the invention will be described in detail thereon. However, the invention is not limited thereto, and exemplary embodiments may be construed as including all modifications, equivalents, or substitutions within the technical concept and scope of the invention. In various aspects, similar reference numerals refer to the same or similar functions. In the drawings, the shape and size of elements may be exaggerated for clarity, and the same reference numerals are used throughout to designate the same or similar elements. In the following detailed description of the invention, reference is made to the accompanying drawings, which illustrate specific embodiments in which the invention can be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure. It should be understood that the various embodiments of the disclosure, though different, are not necessarily mutually exclusive. For example, specific features, structures, and characteristics described herein in conjunction with one embodiment may be implemented in other embodiments without departing from the spirit and scope of the disclosure. Furthermore, it should be understood that the position or arrangement of various elements within each disclosed embodiment may be modified without departing from the spirit and scope of the disclosure. Therefore, the following detailed description should not be considered limiting, and the scope of the disclosure, and the entire scope equivalent to the scope claimed in the claims, is defined and properly interpreted only by the appended claims.

[0073] The terms "first," "second," etc., used in this specification may be used to describe various components, but these components should not be construed as limited to these terms. These terms are only used to distinguish one component from other components. For example, a "first" component may be referred to as a "second" component without departing from the scope of the invention, and a "second" component may similarly be referred to as a "first" component. The term "and / or" includes a combination of multiple items or any one of multiple items.

[0074] When it is said that a component is "coupled" or "connected" to another component, it may mean that it is directly coupled or connected to another component, but it should be understood that there may be another component between them. On the other hand, when it is said that a component is "directly coupled" or "directly connected" to another component, it should be understood that there are no other components between them.

[0075] Furthermore, the constituent parts shown in the embodiments of the present invention are illustrated independently to represent different functional characteristics. Therefore, this does not mean that each constituent part is constructed as a separate hardware or software unit. In other words, for convenience, each constituent part includes every one of the listed constituent parts. Thus, at least two constituent parts of each constituent part can be combined to form one constituent part, or a constituent part can be divided into multiple constituent parts to perform each function. Embodiments combining each constituent part and embodiments in which a constituent part is divided are also included within the scope of the present invention without departing from its spirit.

[0076] The terminology used in this specification is for describing particular embodiments only and is not intended to limit the invention. Expressions used in the singular include expressions used in the plural unless they have a clearly different meaning in the context. It should be understood in this specification that terms such as “comprising,” “having,” etc., are intended to indicate the presence of features, numbers, steps, actions, elements, portions, or combinations thereof disclosed in the specification, and are not intended to exclude the possibility that one or more other features, numbers, steps, actions, elements, portions, or combinations thereof may be present or may be added. In other words, when a particular element is referred to as “comprising,” elements other than the corresponding element are not excluded, but rather additional elements may be included within the embodiments of the invention or within the scope of the invention.

[0077] Furthermore, some components may not be essential for performing the basic functions of the invention, but rather optional components that only improve its performance. In addition to components used to improve performance, the invention can be implemented by including only the essential, indispensable components for carrying out the invention. Structures containing only indispensable components, besides optional components used only to improve performance, are also included within the scope of the invention.

[0078] In the following, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. In describing exemplary embodiments of the present invention, well-known functions or structures will not be described in detail, as they may unnecessarily obscure the understanding of the invention. The same component elements in the drawings are denoted by the same reference numerals, and repeated descriptions of the same elements will be omitted.

[0079] Additionally, in the following text, "image" can refer to the pictures that make up the video, or it can refer to the video itself. For example, "encoding and / or decoding images" can refer to "encoding and / or decoding video," or it can refer to "encoding and / or decoding individual images among the images that make up the video." Here, "picture" can refer to the image itself.

[0080] Encoder: can refer to an encoding device.

[0081] Decoder: can refer to a decoding device.

[0082] Explanation: It can refer to determining the value of a syntax element by performing entropy decoding, or it can refer to an entropy decoder.

[0083] A block can refer to a sample of an M x N matrix, where M and N are positive integers. A block can also refer to a sample matrix of a two-dimensional matrix.

[0084] Unit: This can refer to a unit used for encoding or decoding an image. When encoding and decoding an image, a unit can be a region generated by partitioning the image. Alternatively, when an image is subdivided and encoded or decoded, a unit can refer to the divided units of an image. During encoding and decoding, predetermined processing can be performed on each unit. A single unit can be divided into smaller sub-units. Depending on its function, a unit can also refer to a block, macroblock (MB), coding unit (CU), prediction unit (PU), transform unit (TU), coded block (CB), prediction block (PB), or transform block (TB). A unit can refer to an object including a luma component block for each block to distinguish it from a block, its corresponding chroma component block, and syntax elements. The unit may have different sizes and shapes. Specifically, the shape of a unit can include two-dimensional forms such as rectangles, cubes, trapezoids, triangles, pentagons, etc. Additionally, the shape of a unit can include geometric shapes. Furthermore, unit information can include at least one of the following: unit type (such as coding unit, prediction unit, transform unit, etc.), unit size, unit depth, and the sequence of unit encoding and decoding.

[0085] Reconstructing adjacent units: can refer to reconstructed units that have already been encoded or decoded in space / time and are adjacent to the encoding / decoding target unit.

[0086] Depth: Indicates the degree of partitioning of a unit. In a tree structure, the highest node can refer to the root node, and the lowest node can refer to the leaf node.

[0087] Symbols: can refer to the syntax elements and encoding parameters of the encoding / decoding target unit, the values ​​of transform coefficients, etc.

[0088] Parameter set: This can correspond to the header information in the structure within the bitstream. At least one of the video parameter set, sequence parameter set, image parameter set, and adaptive parameter set can be included in the parameter set. Additionally, the parameter set may include information from slice headers and tile headers.

[0089] Bitstream: can refer to a string of bits that includes encoded image information.

[0090] Encoding parameters can include not only information encoded by the encoder and sent to the decoder along with syntax elements, but also information that can be derived during encoding or decoding, or parameters necessary for encoding and decoding. For example, encoding parameters can include at least one of the following values ​​and / or statistics: intra-frame prediction mode, inter-frame prediction mode, intra-frame prediction direction, motion information, motion vector, reference image index, inter-frame prediction direction, inter-frame prediction indicator, reference image list, motion vector predictor, motion merging candidate, transform type, transform size, information on whether an additional transform is used, filter information within the loop, information on the presence of residual signals, quantization parameters, context model, transform coefficients, transform coefficient levels, encoded block pattern, encoded block flags, image display / output order, slice information, tile information, image type, information on whether a motion merging mode is used, information on whether a skip mode is used, block size, block depth, block partitioning information, cell size, cell partitioning information, etc.

[0091] Prediction unit: This can refer to the basic unit used when performing inter-frame or intra-frame prediction, as well as when compensating for prediction. A prediction unit can be divided into multiple partitions. Each partition can also be a basic unit used when performing inter-frame or intra-frame prediction, as well as when compensating for prediction. A partitioned prediction unit can also refer to a prediction unit in general. Furthermore, a single prediction unit can be divided into smaller sub-units. Prediction units can have various sizes and shapes. Specifically, the shape of a unit can include two-dimensional forms such as rectangles, squares, trapezoids, triangles, pentagons, etc. Additionally, the shape of a unit can include geometric shapes.

[0092] Prediction unit partitioning: can refer to the partitioning form of prediction units.

[0093] Reference image list: This can refer to a list that includes at least one reference image for inter-frame prediction or motion compensation. The type of reference list can include a combined list (LC), L0 (list 0), L1 (list 1), L2 (list 2), L3 (list 3), etc. At least one reference image list can be used for inter-frame prediction.

[0094] Inter-frame prediction indicator: This can indicate the inter-frame prediction direction (one-way prediction, two-way prediction) of the encoded / decoded target block. Alternatively, the indicator can indicate the number of reference images used to generate the prediction blocks for the encoded / decoded target block, or the number of prediction blocks used when motion compensation is performed on the encoded / decoded target block.

[0095] Reference image index: can refer to the index of a specific image within the reference image list.

[0096] Reference image: This can refer to a reference image used by a specific unit for inter-frame prediction or motion compensation. Alternatively, reference picture can refer to a reference image.

[0097] Motion vector: Refers to a two-dimensional matrix used for inter-frame prediction or motion compensation, or it can be the offset between the encoded / decoded target image and the reference image. For example, (mvX, mvY) can indicate a motion vector, where mvX can be the horizontal component and mvY can be the vertical component.

[0098] Motion vector candidate: can refer to a cell that becomes a prediction candidate when predicting motion vectors, or can refer to the motion vector of a cell.

[0099] Motion vector candidate list: This can refer to a list of candidates configured with motion vectors.

[0100] Motion vector candidate index: can refer to an indicator that points to a motion vector candidate in the list of motion vector candidates, or it can refer to an index of the motion vector predictor.

[0101] Motion information: can refer to at least one of the following: motion vector, reference image index, inter-frame prediction indicator, reference image list information, reference image, motion vector candidate, motion vector candidate index, etc.

[0102] Transform unit: This refers to a basic unit when performing transformations such as transform coefficients, inverse transforms, quantization, dequantization, and encoding / decoding of residual signals. A single unit can be divided into smaller subunits. These units may have different sizes and shapes. Specifically, the shape of a unit can include two-dimensional forms such as rectangles, squares, trapezoids, triangles, pentagons, etc. Additionally, the shape of a unit can also include geometric shapes.

[0103] Scaling can refer to the process of multiplying a factor by the levels of transformation coefficients, with the result generating transformation coefficients. Scaling can also refer to inverse quantization.

[0104] Quantization parameter: This can refer to the value used to scale the transform coefficient levels during quantization and inverse quantization. Here, the quantization parameter can be a value mapped to the quantization step size.

[0105] Differential quantization parameter: can refer to the residual value between the predictive quantization parameter and the quantization parameter of the encoding / decoding target unit.

[0106] Scan: This can refer to a method of sorting the coefficients within a block or matrix. For example, sorting a two-dimensional matrix into a one-dimensional matrix can refer to a scan or inverse scan.

[0107] Transformation coefficients: These can be coefficient values ​​generated after performing a transformation. In this invention, the transformation coefficient levels quantized by applying quantization to the transformation coefficients can be included in the transformation coefficients.

[0108] Non-zero transform coefficients: can refer to transform coefficients whose values ​​or magnitudes are not zero.

[0109] Quantization matrix: This can refer to the matrix used for quantization and inverse quantization to improve image quality. It can also refer to a scaling list.

[0110] Quantization matrix coefficients: can refer to each element of the quantization matrix. Quantization matrix coefficients can also refer to matrix coefficients.

[0111] Default matrix: can refer to a predefined quantization matrix that is defined in advance in the encoder and decoder.

[0112] Non-default matrix: can refer to the quantization matrix sent / received by the user and not defined in advance in the encoder and decoder.

[0113] Figure 1 This is a block diagram illustrating the configuration of an image encoding apparatus applying an embodiment of the present invention.

[0114] The encoding device 100 can be a video encoding device or an image encoding device. The video may include at least one image. The encoding device 100 can encode at least one image of the video in chronological order.

[0115] refer to Figure 1 The encoding device 100 may include a motion prediction unit 111, a motion compensation unit 112, an intra-frame prediction unit 120, a switch 115, a subtractor 125, a transform unit 130, a quantization unit 140, an entropy coding unit 150, an inverse quantization unit 160, an inverse transform unit 170, an adder 175, a filter unit 180, and a reference image buffer 190.

[0116] The encoding device 100 can encode the input image in intra-frame mode, inter-frame mode, or both. Furthermore, the encoding device 100 can generate a bitstream by encoding the input image and can output the generated bitstream. When intra-frame mode is used as the prediction mode, switch 115 can switch to intra-frame mode. When inter-frame mode is used as the prediction mode, switch 115 can switch to inter-frame mode. Here, intra-frame mode can be referred to as intra-frame prediction mode, and inter-frame mode can be referred to as inter-frame prediction mode. The encoding device 100 can generate prediction signals for input blocks of the input image. The prediction signal, as a block unit, can be referred to as a prediction block. Additionally, after generating the prediction block, the encoding device 100 can encode the residual value between the input block and the prediction block. The input image can be referred to as the current image, which is the target of the current encoding. The input block can be referred to as the current block or the encoding target block, which is the target of the current encoding.

[0117] When the prediction mode is intra-frame mode, the intra-frame prediction unit 120 can use the pixel values ​​of the previous coded blocks adjacent to the current block as reference pixels. The intra-frame prediction unit 120 can perform spatial prediction by using the reference pixels used for spatial prediction, and can generate prediction samples for the input block by using spatial prediction. Here, intra-frame prediction may refer to intra-frame prediction.

[0118] When the prediction mode is inter-frame mode, the motion prediction unit 111 can search for the region that best matches the input block of the reference image during motion prediction, and can derive the motion vector by using the searched region. The reference image can be stored in the reference image buffer 190.

[0119] The motion compensation unit 112 can generate prediction blocks by performing motion compensation using motion vectors. Here, the motion vectors can be two-dimensional vectors used in inter-frame prediction. Alternatively, the motion vectors can indicate the offset between the current image and the reference image. Here, inter-frame prediction may refer to inter-frame prediction.

[0120] When the value of the motion vector is not an integer, the motion prediction unit 111 and the motion compensation unit 112 can generate prediction blocks by applying interpolation filters to a portion of the reference image. To perform inter-frame prediction or motion compensation, the motion prediction method and motion compensation method of the prediction units included in the coding unit can be determined based on the coding unit in skip mode, merge mode, and AMVP mode. Furthermore, inter-frame prediction or motion compensation can be performed according to the mode.

[0121] Subtractor 125 can generate a residual block by using the residual value between the input block and the prediction block. The residual block can be referred to as the residual signal.

[0122] Transform unit 130 can generate transform coefficients by transforming the residual block and can output the transform coefficients. Here, the transform coefficients can be coefficient values ​​generated by transforming the residual block. In transform skip mode, transform unit 130 can skip the transformation of the residual block.

[0123] The quantized transformation coefficient level can be generated by applying quantization to the transformation coefficients. In the following, in this embodiment of the invention, the quantized transformation coefficient level can be referred to as the transformation coefficient.

[0124] The quantization unit 140 can generate a quantized transformation coefficient level by quantizing the transformation coefficients according to the quantization parameters, and can output the quantized transformation coefficient level. Here, the quantization unit 140 can quantize the transformation coefficients using a quantization matrix.

[0125] Based on the probability distribution, the entropy coding unit 150 can generate a bitstream by performing entropy coding on values ​​calculated by the quantization unit 140 or coding parameter values ​​calculated in the coding process, and can output the bitstream. The entropy coding unit 150 can perform entropy coding on information used for decoding the image and information about the image's pixels. For example, the information used for decoding the image may include syntax elements, etc.

[0126] When entropy coding is applied, the size of the bitstream of the encoded target symbol is reduced by allocating a small number of bits to symbols with high occurrence probabilities and a large number of bits to symbols with low occurrence probabilities. Therefore, entropy coding can increase the compression performance of image coding. For entropy coding, the entropy coding unit 150 can use coding methods such as exponential Golomb, context-adaptive variable-length coding (CAVLC), and context-adaptive binary arithmetic coding (CABAC). For example, the entropy coding unit 150 can perform entropy coding by using a variable-length code / table (VLC). Furthermore, the entropy coding unit 150 can derive a binarization method for the target symbol and a probabilistic model for the target symbol / bin, and can subsequently use the derived binarization method or the derived probabilistic model to perform arithmetic coding.

[0127] To encode the transform coefficient levels, the entropy coding unit 150 can transform the two-dimensional block-form coefficients into a one-dimensional vector form using a transform coefficient scanning method. For example, the two-dimensional block-form coefficients can be transformed into a one-dimensional vector form by scanning the block coefficients using an upper-right scan. Depending on the size of the transform unit and the intra-frame prediction mode, a vertical scan along the column direction and a horizontal scan along the row direction can be used instead of an upper-right scan. In other words, the scanning method can be determined among upper-right scan, vertical scan, and horizontal scan based on the size of the transform unit and the intra-frame prediction mode.

[0128] Encoding parameters can include not only information encoded by the encoder and transmitted to the decoder along with syntax elements, but also information that can be derived during encoding or decoding, or parameters necessary for encoding and decoding. For example, encoding parameters can include at least one of the following values ​​or statistics: intra-frame prediction mode, inter-frame prediction mode, intra-frame prediction direction, motion information, motion vector, reference image index, inter-frame prediction direction, inter-frame prediction indicator, reference image list, motion vector predictor, motion merging candidate, transform type, transform size, information on whether to use additional transforms, filter information within the loop, information on the presence of residual signals, quantization parameters, context model, transform coefficients, transform coefficient levels, encoded block pattern, encoded block flags, image display / output order, slice information, tile information, image type, information on whether to use motion merging mode, information on whether to use skip mode, block size, block depth, block partitioning information, cell size, cell partitioning information, etc.

[0129] The residual signal can refer to the difference between the original signal and the predicted signal. Alternatively, the residual signal can be a signal generated by transforming the difference between the original signal and the predicted signal. Alternatively, the residual signal can be a signal generated by transforming and quantizing the difference between the original signal and the predicted signal. A residual block can be a residual signal, which is a block unit.

[0130] When the encoding device 100 performs encoding using inter-frame prediction, the encoded current image can be used as a reference image for other images to be processed later. Therefore, the encoding device 100 can decode the encoded current image and store the decoded image as a reference image. To perform decoding, inverse quantization and inverse transform can be performed on the encoded current image.

[0131] The quantization coefficients can be dequantized by the inverse quantization unit 160 and inverse transformed by the inverse transform unit 170. The dequantized and inverse transformed coefficients can be added to the prediction block by the adder 175, thereby generating the reconstruction block.

[0132] The reconstructed block can be processed by filter unit 180. Filter unit 180 can apply at least one of a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the reconstructed block or the reconstructed image. Filter unit 180 may be referred to as an in-loop filter.

[0133] Deblocking filters remove block distortion that occurs at the boundaries between blocks. To determine whether a deblocking filter is being operated, it can be determined based on the pixels included in several rows or columns within the block. When a deblocking filter is applied to a block, a strong or weak filter can be applied depending on the desired deblocking intensity. Furthermore, when applying a deblocking filter, horizontal and vertical filtering can be processed in parallel while performing vertical and horizontal filtering.

[0134] Sample-adaptive offset adds the optimal offset value to the pixel value to compensate for coding errors. Sample-adaptive offset can use pixels to correct the offset between the deblocked filtered image and the original image. To perform offset correction on a specific image, one can use methods that consider the edge information of each pixel and apply the offset correction, or divide the image's pixels into a predetermined number of regions, determine the regions to be offset corrected, and apply the offset correction to the determined regions.

[0135] Adaptive loop filters can perform filtering based on values ​​obtained by comparing the reconstructed image and the original image. The pixels of the image can be partitioned into predetermined groups, a single filter can be determined for each group, and different filtering can be performed in each group. Information regarding whether to apply an adaptive loop filter can be sent to each coding unit (CU). The shape and filter coefficients of the adaptive loop filter applied to each block can vary. Alternatively, an adaptive loop filter with the same form (fixed form) can be applied regardless of the characteristics of the target block.

[0136] The reconstructed block that has passed through filter unit 180 can be stored in reference image buffer 190.

[0137] Figure 2 This is a block diagram illustrating the configuration of an image decoding apparatus to which an embodiment of the present invention is applied.

[0138] The decoding device 200 can be a video decoding device or an image decoding device.

[0139] refer to Figure 2 The decoding device 200 may include an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 230, an intra-frame prediction unit 240, a motion compensation unit 250, an adder 255, a filter unit 260, and a reference image buffer 270.

[0140] The decoding device 200 can receive the bitstream output from the encoding device 100. The decoding device 200 can decode the bitstream in intra-frame mode or inter-frame mode. Furthermore, the decoding device 200 can generate a reconstructed image through decoding and can output the reconstructed image.

[0141] When using intra-frame mode as the prediction mode used in decoding, the switch can be switched to intra-frame. When using inter-frame mode as the prediction mode used in decoding, the switch can be switched to inter-frame.

[0142] The decoding device 200 can obtain the reconstructed residual block from the input bit stream and can generate the prediction block.

[0143] When the reconstructed residual block and the prediction block are obtained, the decoding device 200 can generate a reconstructed block as the decoding target block by adding the reconstructed residual block and the prediction block. The decoding target block can be referred to as the current block.

[0144] The entropy decoding unit 210 can generate symbols by performing entropy decoding on the bitstream according to a probability distribution. The generated symbols may include symbols in the form of quantized transform coefficient levels.

[0145] Here, the entropy decoding method can be similar to the entropy encoding method described above. For example, the entropy decoding method can be the inverse process of the entropy encoding method described above.

[0146] To decode the transform coefficient levels, the entropy decoding unit 210 can transform one-dimensional block-form coefficients into two-dimensional vector forms using a transform coefficient scanning method. For example, by using an upper-right scan to scan the block coefficients, the one-dimensional block-form coefficients can be transformed into two-dimensional vector forms. Depending on the size of the transform unit and the intra-frame prediction mode, vertical and horizontal scans can be used instead of upper-right scans. In other words, the scanning method can be determined among upper-right scans, vertical scans, and horizontal scans based on the size of the transform unit and the intra-frame prediction mode.

[0147] The quantized transform coefficient levels can be dequantized by inverse quantization unit 220 and inverse transformed by inverse transform unit 230. The quantized transform coefficient levels are dequantized and inverse transformed to generate reconstructed residual blocks. Here, inverse quantization unit 220 can apply a quantization matrix to the quantized transform coefficient levels.

[0148] When using intra-frame mode, intra-frame prediction unit 240 can generate prediction blocks by performing spatial prediction using pixel values ​​of previous decoded blocks around the target decoded block.

[0149] When using inter-frame mode, motion compensation unit 250 can generate prediction blocks by performing motion compensation using both motion vectors and a reference image stored in reference image buffer 270. When the value of the motion vector is not an integer, motion compensation unit 250 can generate prediction blocks by applying an interpolation filter to a portion of the reference image. To perform motion compensation, the motion prediction method and motion prediction compensation method of the prediction units included in the coding unit can be determined based on the coding unit in skip mode, merge mode, and AMVP mode. Alternatively, inter-frame prediction or motion compensation can be performed according to the mode. Here, the current image reference mode can mean a prediction mode that uses a previously reconstructed region within the current image containing the target block. The previously reconstructed region may not be adjacent to the target block. To specify the previously reconstructed region, a fixed vector can be used for the current image reference mode. Additionally, a flag or index indicating whether the target block is a block decoded in the current image reference mode can be sent by signaling, and this flag or index can be derived using the reference image index of the target block. The current image used for the current image reference mode can exist at a fixed position (e.g., the position where refIdx = 0 or the last position) within the list of reference images used for the target block. Additionally, the image can be variably positioned within the list of reference images, and for this purpose, an additional reference image index can be sent to indicate the position of the current image.

[0150] Adder 255 can add the reconstructed residual blocks to the prediction block. The resulting block, generated by adding the reconstructed residual blocks and the prediction block, can be passed through filter unit 260. Filter unit 260 can apply at least one of a deblocking filter, a sampling adaptive offset, and an adaptive loop filter to the reconstructed block or the reconstructed image. Filter unit 260 can output the reconstructed image. The reconstructed image can be stored in reference image buffer 270 and can be used in inter-frame prediction.

[0151] Figure 3 It is a diagram that schematically represents the partitioning structure of an image during encoding and decoding. Figure 3 An example of partitioning a single cell into multiple cells at a lower level is illustrated.

[0152] To effectively partition an image, coding units (CUs) can be used during both encoding and decoding. A unit can refer to 1) a syntax element and 2) a block containing a sample image. For example, "partition of a unit" can refer to "partition of the block corresponding to that unit." Block partitioning information can include depth information of the unit. Depth information can indicate the number of partitions in the unit and / or the degree of partitioning.

[0153] refer to Figure 3Image 300 is partitioned in order of maximum coding unit (hereinafter referred to as LCU), and the partitioning structure is determined based on the LCU. Here, LCU can be used as coding tree unit (CTU). A single unit can include depth information based on the tree structure and can be partitioned hierarchically. Each of the lower-level partitioning units can include depth information. The depth information indicates the number and / or degree of partitioning in the unit, and therefore can include lower-level unit size information.

[0154] The partitioning structure refers to the distribution of coding units (CUs) within the LCU 310. A CU can be a unit used to efficiently encode an image. The distribution can be determined based on whether a single CU will be partitioned into multiple CUs (including positive integers greater than 2, such as 2, 4, 8, 16, etc.). The width and height of each partitioned CU can be half the width and half the height of a single CU. Alternatively, depending on the number of partitioned units, the width and height of each partitioned CU can be smaller than the width and height of a single CU. Similarly, a partitioned CU can be recursively partitioned into multiple CUs, each CU being half the width and height of the partitioned CUs.

[0155] Here, the partitioning of a CU can be performed recursively until a predetermined depth is reached. The depth information can be information indicating the size of the CU. The depth information for each CU can be stored therein. For example, the depth of an LCU can be 0, and the depth of a minimum coding unit (SCU) can be a predetermined maximum depth. Here, an LCU can be a CU with the maximum CU size as described above, and an SCU can be a CU with the minimum CU size.

[0156] Each time LCU 310 is partitioned and its width and height are reduced, the depth of the CU increases by 1. A CU that has not yet been partitioned can have a size of 2N×2N for each depth, and a CU that has been partitioned can be partitioned from a CU of size 2N×2N into multiple CUs, where each of the multiple CUs has a size of N×N. Each time the depth increases by 1, the size of N is halved.

[0157] refer to Figure 3 The size of an LCU with a minimum depth of 0 can be 64×64 pixels, and the size of an SCU with a maximum depth of 3 can be 8×8 pixels. Here, an LCU with 64×64 pixels can be represented by depth 0, a CU with 32×32 pixels can be represented by depth 1, a CU with 16×16 pixels can be represented by depth 2, and an SCU with 8×8 pixels can be represented by depth 3.

[0158] Furthermore, information about whether a specific CU will be partitioned can be represented by 1 bit of partition information for each CU. All CUs, except the SCU, may contain partition information. For example, when a CU is not partitioned, the partition information can be 0. Alternatively, when a CU is partitioned, the partition information can be 1.

[0159] Figure 4 This is a diagram showing the form of a prediction unit (PU) that can be included in a CU.

[0160] A CU that is no longer partitioned from the CU of the LCU partition can be partitioned into at least one PU. This process can also refer to partitioning.

[0161] A prediction unit (PU) can be the basic unit of prediction. A PU can be encoded and decoded in any of the skip mode, inter-frame prediction mode, and intra-frame prediction mode. A PU can be partitioned in various forms according to each mode.

[0162] like Figure 4 As shown, in skip mode, there may be no partitions within the CU. Alternatively, a 2N×2N mode 410 with the same size as the CU can be supported without partitions within the CU.

[0163] In the inter-frame prediction mode, eight partitioning formats can be supported within the CU, such as 2N×2N mode 410, 2N×2N mode 415, N×2N mode 420, N×N mode 425, 2N×nU mode 430, 2N×nD mode 435, nL×2N mode 440 and nR×2N mode 445.

[0164] Figure 5 This is a diagram illustrating the form of a transformation unit (TU) that may be included in a CU.

[0165] A transform unit (TU) can be a basic unit within a CU used for transform, quantization, inverse transform, and inverse quantization processes. A TU can be rectangular or square in form. The size and / or form of the CU can be determined independently.

[0166] A CU that is no longer partitioned from the LCU partition can be partitioned into one or more TUs. Here, the partitioning structure of a TU can be a quadtree structure. For example, ... Figure 5As shown, depending on the quadtree structure, a single CU 510 can be partitioned once or multiple times, such that the CU 510 is formed by TUs of various sizes. Alternatively, a single CU 510 can be partitioned into at least one TU based on the number of horizontal and / or vertical lines used to partition the CU. The CU can be partitioned into TUs that are symmetrical to each other, or it can be partitioned into TUs that are asymmetrical to each other. To partition into asymmetrical TUs, information about the size and shape of the TUs can be transmitted by signal, or this information can be derived from the information about the size and shape of the CUs.

[0167] While performing the transformation, the residual block can be transformed using one of the predetermined methods. For example, the predetermined methods may include Discrete Cosine Transform (DCT), Discrete Sine Transform (DST), or Karhunen-Loève Transform (KLT). To determine the method for transforming the residual block, the method can be determined by using at least one of the inter-frame prediction mode information of the prediction unit, the intra-frame prediction mode information of the prediction unit, or the size and form of the transform block. Alternatively, in some cases, information indicating the method can be transmitted via signaling.

[0168] Figure 6 This is a diagram illustrating an example of an intra-frame prediction mode.

[0169] The number of intra-prediction modes can vary depending on the size of the prediction unit (PU), or it can be fixed at N, independent of the PU size. Here, N can include 35 and 67, or it can be a positive integer greater than 1. For example, a predetermined intra-prediction mode for the encoder / decoder may include 2 non-angular modes and 65 angular modes, such as... Figure 6 As shown. The two non-angular modes can include DC mode and planar mode.

[0170] The number of intra-prediction modes can vary depending on the type of color component. For example, the number of intra-prediction modes can vary regardless of whether the color component is a luma signal or a chrominance signal.

[0171] The PU can be a square of size NxN or 2Nx2N. The NxN size may include 4x4, 8x8, 16x16, 32x32, 64x64, 128x128, etc. Alternatively, the PU can be of size MxN. Here, M and N can be positive integers greater than 2, and M and N can be different numbers. The unit of the PU can be the size of at least one of CU, PU, ​​and TU.

[0172] Intra-frame coding and / or decoding can be performed using sampled values ​​or coding parameters included in adjacent reconstruction units.

[0173] In intra-frame prediction, a prediction block can be generated by applying a reference sampling filter to a reference pixel using at least one of the sizes of the encoded / decoded target block. The type of reference filter applied to the reference pixel may differ. For example, the reference filter may vary depending on the intra-frame prediction mode of the encoded / decoded target block, the size / form of the encoded / decoded target block, or the location of the reference pixel. "The type of reference filter may differ" can refer to variations in the filter coefficients, the number of filter taps, the filter strength, or the number of filtering operations.

[0174] To perform intra-prediction, the intra-prediction mode of the current prediction unit can be predicted using the intra-prediction modes of its neighboring prediction units. When the intra-prediction mode of the current prediction unit is predicted using the intra-prediction mode information of its neighboring prediction units, and the two modes are identical, information indicating that both modes are identical can be sent using a predetermined flag. Alternatively, when the modes are different, all prediction mode information within the encoded / decoded target block can be encoded using entropy coding.

[0175] Figure 7 This is a diagram illustrating an example of inter-frame prediction processing.

[0176] Figure 7 A rectangle can refer to an image (or picture). Additionally, Figure 7 The arrows indicate the prediction direction. In other words, images can be encoded and / or decoded according to the arrow direction. Based on the encoding type, each image can be classified as an I-image (intra-frame image), P-image (one-way prediction image), and B-image (two-way prediction image), etc. Each image can be encoded and decoded according to its encoding type.

[0177] When the target image to be encoded is an I-image, intra-frame encoding of the target image itself can be performed simultaneously with inter-frame prediction. When the target image to be encoded is a P-image, it can be encoded using inter-frame prediction or motion compensation with reference images in the forward direction. When the target image to be encoded is a B-image, it can be encoded using inter-frame prediction or motion compensation with reference images in both the forward and reverse directions. Alternatively, it can be encoded using inter-frame prediction with reference images in both the forward and reverse directions. Here, in the case of inter-frame prediction mode, the encoder can perform inter-frame prediction or motion compensation, and the decoder can perform motion compensation in response to the encoder. Images of P-images and B-images encoded and / or decoded using reference images are used for inter-frame prediction.

[0178] The following describes in detail the inter-frame prediction according to the embodiments.

[0179] Inter-frame prediction or motion compensation can be performed using a reference image and motion information. Alternatively, inter-frame prediction can utilize the skip mode described above.

[0180] The reference image can be at least one of a previous image or a subsequent image of the current image. Here, in inter-frame prediction, blocks of the current image can be predicted based on the reference image. The region within the reference image can be specified using the reference image index refIdx, which indicates the reference image, and a motion vector, as described later.

[0181] In inter-frame prediction, a reference block corresponding to the current block within a reference image can be selected. The predicted block for the current block can be generated using the selected reference block. The current block can be the currently encoded or decoded target block within a block of the current image.

[0182] Motion information can be derived from the inter-frame prediction processing of the encoding device 100 and the decoding device 200. Furthermore, the derived motion information can be used for inter-frame prediction. Here, the encoding device 100 and the decoding device 200 can improve the efficiency of encoding and / or decoding by using motion information of reconstructed neighboring blocks and / or juxtaposed blocks (col blocks). A juxtaposed block can be a block in the reconstructed juxtaposed picture (col picture) that corresponds spatially to the encoding / decoding target block. Reconstructed neighboring blocks can be blocks within the current picture, as well as reconstructed blocks obtained through encoding and / or decoding. Additionally, a reconstructed block can be a block adjacent to the encoding / decoding target block, and / or a block located at the outer corner of the encoding / decoding target block. Here, a block located at the outer corner of the encoding / decoding target block can be a block adjacent in the vertical direction, and a block adjacent in the vertical direction is adjacent to the encoding / decoding target block in the horizontal direction. Alternatively, a block located at the outer corner of the encoding / decoding target block can be a block adjacent in the horizontal direction, and a block adjacent in the horizontal direction is adjacent to the encoding / decoding target block in the vertical direction.

[0183] Each of the encoding device 100 and the decoding device 200 can determine a predetermined relative position based on a block existing in the spatial location corresponding to the current block of the juxtaposed image. The predetermined relative position can be located inside and / or outside the block existing in the spatial location corresponding to the current block. Furthermore, the encoding device 100 and the decoding device 200 can derive the juxtaposed block based on the determined relative position. Here, the juxtaposed image can be at least one image from a list of reference images.

[0184] The method for deriving motion information can vary depending on the prediction mode of the encoded / decoded target block. For example, prediction modes applied to inter-frame prediction may include Advanced Motion Vector Predictor (AMVP) mode, merging mode, etc. Here, merging mode can refer to motion merging mode.

[0185] For example, in the case of applying the Advanced Motion Vector Predictor (AMVP) mode, the encoding device 100 and the decoding device 200 can generate a list of predicted motion vector candidates by using the recovered motion vectors of neighboring blocks and / or the motion vectors of juxtaposed blocks. In other words, the recovered motion vectors of neighboring blocks and / or the motion vectors of juxtaposed blocks can be used as predicted motion vector candidates. Here, the motion vectors of juxtaposed blocks can refer to temporal motion vector candidates, and the motion vectors of recovered neighboring blocks can refer to spatial motion vector candidates.

[0186] Encoding device 100 can generate a bitstream, and the bitstream can include motion vector candidate indices. In other words, encoding device 100 can entropy encode the motion vector candidate indices to generate the bitstream. The motion vector candidate index can indicate the optimal predicted motion vector selected from the predicted motion vector candidates included in the motion vector candidate list. The motion vector candidate index can be transmitted from encoding device 100 to decoding device 200 via the bitstream.

[0187] The decoding device 200 can perform entropy decoding on the motion vector candidate index through the bit stream, and select the motion vector candidate of the decoding target block from the motion vector candidates included in the motion vector candidate list by using the entropy-decoded motion vector candidate index.

[0188] Encoding device 100 can calculate the motion vector difference (MVD) between the motion vector of the target block and the motion vector candidates, and can entropy encode the motion vector difference (MVD). The bitstream can include the entropy-encoded MVD. The MVD is sent to decoding device 200 via the bitstream. Here, decoding device 200 can entropy decode the MVD from the bitstream. Decoding device 200 can derive the motion vector of the target block by summing the decoded MVD and the motion vector candidates.

[0189] The bitstream may include a reference image index indicating a reference image. The reference image index can be entropy encoded and transmitted from the encoding device 100 to the decoding device 200 via the bitstream. The decoding device 200 can predict the motion vector of the current block using motion information from neighboring blocks, and can derive the motion vector of the target block by using the predicted motion vector and its residual value. The decoding device 200 can generate a predicted block of the target block based on the derived motion vector and the reference image index information.

[0190] As another method for deriving motion information, a merging pattern can be used. A merging pattern can refer to merging motion across multiple blocks. It can also refer to applying motion information from a single block to another block. When a merging pattern is applied, the encoding device 100 and the decoding device 200 can generate a merging candidate list by using the recovered motion information from adjacent blocks and / or the motion information from juxtaposed blocks. Here, the motion information can include at least one of 1) motion vectors, 2) reference image indices, and 3) inter-frame prediction indicators. The prediction indicator can indicate unidirectional (LO prediction, L1 prediction) or bidirectional prediction.

[0191] Here, the merging pattern can be applied in the units of the coding unit or the prediction unit (PU). When the merging pattern is executed by the CU unit or the PU unit, the encoding device 100 can generate a bitstream by entropy encoding predetermined information and send the bitstream to the decoding device 200. The bitstream may include predetermined information. The predetermined information may include 1) a merging flag indicating whether the merging pattern is used for each block partition, and 2) a merging index including information on which block among the adjacent blocks adjacent to the target block is merged. For example, the adjacent blocks adjacent to the target block may include the left adjacent block of the current block, the upper adjacent block of the target block, the temporally adjacent block of the target block, etc.

[0192] The merge candidate list can represent a list containing motion information. The merge candidate list can be generated before executing the merge mode. The motion information stored in the merge candidate list can be at least one of the following: motion information of adjacent blocks adjacent to the encoding / decoding target block, motion information of juxtaposed blocks corresponding to the encoding / decoding target block in the reference image, newly generated motion information by combining motion information pre-existing in the merge motion candidate list, and zero merge candidates. Here, the motion information of adjacent blocks adjacent to the encoding / decoding target block can refer to spatial merge candidates, and the motion information of juxtaposed blocks corresponding to the encoding / decoding target block in the reference image can refer to temporal merge candidates.

[0193] In the skip mode, motion information from adjacent blocks is applied to the target block for encoding / decoding. The skip mode can be one of other modes used for inter-frame prediction. When using the skip mode, the encoding device 100 can generate a bitstream by entropy encoding information from adjacent blocks that can be used to encode the target block, and send this bitstream to the decoding device 200. The encoding device 100 may not send other information, such as syntax information, to the decoding device 200. The syntax information may include at least one of residual information of motion vectors, a coding block flag, and a transform coefficient level.

[0194] Figures 8 to 18 This is a diagram illustrating a method for generating a geometrically modified image by geometrically modifying an image.

[0195] Geometric modification of an image can refer to geometrically altering the image's light information. Light information can refer to the brightness, color, or chromaticity of each point in the image. Alternatively, light information can refer to the pixel values ​​in a digital image. Geometric modification can refer to parallel movement of each point within the image, image rotation, image resizing, etc.

[0196] Figures 8 to 12 These are diagrams illustrating the geometric modifications to the images according to the present invention. In each diagram, (x, y) refers to a point in the original image before modification. (x', y') refers to the point corresponding to (x, y) after modification. Here, the corresponding point refers to the point where the light information of (x, y) is shifted through geometric modification.

[0197] Figure 8 This is a diagram illustrating a transfer modification of an embodiment of geometric modification of an image according to the present invention.

[0198] exist Figure 8 In this context, tx represents the displacement of each point that has been transferred along the x-axis, and ty represents the displacement of each point that has been transferred along the y-axis. Therefore, the point (x', y') within the image is derived by adding tx and ty to the point (x, y), which is the point within the image before the modification. The transformation modification can be... Figure 8 The matrix shown is used to represent this.

[0199] Figure 9 This is a diagram illustrating the size modification of an embodiment of geometric modification of an image according to the present invention.

[0200] exist Figure 9 In this context, sx represents the size modification factor along the x-axis, while sy represents the size modification factor along the y-axis. The size modification factor represents the ratio of the original image size to the modified image size. When the size modification factor equals 1, it means the original image size is equal to the modified image size. When the size modification factor is greater than 1, it means the modified image size has been enlarged. When the size modification factor is less than 1, it means the modified image size has been reduced. The size modification factor always has a value greater than 0. Therefore, by multiplying sx and sy by a point (x, y) in the original image, we can derive a point (x', y') in the modified image after the size change. Size modification can be used... Figure 9 The matrix representation is shown below.

[0201] Figure 10 This is a diagram illustrating a rotational modification of an embodiment of geometric modification of an image according to the present invention.

[0202] exist Figure 10 In this context, θ refers to the rotation angle of the image. Figure 10In this embodiment, rotation is performed centered on point (0,0) of the original image. The point (x', y') within the rotated image can be derived using θ and trigonometric functions. Rotation modification can be... Figure 10 The matrix representation shown is shown.

[0203] Figure 11 This is a diagram illustrating an embodiment of geometric modification of an image according to the present invention, showing an affine modification.

[0204] Affine modification refers to the situation where transfer, size, and rotation modifications are performed in combination. The geometric modifications of an affine modification can vary depending on the order in which the transfer, size, and / or rotation modifications are applied to the image. Depending on the order in which the multiple modifications that make up the affine modification and the complex of each modification are applied, the image can be modified in the form of tilting, as well as transfer, size, and rotation modifications.

[0205] exist Figure 11 In the middle, M i It can be a 3x3 matrix used for transformation modifications, size modifications, or rotation modifications. Depending on the order of the modifications that make up the affine modification, a 3x3 matrix can be obtained by multiplying each matrix used for modification with the other. Figure 11 In this context, matrix A can be represented by the path from matrix M1 to matrix M. n The affine modification is a 3x3 matrix obtained by multiplying the matrices. Matrix A can consist of elements a1 to a6. Matrix p represents the points in the image before modification, where the modification is represented by a matrix. Matrix p' represents the points in the image after modification and corresponds to the point p in the image before modification. Therefore, the affine modification can be expressed as the matrix equation p' = Ap.

[0206] Figure 12 This is a diagram illustrating a projection modification of an embodiment of geometric modification of an image according to the present invention.

[0207] Projective modifications can be extended affine modifications, where perspective modifications are added to the affine modifications. When an object in 3D space is projected onto a 2D plane, perspective modifications may occur depending on the camera's or observer's viewing angle. In perspective modifications, distant objects are represented as small, while closer objects are represented as large.

[0208] exist Figure 12 In this context, matrix H can be used for projection modification. The elements h1 to h6 that constitute matrix H can correspond to the elements that constitute... Figure 11 The elements a1 to a6 of matrix A are affine modifications. Therefore, projection modifications can include affine modifications. The elements h7 and h8 constituting matrix H can be elements related to perspective modifications.

[0209] Geometric modification of an image is a method of modifying the geometry of an image to a specific form. Points in the geometrically modified image can be computed to correspond to points in the original image using geometric modifications defined in a matrix. Conversely, homography refers to a method of deriving a common geometric modification matrix from two images that have corresponding points in each other.

[0210] Figure 13 This is a diagram illustrating an example of a method for implementing homography according to the present invention.

[0211] Homography can be used to derive the geometric modifications between two images by identifying two corresponding points within each other. This can be achieved using feature point matching. Feature points in an image refer to points within that image that possess descriptive features.

[0212] In steps S1301 and S1302, the homography implementation method can extract feature points from the original image and the geometrically modified image. Feature points can be extracted differently depending on the extraction method or the intended use. Points where brightness values ​​change significantly within the image, the center points of regions with specific shapes, or the outer corners of objects within the image can be used as feature points. Feature points can be extracted using algorithms such as Scale Invariant Feature Transform (SIFT), Accelerated Robust Feature Transform (SURF), and blob detection.

[0213] In step S1303, the homography implementation method can match feature points based on feature points extracted from the original image and the geometrically modified image. Specifically, each extracted feature point is descriptive, and feature points between two images can be matched by finding points with similar descriptive information. The matched feature points can be used as points corresponding to each other in the original image and the geometrically modified image.

[0214] However, feature point matching may not match points that actually correspond to each other. Therefore, in step S1304, valid feature points can be selected from the derived feature points. The method for selecting valid feature points can vary depending on the calculation algorithm. For example, methods such as excluding feature points that do not meet the baseline based on descriptive information, excluding feature points with very low consistency based on the distribution of matching results, or using the Random Sample Consensus (RANSAC) algorithm can be used. The homography implementation method can selectively execute step S1304 based on the feature point matching results. In other words, step S1304 may be omitted depending on the situation. Alternatively, steps S1303 and S1304 can be combined. Or, the homography implementation method can perform the valid feature point matching process without executing steps S1303 and S1304.

[0215] In step S1305, the homography implementation method can derive the relationship between the original image and the geometrically modified image by using the selected valid points. In step S1306, the homography implementation method can derive the geometric matrix by using the derived formula. Alternatively, the homography implementation method may omit step S1306 and output information about the derived formula obtained in step S1305, other than the geometrically modified matrix, in a different form.

[0216] Figure 14 This is an exemplary method according to the present invention for deriving the relationship between two corresponding points in two images.

[0217] Geometric modifications to the image can be performed using a 3×3 matrix H. Therefore, a simultaneous equation consisting of elements h1 to h9 of matrix H as unknowns can be derived from the matrix formula p' = Hp. Here, p represents a point within the original image, and p' represents the point within the geometrically modified image corresponding to point p. The equations can be easily computed by fixing H9 to 1, by dividing all elements of matrix H by h9. Furthermore, the number of unknowns can be reduced from 9 to 8.

[0218] Figure 14 The elements k1 to k8 correspond to the values ​​of h1 to h8 divided by h9. Changing h9 to 1, and changing h1 to h8 to k1 to k8 respectively, performs the same geometric modification. Therefore, it may be necessary to compute 8 unknowns. Figure 14 In this context, the final formula for a pair of points that match each other can be expressed in two forms: x' and y'. At least four pairs of matching points may be required because there are eight unknown values. However, as mentioned above, a pair of points may not match each other. Or, a pair of points may mismatch. This error can occur even if valid feature points are selected. This error can be reduced by using many pairs of matching points while calculating the geometric modification matrix. Therefore, considering these features, the number of pairs of points to be used can be determined.

[0219] Figure 15 This is a diagram illustrating a method for generating a geometrically modified image based on a geometric modification matrix and an original image according to the present invention.

[0220] like Figure 15 As shown, by using the light information of points within the original image, the generation of a geometrically modified image can correspond to the generation of the light information of the corresponding points within the geometrically modified image. Figure 15In this context, (x0, y0), (x1, y1), and (x2, y2) refer to different points within the original image. Additionally, (x'0, y'0), (x'1, y'1), and (x'2, y'2) are points within the geometrically modified image corresponding to (x0, y0), (x1, y1), and (x2, y2), respectively. Function f calculates the corresponding x' coordinates on the x-axis within the geometrically modified image using the points (x, y) in the original image and the additional information α for geometric modification. Function g calculates the corresponding y' coordinates on the y-axis within the geometrically modified image using the points (x, y) in the original image and the additional information β for geometric modification. When (x, y), (x', y'), function f, and function g are expressed as matrix formulas, matrix H can refer to the geometric modification method. Therefore, points corresponding to each other in the original and geometrically modified images can be found using matrix H.

[0221] Figure 15 The geometric modification method can be problematic in discrete-sampled image signals because optical information is only included in points with integer coordinates in the discrete image signal. Therefore, when a point within the geometrically modified image and corresponding to a point in the original image has real coordinates, the optical information with the nearest integer coordinates is assigned to the point within the geometrically modified image. Consequently, optical information may be superimposed onto a portion of the points with real coordinates within the geometrically modified image, or optical information may not be assigned at all. In such cases, inverse mapping can be used.

[0222] Figure 16 This is a diagram illustrating a method for generating a geometrically modified image using inverse mapping according to the present invention.

[0223] Figure 16 The dashed rectangular region represents the actually observed region. Points within the original image corresponding to each point within the dashed rectangular region can be derived. Therefore, the light information of the original image can be assigned to all points within the geometrically modified image. However, the point (x3, y3) corresponding to (x'3, y'3) may be located outside the original image. In this case, the light information of the original image may not be assigned to the point (x'3, y'3). Among the points for which no light information from the original image has been assigned, the light information of the adjacent points in the original image can be assigned. In other words, the light information of the nearest point within the original image (e.g., (x4, y4)) can be assigned.

[0224] Figure 17 The figure illustrates a method for generating a geometrically modified image based on a geometric modification matrix and an original image according to the present invention, wherein the geometric modification matrix may correspond to geometric modification information.

[0225] In step S1701, the generation method may receive an input original image, a geometric modification matrix, and / or information about the current point of the geometric modification image. The generation method can calculate a point in the original image corresponding to the current point of the geometric modification image using the original image and the geometric modification matrix. The calculated corresponding point in the original image can be a real-valued point with real coordinates.

[0226] In step S1702, the generation method can determine whether the calculated corresponding point is located inside the original image.

[0227] In step S1702, when the calculated corresponding point is not located inside the original image, in step S1703, the generation method can use the corresponding point to change the point closest to the calculated corresponding point in the original image.

[0228] In step S1702, when the calculated corresponding point is located inside the original image, the generation method can execute step S1704. When the calculated corresponding point is changed in step S1703, the generation method can execute step S1704.

[0229] In step S1704, when the corresponding point has real coordinates, the generation method can identify the closest point with integer coordinates. When the corresponding point has integer coordinates, the generation method can skip steps S1704 and S1705 and execute step S1706.

[0230] In step S1705, the generation method can generate light information of points with real coordinates by inserting light information (e.g., pixel values) of recognition points with integer coordinates. As an insertion method, Lanczos interpolation, S-spline interpolation, or bicubic interpolation can be used.

[0231] In step S1706, the generation method can check whether all points within the geometrically modified image have completed their geometric modifications. Then, the generation method can finally output the generated geometrically modified image.

[0232] If it is determined in step S1706 that the geometric modification is not completed, in step S1707, the generation method can change the current point of the geometric modification image to another point, and can repeat steps S1701 to S1706.

[0233] Figure 18 Reference numerals illustrating embodiments of the present invention are shown. Figure 17 A graph illustrating bilinear interpolation, one of the various interpolation methods explained.

[0234] exist Figure 18 In this context, real coordinates (x, y) can correspond to... Figure 17The real number corresponding points mentioned in step S1704. The four points (i,j), (i,j+1), (i+1,j), and (i+1,j+1) adjacent to the coordinates (x,y) can correspond to Figure 17 The nearest point with integer coordinates mentioned in step S1704. I(x, y) can refer to the light information of point (x, y), such as brightness. a refers to the x-axis distance between i and x, and b refers to the y-axis distance between j and y. 1-a refers to the x-axis distance between i+1 and x, and 1-b refers to the y-axis distance between j+1 and y. The light information of point (x, y) can be calculated from the light information of points (i, j), (i, j+1), (i+1, j), and (i+1, j+1) by using the ratio of a to 1-a in the x-axis and the ratio of b to 1-b in the y-axis.

[0235] When the encoder's inter-frame prediction unit performs motion prediction, it can predict the encoded target region (current region or current block) within the encoded target image (current image) by referring to a reference image. Here, when the time interval between the reference image and the encoded target image is large, or when rotation, zooming, scaling, or global motion such as a change in the target's viewpoint has occurred between the two images, the pixel consistency between the two images decreases. Therefore, prediction accuracy may decrease and coding efficiency may decrease. In this case, the encoder can calculate the change in motion between the encoded target image and the reference image and geometrically modify the reference image so that it has a similar form to the encoded target image. See [reference needed]. Figures 8 to 18 A reference image can be geometrically modified. The reference image can be geometrically modified at the frame, slice, and / or block level. An image generated by geometrically modifying a reference image can be defined as a geometrically modified image. Motion prediction accuracy can be improved by referencing a geometrically modified image instead of a reference image. The entirety or a portion of a reference image can be defined as a reference image, and a geometrically modified reference image can be defined as a geometrically modified image.

[0236] Motion information can be generated by performing inter-frame prediction for the current block. The generated motion information can be encoded and included in the bitstream. Motion information may include motion vectors, the number of reference images, reference orientation, etc. Motion information can be encoded in various forms of units that configure the images. For example, motion information can be encoded as prediction units (PUs). In this invention, motion information may refer to inter-frame prediction information.

[0237] According to embodiments of the present invention, the target image can be predicted not only using a reference image but also using a geometrically modified image. When using a geometrically modified image, geometric modification information can be additionally encoded / decoded. Geometric modification information can refer to all kinds of information used to generate the geometrically modified image based on a reference image or a portion thereof. For example, geometric modification information may include global motion vectors, geometric transfer modification matrices, geometric size modification matrices, geometric affine modification matrices, and / or geometric projection modification matrices.

[0238] The encoder can generate geometric modification information based on the relationship between the target image and a reference image. This geometric modification information can be used to generate a geometrically modified image, making the reference image geometrically similar to the target image. The encoder can then perform optimized encoding using both the reference image and the geometrically modified image.

[0239] Figure 19 This is a diagram illustrating motion prediction using reference images and / or geometrically modified images according to embodiments of the present invention.

[0240] Figure 19 (a) in the figure is a motion prediction diagram of the encoder according to an embodiment of the present invention. Figure 19 (b) in the figure is a graph of motion prediction by the decoder according to an embodiment of the present invention.

[0241] First, refer to Figure 19 (a) in the example describes the motion prediction of the encoder.

[0242] In steps S1901 and S1902, the encoder can specify the target image and reference images. The encoder can select a reference image from the list of reference images. The list of reference images can be configured to contain reference images stored in a reference image buffer (reconstruction image buffer).

[0243] In step S1903, the encoder can generate geometric modification information based on the encoded target image and the reference image. The geometric modification information can be generated using the reference image. Figures 8 to 18 The method described is used to generate it.

[0244] In step S1904, the encoder can generate a geometrically modified image from the reference image based on the generated geometric modification information. The encoder can then store the generated geometrically modified image in a geometrically modified image buffer.

[0245] In step S1905, the encoder can perform motion prediction by referring to a reference image and / or a geometrically modified image.

[0246] In step S1906, the encoder can store or update the best prediction information with optimal coding efficiency based on the reference signal. Rate distortion cost (RD Cost) can be used as an indicator to determine the optimal coding efficiency.

[0247] In step S1907, the encoder can determine whether the above processing has been applied to all reference images in the reference image list. If not, the above steps can be repeated after moving to step S1902. If so, the encoder can encode the finally determined optimal (or best) motion prediction information and / or geometric modification information in step S1908. The encoder can only encode and send geometric modification information when the geometric modification image is used for motion prediction.

[0248] As Figure 19 As a result of motion prediction in (a), the encoder can generate motion vectors and information about the reference image as inter-frame prediction information. The information about the reference image may include information for identifying the reference image used for inter-frame prediction (e.g., reference image index), and / or information indicating whether a geometrically modified image is referenced (e.g., geometric modification usage information). This information can be sent in various units. For example, it can be sent in units of images, slices, coding units (CUs), or prediction units (PUs).

[0249] Figure 19 Each step in (a) can be applied to a portion of the target image and a portion of the reference image. Here, the target image can correspond to the target image, the reference image can correspond to the reference image, and the geometrically modified image can correspond to the geometrically modified image.

[0250] Next, refer to Figure 19 (b) in the example describes the motion prediction of the decoder.

[0251] In step S1911, the decoder can receive and decode the inter-frame prediction information of the current block.

[0252] In step S1912, the decoder can select a reference image based on inter-frame prediction information. For example, the decoder can use a reference image index to select a reference image.

[0253] In step S1913, the decoder can determine whether to use geometrically modified images based on inter-frame prediction information. The decoder can use the geometrically modified information used in step S1913.

[0254] When it is determined in step S1913 that the geometrically modified image is not used, in step S1915, the decoder can perform inter-frame prediction by referring to the reference image selected in step S1912.

[0255] When it is determined in step S1913 that a geometrically modified image is to be used, in step S1914, the decoder can generate a geometrically modified image by geometrically modifying the reference image selected in step S1912. The decoder can generate a geometrically modified image by geometrically modifying a portion of the reference image. In this case, a portion of the reference image can correspond to the reference image, and the geometrically modified image can correspond to the geometrically modified image.

[0256] In step S1915, the decoder can perform inter-frame prediction based on the geometrically modified image generated in step S1914.

[0257] Figure 20 This is a block diagram illustrating the configuration of an image encoding apparatus to which another embodiment of the present invention is applied.

[0258] Figure 20 The encoding apparatus shown may include a geometrically modified image generation unit 2010, a geometrically modified image prediction unit 2015, an extended intra-frame prediction unit 2020, a subtractor 2025, a transform unit 2030, a quantization unit 2040, an entropy coding unit 2050, an inverse quantization unit 2060, an inverse transform unit 2070, an adder 2075, a deblocking filter unit 2080, and a sampling adaptive offset unit 2090.

[0259] The geometric modification image generation unit 2010 can generate a geometric modification image 2012 by calculating geometric modification information. The calculated geometric modification information reflects the changes in pixel values ​​between the encoded target image 2011 and the reference images in the reference image list constructed from the reconstructed image buffer 2013. The generated geometric modification image 2012 can be stored in the geometric modification image buffer 2016.

[0260] The geometric modification image prediction unit 2015 may include a geometric modification image buffer 2016 and an inter-frame predictor 2017. The geometric modification image buffer 2016 may store geometric modification images generated in the geometric modification image generation unit 2010. The inter-frame predictor 2017 can perform motion prediction using both the geometric modification images stored in the geometric modification image buffer and reference images from a reference image list constructed from the reconstructed image buffer 2013. When referencing the geometric modification images while performing motion prediction, the geometric modification information used to generate the geometric modification images can be sent to the entropy coding unit 2050 and encoded therein.

[0261] The Extended Intra Prediction Unit 2020 can perform extended intra prediction by referencing the geometrically modified image and the encoded / decoded signal of the current image.

[0262] refer to Figure 20 The configuration of the encoding device described is only one of the various embodiments of the present invention and is not limited thereto. Figure 20 Some configurations of the encoding apparatus shown can be combined with other configurations or omitted. Alternatively, other configurations can be added. Furthermore, a portion of the multiple configurations included in the geometry-modified image generation unit 2010 and the geometry-modified image prediction unit 2015 can be configured independently of the geometry-modified image generation unit 2010 and the geometry-modified image prediction unit 2015. Alternatively, it can be included in a sub-configuration of another configuration, or combined with another configuration.

[0263] Figure 21 This is a block diagram illustrating the configuration of an image decoding apparatus to which another embodiment of the present invention is applied.

[0264] Figure 21 The decoding device shown may include an entropy decoding unit 2110, an inverse quantization unit 2120, an inverse transform unit 2130, a subtractor 2135, a filter unit 2140, an extended intra-frame prediction unit 2150, a geometrically modified image predictor 2160, and a geometrically modified image generator 2170. The decoding device can output a decoded image 2180 by receiving a bitstream 2100.

[0265] The geometric modification image generator 2170 can generate a geometric modification image 2172 by using geometric modification information extracted from bitstream 2100 and entropy decoded, and reference images from a list of reference images constructed from the reconstructed image buffer 2171.

[0266] The geometrically modified image prediction unit 2160 may be configured with a geometrically modified image buffer 2161 for storing geometrically modified images and an inter-frame predictor 2162.

[0267] The geometrically modified image 2172 generated in the geometrically modified image generator 2170 can be stored in the geometrically modified image buffer 2161. The geometrically modified image 2172 stored in the geometrically modified image buffer 2161 can be used as a reference signal in the inter-frame predictor 2162.

[0268] The inter-frame predictor 2162 can reconstruct the decoded target image based on information sent from the encoding device by using a reference image and / or a geometrically modified image as a reference signal for motion prediction.

[0269] refer to Figure 21 The configuration of the decoding device described is merely one of the various embodiments of the present invention, and is not limited thereto. Figure 21Some configurations of the decoding device shown can be combined with other configurations or omitted. Alternatively, other configurations can be added. Furthermore, a portion of the multiple configurations included in the geometry-modified image generator 2170 and the geometry-modified image prediction unit 2160 can be configured independently of the geometry-modified image generator 2170 and the geometry-modified image prediction unit 2160. Alternatively, it can be included in a sub-configuration of another configuration, or combined with another configuration.

[0270] According to the present invention, geometric modification usage information may not be encoded and can be predicted. Geometric modification usage information may be information used to indicate whether a reference used in inter-frame prediction is a geometrically modified image generated through geometric modification.

[0271] Video encoding / decoding apparatuses using geometrically modified images need information to identify whether a reference image or a geometrically modified reference image is used for inter-frame prediction of the current block. Additionally, when one or more reference images are used for inter-frame prediction, all reference images can be geometrically modified reference images. Alternatively, a portion of the reference images can be geometrically modified reference images. Alternatively, none of the reference images can be geometrically modified reference images. Geometric modification usage information can be sent for each block unit, indicating whether the reference image is used as is or by geometrically modifying it during inter-frame prediction. Therefore, a large amount of data can be used to encode the geometric modification usage information. The block unit used to send the geometric modification usage information can include all types of units used for encoding images. For example, the unit can be a macroblock, coding unit (CU), prediction unit (PU), etc.

[0272] While encoding / decoding the current block, this invention can predict the geometric modification usage information of the current block based on geometric modification usage information derived from blocks that are temporally or spatially adjacent to the current block. Therefore, additional data transmission is omitted by predicting the geometric modification usage information of the current block.

[0273] Figure 22 This is a diagram illustrating the operation of an encoder including a geometric modification information prediction unit according to an embodiment of the present invention.

[0274] according to Figure 22 The encoder of the embodiment shown may include an image relationship recognition unit 2210, an image geometry modification unit 2220, an inter-frame prediction unit 2230, and / or a reconstructed image buffer 2250.

[0275] The image relationship recognition unit 2210 can recognize the relationship between the encoded target image 2260 and the reference image stored in the reconstructed image buffer 2250, and generate image relationship information. The image relationship information can refer to information that can be used to modify the reference image to be similar to the encoded target image. In this invention, the image relationship information can refer to geometric modification information. For example, the image relationship information may include a geometric modification matrix, and / or global motion information, etc.

[0276] The image geometry modification unit 2220 can generate a geometry-modified reference image from the reference image based on the inter-image relationship information generated in the inter-image relationship recognition unit 2210. Figure 15 This is an example of generating a geometrically modified reference image. In Figure 15 In the illustrated embodiment, matrix H can correspond to inter-image relationship information. (Reference image) Figure 15 Each pixel of the original image can be compared with a geometrically modified reference image. Figure 15 Position matching of geometrically modified images.

[0277] The inter-frame prediction unit 2230 can determine the reference image with optimal coding efficiency by using each reference image and / or each corresponding geometrically modified reference image stored in the reconstructed image buffer 2250, and generate reference image information 2280. Reference image information 2280 may include all kinds of information specifying the reference image used for inter-frame prediction by the encoder. For example, reference image information 2280 may include information specifying the reference image within the reference image buffer (i.e., the reference image index), motion vectors indicating the reference block within the reference image referenced by the encoded target block, etc. The inter-frame prediction unit 2230 may use pixel information (pixel values) of the reference image and / or its corresponding geometrically modified reference image, as well as mode information 2270. Mode information may refer to all kinds of information used for encoding / decoding each reference image and / or its corresponding geometrically modified reference image. Mode information may include the partitioning structure of the reference image and / or prediction information (such as motion information), etc.

[0278] Inter-frame prediction unit 2230 can perform inter-frame prediction for encoding target image 2260 by using each reference image and / or the corresponding geometrically modified reference image. Inter-frame prediction unit 2230 may not use all reference images, but rather a subset of reference images that produce good results.

[0279] To decode an image, the decoder may need information indicating to the encoder which reference image to use among the reference images and / or corresponding geometrically modified reference images. For example, the decoder may need optimal (best) prediction reference image information (e.g., the number or index of optimal prediction reference images) 2280 and / or information indicating whether the geometrically modified reference image is used for inter-frame prediction geometric modification usage 2290.

[0280] The geometric modification usage information predictor 2240 can predict the geometric modification usage information 2290 and encode the geometric modification usage information 2290 based on the prediction result. In this invention, the geometric modification usage information predictor 2240 can be included in the inter-frame prediction unit 2230 and predict the geometric modification usage information 2290.

[0281] The operations performed in the encoder can be performed identically in the decoder to accurately reconstruct the information. Therefore, when the geometric modification usage information used in the encoder is predicted identically in the decoder, the geometric modification usage information can be predicted without sending additional data.

[0282] Figure 23 This is a diagram illustrating the operation of an encoder including a geometric modification information prediction unit (predictor) according to an embodiment of the present invention.

[0283] according to Figure 23 The encoder of the illustrated embodiment may include a reconstructed image buffer 2310, an inter-image relationship recognition unit 2320, an image geometric modification unit 2330, a geometrically modified image buffer 2340, an inter-frame prediction unit 2350, and / or a selective data transmission unit 2370.

[0284] Image relationship recognition unit 2320 can correspond to Figure 22 Image relationship recognition unit 2210. Image relationship recognition unit 2320 can recognize the relationship between the encoded target image 2300 and each reference image, and output the recognized information as image relationship information. Image relationship information can be used by image geometry modification unit 2330.

[0285] Image geometry modification unit 2330 can correspond to Figure 22Image geometry modification unit 2220. Image geometry modification unit 2330 can generate geometry-modified reference images (geometric modification images) from each reference image by using the inter-image relationship information identified by inter-image relationship recognition unit 2320. Geometry-modified reference image buffer (geometric modification image buffer) 2340 can manage geometry-modified reference images. Geometry-modified reference image buffer 2340 can be managed using reconstructed image buffer 2310. The geometry-modified reference images generated by image geometry modification unit 2330 can be used for inter-frame prediction of encoded target images.

[0286] Inter-frame prediction unit 2350 can correspond to Figure 22 Inter-frame prediction unit 2230. Inter-frame prediction unit 2350 can perform inter-frame prediction by using reference images and / or geometrically modified reference images.

[0287] Inter-frame prediction unit 2350 may include a geometric modification usage information predictor 2360. Geometric modification usage information 2382 may indicate whether the reference picture used for inter-frame prediction is a geometrically modified picture generated through geometric modification. Geometric modification usage information predictor 2360 may predict information regarding whether a geometrically modified reference picture is used for a coded target block. Information regarding the use of geometrically modified reference pictures may refer to various types of information identifying the geometrically modified reference picture among all reference pictures used for inter-frame prediction.

[0288] The geometric modification information predictor 2360 may include a neighboring block identification unit 2361, a prediction candidate list generator 2362, and / or a prediction candidate selector 2363.

[0289] Neighboring blocks adjacent to the coded target block (current block) can be used to predict geometric modifications using information 2832. Neighboring blocks can include blocks that are temporally and spatially adjacent to the current block. Temporally adjacent blocks can include... Figure 29 The juxtaposed blocks (M and / or H). Spatially adjacent blocks may include Figure 29 Blocks A0 to B2. Alternatively, except... Figure 29 In addition to blocks A0 to B2, spatially adjacent blocks may include blocks adjacent to the current block. Alternatively, temporally adjacent blocks may include blocks directly adjacent to the current block and blocks separated from the current block by a predetermined distance.

[0290] The predicted similarity between neighboring blocks and the current block can be used to predict geometric modification usage information 2832. Therefore, in order to predict geometric modification usage information 2832, the neighboring block identification unit 2261 can identify at least one neighboring block that is temporally and / or spatially adjacent to the current block. The neighboring block identification unit 2261 can identify neighboring blocks by checking whether a neighboring block exists, whether a neighboring block is available, and / or whether a neighboring block is an inter-frame prediction block, etc. The availability of a neighboring block can refer to whether information about the neighboring block can be referenced. For example, if a neighboring block and the current block belong to different tiles and / or slices, or if the neighboring block is not referenced through parallel processing or for other reasons, the neighboring block can be determined to be unavailable.

[0291] The prediction candidate list generator 2362 can generate a prediction candidate list based on the availability of neighboring blocks identified by the neighboring block identifier 2361 and / or the priority of neighboring blocks for prediction.

[0292] The prediction candidate selector 2363 can select one prediction candidate from the prediction candidate list. The prediction candidate selector 2363 can choose the prediction candidate with the best prediction result from the prediction candidates included in the prediction candidate list. When the prediction candidate with the best prediction is selected, the inter-frame prediction information of the current block can be predicted using the inter-frame prediction information of the selected prediction candidate. The inter-frame prediction information may include reference image recognition information, geometric modification usage information, and / or motion vectors. The geometric modification usage information predictor 2360 can be used in merge mode and other inter-frame prediction modes. Additionally, multiple geometric modification usage information predictors 2360 can be provided based on the number of inter-frame prediction modes.

[0293] When predictions of adjacent blocks that are temporally and / or spatially adjacent to the current block are not available, or when coding efficiency is poor, the best prediction reference image search unit 2355 can be used. The best prediction reference image search unit 2355 may not predict the geometric modification usage information 2382 and directly search for reference images that can derive the best prediction result for the current block. The best prediction reference image search unit 2355 can generate best prediction reference image information 2381 and geometric modification usage information 2382 based on the searched reference images.

[0294] The best prediction reference image information 2381 can refer to information about reference images referenced during inter-frame prediction, and can include the reference image number (index). When performing inter-frame prediction, the best prediction reference image can be the reference image with the best coding efficiency among the reference image candidates.

[0295] The selective data transmission unit 2370 can select and encode information to be sent to the decoder from the information generated by the inter-frame prediction unit based on the prediction method used in the inter-frame prediction unit, and transmit the encoded information via a bitstream. The prediction method information 2380 can indicate the prediction method used in the inter-frame prediction unit and can be encoded and transmitted via a bitstream. The prediction method may include prediction methods for performing inter-frame prediction, including merge mode and AMVP mode.

[0296] When performing inter-frame prediction for the current block by predicting geometric modification usage information, this information may not need to be sent. The decoder can predict and obtain the geometric modification usage information using the same method as the encoder. In this case, the selective data transmission unit does not need to encode the geometric modification usage information. Therefore, the amount of data transmitted to the decoder can be reduced.

[0297] Figure 24 This is a diagram illustrating a method for generating geometric modification information according to an embodiment of the present invention.

[0298] In step S2401, the generation method can determine the predictability of geometric modification usage information for the encoded target block (current block). This determination can be performed based on the presence and availability of neighboring blocks adjacent to the encoded target block in time and / or space. For example, when all neighboring blocks are predicted by intra-frame prediction, reference picture information and / or geometric modification usage information of neighboring blocks are unavailable, so the geometric modification usage information of the encoded target block can be determined as unpredictable.

[0299] When the generation method determines in step S2401 that the geometric modification usage information is unpredictable, in step S2402, the generation method can generate the geometric modification usage information. The generation of the geometric modification usage information in step S2402 can be performed by directly searching for a reference image that can yield the best prediction result when predicting the current block between frames, such as... Figure 23 The best prediction reference image search unit is like 2355.

[0300] When the generation method determines in step S2401 that the geometric modification usage information is predictable, in step S2403, the generation method can predict the geometric modification usage information. In step S2403, the generation method can predict the geometric modification usage information based on the inter-frame prediction information of neighboring blocks adjacent to the current block, such as... Figure 23 The geometric modifications were done using the Information Predictor 2360, just as it was done.

[0301] In step S2404, the generation method can determine whether to correct the geometric modification usage information. The geometric modification usage information predicted in step S2403 may not actually match the optimal geometric modification usage information. This mismatch may degrade coding performance. Step S2404 can determine whether such a mismatch exists.

[0302] The encoder can detect performance degradation caused by predicted geometric modification usage information. For example, the encoder may include optimal geometric modification usage information that derives best coding performance by calculating rate-distortion cost (hereinafter referred to as RD cost).

[0303] When the generation method determines in step S2404 that it does not correct the geometric modification usage information—in other words, when the predicted geometric modification usage information has good (or optimal) performance—the generation method may not send the geometric modification usage information of the current block to the decoder. The decoder can predict the geometric modification usage information using the same method used in the encoder. Therefore, the predicted geometric modification usage information can also be used as the geometric modification usage information of the current block.

[0304] When the generation method determines in step S2404 that the geometric modification usage information needs to be corrected—in other words, the predicted geometric modification usage information differs from the optimal geometric modification usage information—the generation method can correct the predicted geometric modification usage information. For example, the generation method can generate the geometric modification usage information using the method in step S2402 instead of using the predicted geometric modification usage information. Alternatively, the generation method can calculate the residual of the symbolized geometric modification usage information based on its frequency of occurrence and can encode only the residual information. Alternatively, the generation method can correct the predicted geometric modification usage information by using the geometric modification usage information of neighboring blocks that are temporally and / or spatially adjacent to the current block.

[0305] Figure 25 This is a configuration diagram of an inter-frame prediction unit including a geometric modification information predictor according to an embodiment of the present invention.

[0306] Figure 25 The inter-frame prediction unit may include a geometrically modified image generator 2510, a geometrically modified image buffer 2520, a geometrically modified usage information predictor 2540, a reconstructed image buffer 2530, and / or an inter-frame predictor 2550.

[0307] The geometrically modified image generator 2510 can generate geometrically modified images based on the relationship between an encoded / decoded target image (the current image) and a reference image. This can be achieved by using a reference image... Figures 8 to 18The method used is to generate a geometrically modified image. Geometric modifications can be performed on a portion or all of a reference image. Additionally, the geometrically modified image generator 2510 can generate an actual geometrically modified image, which is generated from a geometrically modified reference image or a virtual geometrically modified image generated based on the current image and the reference image. A virtual geometrically modified image can refer to a geometrically modified image that includes information capable of deriving the light information of each pixel within the geometrically modified image, rather than the light information of each pixel within the geometrically modified image (e.g., brightness, color, and / or chromaticity). The generated geometrically modified image can be stored in a geometrically modified image buffer 2520.

[0308] The geometric modification usage information predictor 2540 can predict the geometric modification usage information of the current block within the current image. The prediction of geometric modification usage information can be performed based on the inter-frame prediction information (geometric modification usage information) of neighboring blocks adjacent to the current block. The prediction of geometric modification usage information can be achieved by using a reference... Figures 22 to 24 The explained method is used to perform this. The predicted geometric modifications can be used by the inter-frame predictor 2550 to perform inter-frame prediction for the current block.

[0309] Inter-frame predictor 2550 can perform inter-frame prediction for the current block using at least one reconstructed image stored in reconstructed image buffer 2530 and / or at least one geometrically modified image stored in geometrically modified image buffer 2520. Inter-frame predictor 2550 can determine whether to reference a reconstructed image or a corresponding geometrically modified image based on geometric modification usage information.

[0310] Additionally, multiple reference images can be referenced while performing inter-frame prediction. Here, there can be multiple reference image information. For example, when inter-frame prediction is bidirectional, two reference images can be referenced. The reference image information of the two reference images can then be sent to the decoder.

[0311] Figure 26 This is a diagram illustrating an example encoding method for using information to explain geometric modifications.

[0312] To accurately predict the current block, the encoding method can use at least one reference image. In this case, the orientation information of the reference image used in inter-frame prediction (prediction orientation information), the information used to identify the reference image (reference image identification information), and / or the information about whether the reference image has been geometrically modified (geometric modification usage information) can be encoded. In step S2610, the encoding method can determine the prediction orientation based on at least one reference image used for inter-frame prediction of the current block. As a result, the number of reference images used for inter-frame prediction can be checked in step S2610.

[0313] In step S2620, the encoding method can determine whether each reference image used in inter-frame prediction has been geometrically modified. In other words, in step S2610, the encoding method can determine individually whether all reference images are geometrically modified reference images. The results can be simplified by restricting some generation cases.

[0314] In step S2630, the encoding method can generate geometric modification usage information for the current block through steps S2610 and S2620. In step S2640, the encoding method can perform entropy encoding on the generated geometric modification usage information. Entropy encoding can refer to all types of encoding based on the frequency of symbol occurrence. Predicted direction information and / or geometric modification usage information can be combined, and each combination can be represented as a single symbol.

[0315] Furthermore, based on the similarity between symbols, the generated symbols can be modified to have a favorable frequency of occurrence for entropy coding by using residuals between symbols. For example, when the generated symbols are 3, 4, 5, and 6, the generated symbols can be changed to 3, 1, 1, and 1 by generating residuals obtained by subtracting previous symbols from the current symbols. Since entropy coding is based on frequency of occurrence, and 1 is a high frequency of occurrence, the modified symbols can lead to an improvement in the efficiency of entropy coding.

[0316] Figure 27 This is a diagram illustrating another example of an encoding method that uses information to explain geometric modifications. (See reference...) Figure 26 The information encoded via inter-frame prediction may include prediction direction information and / or information about whether each reference image has been geometrically modified. Prediction direction information and / or information about whether each reference image has been geometrically modified may be combined, and each combination may be encoded in a single symbol.

[0317] Figure 27 An example of encoding with two maximum prediction directions is shown.

[0318] exist Figure 27 In this context, HIdx can be a symbol corresponding to a combination of prediction direction information and / or information about whether each reference image is a geometrically modified reference image.

[0319] In step S2710, the encoding method can determine whether to use bidirectional prediction for inter-frame prediction of the encoded target block. When bidirectional prediction is not used, the encoding method can proceed to step S2740.

[0320] When bidirectional prediction is used, in step S2720, the encoding method can determine whether the geometrically modified image is used for both directions of bidirectional prediction.

[0321] In step S2730, when the geometrically modified image is used for two directions of bidirectional prediction in step S2720, the encoding method can determine HIdx as 3.

[0322] In step S2740, when the geometrically modified image in step S2720 does not have two directions for bidirectional prediction, the encoding method can determine whether the reference image used for prediction in the first direction is a geometrically modified reference image. In step S2750, when the reference image used for prediction in the first direction is a geometrically modified reference image, the encoding method can set HIdx to 1.

[0323] In step S2760, when the reference image used for prediction in the first direction is not a geometrically modified reference image, the encoding method can determine whether the reference image used for prediction in the second direction is a geometrically modified reference image. In step S2780, when the reference image used for prediction in the second direction is a geometrically modified reference image, the encoding method can determine HIdx to be 2.

[0324] In step S2770, when the reference image used for second direction prediction in step S2760 is not a geometrically modified reference image, the encoding method can determine HIdx as 0.

[0325] In step S2790, the encoding method can perform entropy encoding on HIdx. The entropy encoding in step S2790 can be performed based on the cumulative occurrence frequency of HIdx.

[0326] Figure 27 This is a diagram illustrating an embodiment of symbols (e.g., HIdx) corresponding to a combination of information that determines and encodes prediction orientation information and / or information about whether each reference image has been geometrically modified. Figure 27 The order of the methods shown can be changed. Additionally, the value assigned to HIdx is the one that distinguishes the combination, and another value can be assigned to HIdx instead of the one described above.

[0327] According to the present invention, a merging mode can be used to efficiently encode inter-frame prediction information. The merging mode uses the motion information of adjacent blocks as is for the motion information of the current block, without any correction. Therefore, the information used for motion correction may not be sent separately to the decoder.

[0328] Figures 28 to 34 This is a diagram illustrating an example of using information to predict geometric modifications by employing merging patterns.

[0329] Figure 28 This is a diagram illustrating an example of encoding an image using a merging pattern.

[0330] exist Figure 28In the diagram, the arrows indicate the corresponding motion vectors of the blocks. Figure 27 Blocks with similar motion vectors can be grouped together, and the grouped blocks can be represented by several different regions, such as... Figure 28 As shown, the merge mode examines the motion information of candidate blocks adjacent to the current block, and merges the candidate block and the current block into the same group when the motion information of the candidate block is similar to that of the current block. Therefore, the motion information of a specific block can be used for another block adjacent to that specific block.

[0331] When a merge mode is applied to the current block, merge information can be encoded, rather than motion information. Merge information may include merge flags indicating whether to perform the merge mode for each block partition, and information about selecting a merge candidate from a list of merge candidates that includes at least one merge candidate adjacent to the current block. See later. Figure 29 Describe the at least one merge candidate. The motion information of the selected merge candidate can be used as the motion information of the current block.

[0332] Figure 29 This is a diagram showing adjacent blocks included in the merge candidate list.

[0333] exist Figure 29 In this context, the current block X corresponding to a single CU is divided into two PUs. The candidate list for merging the two PUs can be configured based on a single CU that includes both PUs. In other words, the candidate list can be configured using adjacent blocks that are adjacent to the CU. For example, when motion information for adjacent block A0 is available, block A0 can be selected and inserted into the candidate list. Figure 29 As shown, adjacent blocks may include A0, B1, B0, A0, and B2, and / or blocks H (or M) located at the same position within the reference image. The merge candidate list may include a predetermined number of candidates. Adjacent blocks may be included in the merge candidate list based on a predetermined order. The predetermined order may be A0→B1→B0→A0→B2→H (or M).

[0334] Here, block H (or M) refers to a candidate block used to obtain temporal motion information. Block H can refer to the block located to the lower right of block X' within the reference image, where block X' is located at the same position as the current CU(X). Block M can refer to one of the blocks within block X' located within the reference image. Based on block X', blocks H and M can be used for temporal candidates to configure the merged candidate list. Block M can be used when motion information for block H is unavailable.

[0335] The merge candidate list may include merge candidates generated based on combinations of merge candidates included in the merge candidate list.

[0336] Each merging candidate may include motion information. According to the present invention, each merging candidate may include not only motion information, but also information regarding whether a geometrically modified reference image is used (geometric modification usage information), or may include a structure for obtaining the information.

[0337] Figures 30 to 33 This is a diagram illustrating a method for using information to predict geometric modifications according to an embodiment of the present invention.

[0338] The encoding apparatus according to the invention can perform inter-frame prediction by referring to at least one reference image. Alternatively, the encoding apparatus can refer to a geometrically modified reference image generated by geometrically modifying the reference image to perform inter-frame prediction accurately. Therefore, the encoding apparatus may need to determine whether to use the geometrically modified reference image while performing inter-frame prediction for the target block. As information regarding whether the geometrically modified reference image is used, geometric modification usage information can be transmitted by signaling. Alternatively, geometric modification usage information can be predicted from adjacent blocks using a method described later.

[0339] Figure 30 This is a diagram explaining the steps involved in generating a geometrically modified reference image.

[0340] Figure 30 Image relationship identifier 3020 can correspond to Figure 22 Image relationship identifier 2210. Image relationship identifier 3020 can derive geometric modification information that can geometrically modify a reference image similar to the encoded target image. Geometric modification information can be expressed as a geometric modification matrix or a relationship between pixel positions.

[0341] Figure 30 The image geometry modification unit 3010 can correspond to Figure 22 Image geometric modification unit 2220. Image geometric modification unit 3010 can generate a geometrically modified reference image by geometrically modifying a reference image using geometric modification information derived from the image relationship identifier 3020. Reference image N refers to a number of reference images N, and geometrically modified reference image N' refers to a geometrically modified reference image generated by geometrically modifying the reference image N.

[0342] Figure 31 This is an example diagram showing the configuration of inter-frame prediction information included in blocks within the encoded target image.

[0343] The reference image buffer can include N reference images. The geometry modification reference image buffer can include N geometry modification reference images, each corresponding to one of the N reference images. Figure 30The method generates the reference image by geometrically modifying it. Therefore, the geometrically modified reference image buffer can have the same number of geometrically modified reference images as the reference image buffer. C1, C2, and C3 are blocks that are encoded / decoded within the target image. Inter-frame prediction for each block can be performed by referencing images from each other. C1', C2', and C3' refer to the regions referenced by C1, C2, and C3, respectively.

[0344] Inter-frame prediction information may include motion vectors and reference image numbers, and / or geometry modification usage information. Motion vectors refer to the positional difference between the current block and the reference block. MV1, MV2, and MV3 refer to motion vectors indicating the positional differences within the frame between C1 and C1', C2 and C2, and C3 and C3', respectively. For example, when the position of C1 within the current frame is (10, 5) and the position of C1' within the reference image N is (13, 7), MV1 becomes (3, 2). Reference image number refers to the number (or index) of the reference image used to specify the reference image within the reference image buffer. In the case of geometry modification reference images, the reference image number refers to the number of the reference image used to generate the geometry modification reference image. Geometry modification usage information indicates whether the block's position reference is located within a geometry modification reference image. Geometry modification usage information of 0 indicates that a geometry modification reference image is used. Alternatively, geometry modification usage information of X indicates that a geometry modification reference image is not used.

[0345] Figure 31 The illustrated embodiment relates to unidirectional prediction, and geometric modification usage information can be represented as O / X. When referencing two or more reference images while performing inter-frame prediction, this can be achieved by defining and referencing... Figure 26 and 27 The information is used to modify the geometry for each corresponding symbol value in all possible combinations described.

[0346] Figure 31 The inter-frame prediction information configured in the configuration can be used as prediction candidates for the inter-frame prediction information of the current block, which will be described later.

[0347] Figure 32 This is a diagram illustrating an embodiment of predicting inter-frame prediction information for the current block X within the current image.

[0348] Inter-frame prediction information for the current block X, including motion vectors, reference image numbers, and / or geometric modification usage information, can be predicted based on the inter-frame prediction information of neighboring blocks adjacent to the current block X. Neighboring blocks adjacent to the current block X may include reference... Figure 29 The described temporal and / or spatially adjacent blocks. Figure 32In the process, blocks C1, C2, and C3, which have already been encoded and are spatially adjacent to the current block X, are used as neighboring blocks. Depending on the merging mode, neighboring blocks C1, C2, and C3 can become merging candidates when predicting inter-frame prediction information for the current block X. Figure 32 The configuration includes 3 merge candidates, but is not limited to this.

[0349] According to the present invention, information on geometric modifications can be included in the information derived through merging pattern prediction. It can be calculated... Figure 32 The prediction cost is calculated for the merged candidates C1, C2, and C3. The prediction cost can be calculated as the ratio of the prediction accuracy derived from the prediction result to the number of bits generated during encoding. Prediction accuracy refers to the similarity between the pixel values ​​of the predicted reference block in the reference image and the current block in the current image. A high prediction cost means that a large number of bits are required to reconstruct a block of the same quality. Generally, the higher the prediction accuracy, the lower the prediction cost.

[0350] Figure 33 This is a diagram illustrating an embodiment of predicting inter-frame prediction information for the current block based on merging candidates.

[0351] Figure 33 The embodiments are based on Figure 32 The inter-frame prediction information for the current block X is predicted from the merge candidate block C2 with the lowest prediction cost. Specifically, the motion vectors, reference image number, and / or geometric modification usage information of block C2 can be set as the motion vectors, reference image number, and / or geometric modification usage information for the current block X. The inter-frame prediction information for the current block X can be predicted from the merge candidate block C2. When using merge mode, the encoder and decoder can generate the same merge candidate list. Therefore, the encoder only sends the decoder information (e.g., merge index) for selecting merge candidates from the merge candidate list, not the inter-frame prediction information for the current block X.

[0352] exist Figure 33 In the embodiment, since the inter-frame prediction information of the current block X is derived from the merge candidate block C2, the predicted block of the current block X and the predicted block of the merge candidate block C2 are adjacent in the same image.

[0353] For adjacent blocks that are temporally and / or spatially contiguous, whether a geometrically modified reference image is used depends on whether they are identical or similar. This invention uses these features to predict information about whether a geometrically modified reference image is used. According to this invention, coding efficiency is improved because geometric modification usage information is not sent. For example, when matching the prediction results of the encoder and decoder using the same prediction method, the transmission of additional information besides the prediction information can be omitted.

[0354] Figure 34This is a diagram illustrating an embodiment of a method for predicting inter-frame prediction information for the current block based on a merged candidate list.

[0355] Figure 34 The list of merge candidates can include five merge candidates. Descriptions and references using motion vectors, reference image numbers, and / or geometric modifications are also included. Figures 30 to 33 The description is the same. Since merge candidate 1 has the best encoding performance among the merge candidates, this method can use the information of merge candidate 1 as the information of the current block.

[0356] Advanced Motion Vector Prediction (AMVP) mode, instead of merging mode, can be used to predict motion vectors. When using AMVP mode, motion information can also be predicted by using neighboring blocks adjacent to the current block. Motion information can be found using the same method as in merging mode; however, AMVP mode can also include the transmission of residual values ​​of motion information, prediction directions, and / or reference image index information. The prediction of geometric modification usage information according to the present invention can also be applied to AMVP mode. For example, AMVP mode can derive the geometric modification usage information of the current block by using the geometric modification usage information of candidate blocks as a merging mode, since AMVP mode finds candidate blocks as examples of merging modes. The geometric modification usage information of candidate blocks can be used as is for the geometric modification usage information of the current block. Alternatively, the residual value between the geometric modification usage information of the current block and the candidate block can be transmitted via a bitstream, and the geometric modification usage information of the current block can be derived from the geometric modification usage information of the candidate blocks and the residual value.

[0357] The inter-frame prediction unit of a decoder that performs inter-frame prediction using two reference images needs to know information about both reference images. The reference images used for inter-frame prediction can be images of decoded images stored in a reference image buffer (reconstructed image buffer), or geometrically modified reference images generated by geometrically modifying decoded images. The reference image buffer manages geometrically modified reference images by matching decoded images with their corresponding geometrically modified reference images.

[0358] Here, the information sent from the encoder to the decoder may include: an identifier that can distinguish reference images within the reference image buffer (matching the already decoded image and its corresponding geometrically modified reference image), and an identifier that can identify the use of the geometrically modified reference image (e.g., geometric modification usage information).

[0359] The geometric modification usage information predictor of this invention can predict the geometric modification usage information of the current block based on the modification usage information of neighboring blocks. Specifically, the geometric modification usage information of the current block can be predicted using statistics on the geometric modification usage information of neighboring blocks. These statistics may include average, median, distribution, maximum, and / or minimum values, etc. Alternatively, the geometric modification usage information of the current block can be predicted using the modification usage information of neighboring blocks selected from among the neighboring blocks. The predicted geometric modification usage information can be used as is for the geometric modification usage information of the current block. Alternatively, the predicted geometric modification usage information can be corrected and then used for the geometric modification usage information of the current block. According to this invention, since the geometric modification usage information of the current block is predicted through the geometric modification usage information of neighboring blocks, the transmission of the geometric modification usage information of the current block can be omitted.

[0360] Furthermore, by using the consistency and statistical frequency of predicted geometric modification usage information to remove or reduce redundant information, geometric modification usage information can be encoded with symbols of small size.

[0361] Furthermore, according to the present invention, a reference image or a geometrically modified reference image (geometrically modified image) generated by geometrically modifying a reference image can be used to predict the encoded / decoded target block. When using a merging mode, the merged candidate block and the encoded / decoded target block can have similar motion information. Therefore, the information identifying the use of the geometrically modified reference image while predicting the encoded / decoded target block can be similar to the geometrically modified usage information of the merged candidate block. According to embodiments of the present invention, geometrically modified usage information and motion information can be derived by using these features.

[0362] Furthermore, according to embodiments of the present invention, image information can be expressed using different symbols. Each symbol may have a biased frequency of occurrence. Entropy coding refers to an encoding method that considers the frequency of occurrence of symbols. According to entropy coding, symbols with higher frequency of occurrence are represented as symbols with smaller sizes, and symbols with lower frequency of occurrence are represented as symbols with larger sizes, thus improving encoding efficiency. The encoder can calculate the frequency of occurrence of each symbol. However, the decoder may not receive information about the frequency of occurrence of each symbol. Therefore, the frequency of occurrence of each symbol can be predicted based on the frequency of occurrence of symbols up to the point where encoding / decoding has been performed.

[0363] Tables 1 and 2 are views illustrating examples of syntax configurations for the coding unit (CU) and prediction unit (PU) according to the present invention. In Tables 1 and 2, a cu_skip_flag being true indicates the use of the merge skip mode. The merge skip mode refers to a merge mode in which the transmission of additional information is further omitted than the transmission of the merge mode. Therefore, like the merge mode, the merge skip mode can be applied to the present invention, and thus there is no need to additionally transmit geometric modification usage information.

[0364] The `merge_flag` indicates whether the merge mode is used. When the merge mode is used and applied to this invention, it is not necessary to send geometry modification usage information. Alternatively, geometry modification usage information can be sent when the merge mode is not used.

[0365] Tables 1 and 2 provide examples of the syntax configurations for CU and PU when the merge mode is applied in this invention. When the merge mode is not used, relevant syntax for including geometric modification usage information can be included in the bitstream.

[0366] For example, the modification_image_reference_type indicates geometric modification usage information and is only sent when merge mode is not used. In the examples in Tables 1 and 2, when merge mode is used, the modification_image_reference_type is not included in the bitstream. Therefore, by not signaling geometric modification usage information when using merge mode, the amount of data transmitted can be reduced.

[0367] [Table 1]

[0368]

[0369]

[0370] [Table 2]

[0371]

[0372]

[0373]

[0374] In the above embodiments, the method is described based on a flowchart having a series of steps or units. However, the present invention is not limited to the order of these steps, but some steps may be performed simultaneously with other steps or in a different order. Furthermore, those skilled in the art should understand that the steps in the flowchart do not exclude each other, and other steps can be added to the flowchart, or some steps can be deleted from the flowchart without affecting the scope of the present invention.

[0375] The above description includes examples of various aspects. It is certainly impossible to describe every possible combination of components or methods in order to describe each aspect, but one skilled in the art will recognize that many further combinations and permutations are possible. Therefore, this subject matter specification is intended to cover all such changes, modifications, and variations that fall within the spirit and scope of the appended claims.

[0376] Computer-readable storage media may include, individually or in combination, program instructions, data files, data structures, etc. Program instructions recorded in a computer-readable storage medium may be any program instructions specifically designed and constructed for this invention or known to those skilled in the art of computer software. Examples of computer-readable storage media include magnetic recording media (such as hard disks, floppy disks, and magnetic tapes); optical data storage media (such as CD-ROMs or DVD-ROMs); magneto-optical media (such as optical disks); and hardware devices specifically configured to store and implement program instructions (such as read-only memory (ROM), random access memory (RAM), and flash memory). Examples of program instructions include not only machine language code formatted by a compiler but also high-level language code that can be implemented by a computer using an interpreter. Hardware devices may be configured to operate (or vice versa) by one or more software modules to perform the processes according to the invention.

[0377] Although the invention has been described with reference to specific entries such as detailed elements and limited embodiments and drawings, these are provided only to aid in a more general understanding of the invention, and the invention is not limited to the embodiments described above. Those skilled in the art will understand that various modifications and changes can be made from the above description.

[0378] Therefore, the spirit of the present invention should not be limited to the above embodiments, and the entire scope of the appended claims and their equivalents shall fall within the scope and spirit of the present invention.

[0379] Industrial applicability

[0380] This invention can be used to encode / decode images.

Claims

1. A method for decoding an image, the method comprising: determining whether a merge mode is applied to a current block; when the merge mode is applied to the current block, determining whether geometric modification information is used for inter prediction of the current block based on geometric modification usage information of the current block; when the geometric modification usage information indicates that the geometric modification information is used for inter prediction of the current block, obtaining a prediction block of the current block within a decoded target picture by performing inter prediction based on a reference picture and the geometric modification information; and obtaining a reconstructed block of the current block by summing the prediction block and a residual block, wherein geometric modification for the geometric modification information comprises affine modification, and wherein the geometric modification information and a reference picture index specifying the reference picture are derived from a merge candidate among a merge candidate list of the current block, and the merge candidate list comprises at least one spatial merge candidate derived from at least one neighboring block of the current block.

2. The method of claim 1, wherein when the geometric modification usage information indicates that the geometric modification information is used, 3 control points of the current block are used to obtain the prediction block of the current block.

3. The method of claim 1, wherein the geometric modification information consists of 6 parameters.

4. The method of claim 1, wherein the geometric modification is one of geometric modification types, and the geometric modification types comprise at least one of geometric shift modification, size modification, rotation modification, affine modification and projection modification.

5. A method for encoding an image, the method comprising: obtaining a prediction block of a current block within an encoded target picture by performing inter prediction based on a reference picture and geometric modification information; determining whether a merge mode is applied to the current block; when the merge mode is applied to the current block, encoding geometric modification usage information of the current block, the geometric modification usage information indicating whether geometric modification information is used for inter prediction of the current block; and obtaining a residual block of the current block by subtracting the prediction block from an original block, wherein geometric modification for the geometric modification information comprises affine modification, wherein the geometric modification information is derived based on a merge candidate among a merge candidate list of the current block, and a reference picture index specifying the reference picture of the current block is the same as a reference picture index of the merge candidate, and wherein the merge candidate list comprises at least one spatial merge candidate derived from at least one neighboring block of the current block.

6. An apparatus for transmitting compressed video data, comprising: a processor configured to obtain the compressed video data; and a transmitter configured to transmit the compressed video data, wherein obtaining the compressed video data comprises: obtaining a prediction block of a current block within an encoded target picture by performing inter prediction based on a reference picture and geometric modification information; determining whether a merge mode is applied to the current block; ​ when the merge mode is applied to the current block, coding geometry modification usage information for the current block, the geometry modification usage information indicating whether geometry modification information is used for inter prediction of the current block; and obtaining a residual block of the current block by subtracting the prediction block from the original block, wherein the geometry modification for the geometry modification information comprises affine modification, wherein the geometry modification information is derived based on a merge candidate among a merge candidate list of the current block, and specifies that a reference picture index of the reference picture of the current block is the same as a reference picture index of the merge candidate, and wherein the merge candidate list comprises at least one spatial merge candidate derived by at least one neighboring block of the current block.

Citation Information

Patent Citations

  • Method for encoding and decoding image information and device using same

    CN103299642A

  • Inter-frame prediction method in hybrid video coding standard

    CN104935938A