Video encoding / decoding method

By deriving motion vectors in the image decoding device and using reconstructed pixel regions for inter-frame prediction, the problem of reduced coding efficiency and block artifacts in high-efficiency video coding with many block motions is solved, achieving more efficient coding and decoding.

CN116193109BActive Publication Date: 2026-07-24IND ACAD COOP GRP OF SEJONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IND ACAD COOP GRP OF SEJONG UNIV
Filing Date
2018-01-16
Publication Date
2026-07-24

Smart Images

  • Figure CN116193109B_ABST
    Figure CN116193109B_ABST
Patent Text Reader

Abstract

The present invention relates to an image encoding / decoding method, and an image decoding method and apparatus suitable for an embodiment of the present invention, wherein after selecting a reconstructed pixel region within an image to which a current block belongs that needs to be decoded, a motion vector of the reconstructed pixel region is derived based on the reconstructed pixel region and a reference image of the current block, and the derived motion vector is selected as a motion vector of the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the application filed with the State Intellectual Property Office on January 16, 2018, entitled "Image Encoding / Decoding Method and Apparatus" with application number 201880007055.1. Technical Field

[0002] This invention relates to an image signal encoding / decoding method and apparatus, and more particularly to an image encoding / decoding apparatus and method utilizing inter-frame prediction. Background Technology

[0003] In recent years, the demand for multimedia data such as video on the Internet has been increasing dramatically. However, the current development speed of channel bandwidth is insufficient to fully meet the rapidly increasing volume of multimedia data. Considering the above situation, the Video Coding Expert Group (VCEG) of the International Telecommunication Union Telecommunication Standardization Sector (ITU-T) and the Moving Picture Expert Group (MPEG) of the International Organization for Standardization / International Electrotechnical Commission (ISO / IEC) developed the first version of the video compression standard, High Efficiency Video Coding (HEVC), in February 2014.

[0004] High-efficiency video coding (HEVC) employs various techniques such as intra-frame prediction, inter-frame prediction, transform, quantization, entropy coding, and loop filtering. In inter-frame prediction within HEVC, efficiency can be improved by applying new techniques such as block merging and Advanced Motion Vector Prediction (AMVP). However, when multiple motions exist within a block, further segmentation of the block can lead to a sharp increase in overhead and a decrease in coding efficiency. Summary of the Invention

[0005] Technical issues

[0006] To address the problems described above, the main objective of this invention is to improve the efficiency of inter-frame prediction by providing an improved inter-frame prediction method.

[0007] Furthermore, the main objective of this invention is to provide a motion vector derivation method using an image decoding device that does not require the image encoding device to transmit motion vector information to the decoding device.

[0008] Furthermore, the main objective of this invention is to provide a method for deriving control point motion vectors using an image decoding device, which eliminates the need for the image encoding device to transmit control point motion vectors to the decoding device during affine inter-frame prediction.

[0009] Furthermore, the main objective of this invention is to provide an inter-frame prediction method that can effectively perform encoding or decoding even when multiple motions exist in a block.

[0010] Furthermore, the main objective of this invention is to reduce block effects that may occur when a block is divided into multiple regions and encoding or decoding is performed using different inter-frame predictions.

[0011] Furthermore, the main objective of this invention is to improve the efficiency of inter-frame prediction by segmenting the current block that needs to be encoded or decoded based on the segmentation structure of the surrounding reconstructed blocks.

[0012] Furthermore, the main objective of this invention is to improve the efficiency of inter-frame prediction by segmenting the reconstructed surrounding image regions used in the encoding or decoding process of the current block based on the segmentation structure of the surrounding reconstructed blocks.

[0013] Furthermore, the main objective of this invention is to improve the encoding or decoding efficiency of images by performing encoding or decoding using the current block or surrounding images segmented in the manner described above.

[0014] Technical solution

[0015] An image decoding method and apparatus according to one embodiment of the present invention, after selecting a reconstructed pixel region within an image to which the current block to be decoded belongs, derives a motion vector of the reconstructed pixel region based on the reconstructed pixel region and a reference image of the current block, and selects the derived motion vector as the motion vector of the current block.

[0016] The reconstructed pixel region can include at least one of the regions adjacent to the upper side of the current block or the regions adjacent to the left side of the current block.

[0017] The motion vector of the reconstructed pixel region can be derived based on the position of the region corresponding to the reconstructed pixel region determined in the reference image.

[0018] An image encoding method and apparatus according to one embodiment of the present invention, after selecting a reconstructed pixel region within an image to which the current block to be encoded belongs, derives a motion vector of the reconstructed pixel region based on the reconstructed pixel region and a reference image of the current block, and selects the derived motion vector as the motion vector of the current block.

[0019] Furthermore, the image encoding method and apparatus according to one embodiment of the present invention can encode the image after generating motion vector derivation indication information on the decoder side.

[0020] The aforementioned decoder-side motion vector derivation indication information can indicate whether to select the motion vector derived in the reconstructed pixel region as the motion vector of the current block.

[0021] An image decoding method and apparatus according to another embodiment of the present invention selects at least one reconstructed pixel region within the image to which the current block to be decoded using affine inter-frame prediction. Then, based on the at least one reconstructed pixel region and the reference image of the current block, motion vectors of the at least one reconstructed pixel region are derived respectively, and the motion vectors of the at least one reconstructed pixel region are selected respectively as motion vectors of at least one control point of the current block.

[0022] The aforementioned at least one reconstructed pixel region can be a region that is adjacent to at least one control point of the aforementioned current block.

[0023] Furthermore, at least one of the aforementioned control points can be located on the upper left, upper right, or lower left side of the current block.

[0024] Furthermore, the motion vector of the control point located on the lower right side of the current block can be decoded based on the motion information contained in the bitstream.

[0025] Furthermore, another embodiment of the image decoding method and apparatus applicable to this invention can decode the motion vector derivation indication information of the decoder-side control point.

[0026] The image decoding method and apparatus according to another embodiment of the present invention can derive the motion vectors of at least one reconstructed pixel region based on the motion vector derivation instruction information of the decoder-side control point.

[0027] An image coding method and apparatus according to another embodiment of the present invention selects at least one reconstructed pixel region within the image to which the current block to be coded using affine inter-frame prediction. Then, based on the at least one reconstructed pixel region and the reference image of the current block, motion vectors of the at least one reconstructed pixel region are derived respectively, and the motion vectors of the at least one reconstructed pixel region are selected respectively as motion vectors of at least one control point of the current block.

[0028] An image decoding method and apparatus according to another embodiment of the present invention can divide the current block to be decoded into multiple regions including a first region and a second region and obtain the prediction block of the first region and the prediction block of the second region. The prediction block of the first region and the prediction block of the second region can be obtained by different inter-frame prediction methods.

[0029] The first region mentioned above can be a region adjacent to a reconstructed image region within the image to which the current block belongs, and the second region mentioned above can be a region not adjacent to a reconstructed image region within the image to which the current block belongs.

[0030] The image decoding method and apparatus according to another embodiment of the present invention can infer the motion vector of the first region based on the reconstructed image region within the image to which the current block belongs and the reference image of the current block.

[0031] An image decoding method and apparatus according to another embodiment of the present invention can divide a region located at the boundary into multiple sub-blocks as a region within the prediction block of the first region or divide a region located at the boundary into multiple sub-blocks as a region within the prediction block of the second region. After generating a prediction block of the first sub-block using motion information of one of the multiple sub-blocks, namely the sub-blocks surrounding the first sub-block, a prediction block of the first sub-block with weighted summation is obtained by performing a weighted summation on the first sub-block and the prediction block of the first sub-block.

[0032] An image coding method and apparatus according to another embodiment of the present invention can divide a current block to be encoded into multiple regions including a first region and a second region and obtain a prediction block of the first region and a prediction block of the second region. The prediction blocks of the first region and the prediction blocks of the second region can be obtained by different inter-frame prediction methods.

[0033] The first region mentioned above can be a region adjacent to the encoded and reconstructed image region within the image to which the current block belongs, and the second region mentioned above can be a region not adjacent to the encoded and reconstructed image region within the image to which the current block belongs.

[0034] The image coding method and apparatus according to another embodiment of the present invention can infer the motion vector of the first region based on the encoded reconstructed image region within the image to which the current block belongs and the reference image of the current block.

[0035] An image coding method and apparatus according to another embodiment of the present invention can divide a region located at the boundary into multiple sub-blocks as a region within the prediction block of the first region or divide a region located at the boundary into multiple sub-blocks as a region within the prediction block of the second region. After generating a prediction block of the first sub-block using motion information of one of the multiple sub-blocks, namely the sub-blocks surrounding the first sub-block, a prediction block of the first sub-block with weighted summation is obtained by performing a weighted summation on the first sub-block and the prediction block of the first sub-block.

[0036] An image decoding method and apparatus according to another embodiment of the present invention can divide the current block into multiple sub-blocks based on the blocks surrounding the current block to be decoded, and decode the multiple sub-blocks of the current block.

[0037] The image decoding method and apparatus according to another embodiment of the present invention can divide the current block into multiple sub-blocks based on the segmentation structure of the surrounding blocks of the current block.

[0038] An image decoding method and apparatus according to another embodiment of the present invention can divide the current block into multiple sub-blocks based on at least one of the following: the number of peripheral blocks, the size of the peripheral blocks, the shape of the peripheral blocks, and the boundaries between the peripheral blocks.

[0039] An image decoding method and apparatus according to another embodiment of the present invention can divide a reconstructed pixel region into sub-block units as a region located around the current block, and decode at least one of a plurality of sub-blocks of the current block using at least one sub-block contained in the reconstructed pixel region.

[0040] The image decoding method and apparatus according to another embodiment of the present invention can divide the reconstructed pixel region into sub-block units based on the segmentation structure of the surrounding blocks of the current block.

[0041] The image decoding method and apparatus according to another embodiment of the present invention can divide the reconstructed pixel region into sub-block units based on at least one of the following: the number of peripheral blocks, the size of the peripheral blocks, the shape of the peripheral blocks, and the boundaries between the peripheral blocks.

[0042] An image encoding method and apparatus according to another embodiment of the present invention divides the current block into multiple sub-blocks based on the blocks surrounding the current block and encodes the multiple sub-blocks of the current block.

[0043] The image encoding method and apparatus according to another embodiment of the present invention can divide the current block into multiple sub-blocks based on the segmentation structure of the surrounding blocks of the current block.

[0044] The image encoding method and apparatus according to another embodiment of the present invention can divide the current block into multiple sub-blocks based on at least one of the number of peripheral blocks, the size of the peripheral blocks, the shape of the peripheral blocks, and the boundaries between the peripheral blocks.

[0045] An image encoding method and apparatus according to another embodiment of the present invention can divide a reconstructed pixel region into sub-block units as a region located around the current block, and encode at least one of a plurality of sub-blocks of the current block using at least one sub-block contained in the reconstructed pixel region.

[0046] The image encoding method and apparatus according to another embodiment of the present invention can divide the reconstructed pixel region into sub-block units based on the segmentation structure of the surrounding blocks of the current block.

[0047] The image encoding method and apparatus according to another embodiment of the present invention can divide the reconstructed pixel region into sub-block units based on at least one of the number of peripheral blocks, the size of the peripheral blocks, the shape of the peripheral blocks, and the boundaries between the peripheral blocks.

[0048] Beneficial effects

[0049] This invention reduces the amount of encoded information generated during video encoding, thereby improving encoding efficiency. Furthermore, adaptive decoding of encoded images enhances image reconstruction efficiency and improves the picture quality of the played images.

[0050] Furthermore, the inter-frame prediction applicable to this invention does not require the image coding device to transmit motion vector information to the decoding device, thus reducing the amount of coding information and thereby improving its coding efficiency.

[0051] Furthermore, this invention can reduce block effects that may occur when a block is divided into multiple regions and encoding or decoding is performed using different inter-frame predictions. Attached Figure Description

[0052] Figure 1 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied.

[0053] Figure 2 This is a schematic diagram used to illustrate a method for generating motion information using motion inference based on existing technology.

[0054] Figure 3 This is a schematic diagram illustrating an instance of a surrounding block that can be used in the process of generating motion information for the current block.

[0055] Figure 4 This is a block diagram illustrating an image decoding apparatus to which one embodiment of the present invention is applied.

[0056] Figure 5a as well as Figure 5b This is a schematic diagram illustrating inter-frame prediction using reconstructed pixel regions applicable to the first embodiment of the present invention.

[0057] Figure 6 This is a sequence diagram for illustrating the inter-frame prediction method to which the first embodiment of the present invention applies.

[0058] Figures 7a to 7c This is a schematic diagram illustrating an embodiment of the reconstructed pixel region.

[0059] Figure 8 This is a sequence diagram illustrating the decision process of an inter-frame prediction method applicable to one embodiment of the present invention.

[0060] Figure 9 This is a schematic diagram illustrating the encoding process used to indicate the inter-frame prediction method.

[0061] Figure 10 It is used to follow such Figure 9 The diagram illustrates the decoding process of decoder-side motion vector derivation (DMVD) indication information encoded in the manner shown.

[0062] Figure 11 This is a schematic diagram used to illustrate affine inter-frame prediction.

[0063] Figure 12a as well as Figure 12b This is a schematic diagram illustrating the derivation of motion vectors for control keys using reconstructed pixel regions according to the second embodiment of the present invention.

[0064] Figure 13 This is a sequence diagram for illustrating the inter-frame prediction method applicable to the second embodiment of the present invention.

[0065] Figure 14 This is a sequence diagram illustrating the decision process of the inter-frame prediction method to which the second embodiment of the present invention applies.

[0066] Figure 15 For indicating passage Figure 14 The diagram illustrates the encoding process of information determined by the inter-frame prediction method.

[0067] Figure 16 It is used to follow such Figure 15 The diagram illustrates the decoding process of decoder-side control point motion vector derivation (DCMVD) indication information encoded in the manner shown.

[0068] Figure 17 This is a sequence diagram illustrating an embodiment of an image decoding method that uses reconstructed pixel regions to derive motion vectors of three control points and generate a predicted block for the current block.

[0069] Figure 18 This is a schematic diagram illustrating a current block that is divided into multiple regions for inter-frame prediction according to the third embodiment of the present invention.

[0070] Figure 19 This is a sequence diagram for illustrating the inter-frame prediction method applicable to the third embodiment of the present invention.

[0071] Figure 20 This is a schematic diagram illustrating an embodiment of motion inference and motion compensation using reconstructed pixel regions.

[0072] Figure 21 This is a sequence diagram illustrating the decision process of an inter-frame prediction method applicable to one embodiment of the present invention.

[0073] Figure 22 This is a schematic diagram illustrating the process of transmitting information indicating the inter-frame prediction method to the image decoding device.

[0074] Figure 23 This is a schematic diagram illustrating the decoding process for information used to indicate which method of inter-frame prediction was used.

[0075] Figure 24 This is a schematic diagram illustrating the method of generating inter-frame prediction blocks using information indicating which method of inter-frame prediction was used.

[0076] Figure 25This is a sequence diagram used to illustrate the method of reducing block effects applicable to the fourth embodiment of the present invention.

[0077] Figure 26 This is a schematic diagram illustrating the applicable method for weighted summation of sub-blocks within a prediction block and the sub-blocks adjacent to the top.

[0078] Figure 27 This is a diagram illustrating the applicable method for weighted summation of sub-blocks within a prediction block and the sub-blocks adjacent to it on the left.

[0079] Figure 28 This is a sequence diagram illustrating the process of determining whether the weighted summation between sub-blocks is applicable.

[0080] Figure 29 It is a sequence diagram illustrating the encoding process of information used to indicate whether a weighted summation between sub-blocks is applicable.

[0081] Figure 30 This is a sequence diagram illustrating the decoding process of information used to indicate whether a weighted summation between sub-blocks is applicable.

[0082] Figure 31a as well as Figure 31b This is a schematic diagram illustrating inter-frame prediction using reconstructed pixel regions according to the fifth embodiment of the present invention.

[0083] Figure 32 This is a schematic diagram illustrating the scenario where additional motion predictions are performed on the current block in units of sub-blocks.

[0084] Figure 33 This is a schematic diagram illustrating an instance where the reconstructed pixel region and the current block are divided into sub-block units.

[0085] Figure 34 This is a sequence diagram illustrating an example of an inter-frame prediction method that utilizes reconstructed pixel regions.

[0086] Figure 35 This is a schematic diagram illustrating an example of how the present invention divides a reconstructed pixel region into sub-blocks using reconstructed blocks existing around the current block.

[0087] Figure 36 This is a schematic diagram illustrating an example of how the present invention divides the current block into multiple sub-blocks using reconstruction blocks existing around the current block.

[0088] Figure 37 This is a sequence diagram illustrating a method for dividing a current block into multiple sub-blocks according to one embodiment of the present invention.

[0089] Figure 38 This is a sequence diagram illustrating a method for dividing a reconstructed region into multiple sub-blocks during the encoding or decoding process of the current block, according to one embodiment of the present invention.

[0090] Figure 39 It is to utilize according to such Figure 36 The illustrated sequence diagram shows an embodiment of an inter-frame prediction method for sub-blocks of the current block segmented in the manner shown.

[0091] Figure 40 According to Figure 39 The diagram illustrates the encoding method for information determined by inter-frame prediction.

[0092] Figure 41 According to Figure 40 The diagram illustrates an example of a decoding method for information encoded using the encoding method shown in the figure.

[0093] Figure 42a And 42b is a schematic diagram for illustrating the sixth embodiment to which the present invention is applicable.

[0094] Figure 43 For reference Figure 42a as well as Figure 42b A sequence diagram illustrating an example of the inter-frame prediction mode determination method applicable to the sixth embodiment of the present invention is provided.

[0095] Figure 44 According to Figure 43 The diagram illustrates the information encoding process determined by the method shown.

[0096] Figure 45 According to Figure 44 The diagram illustrates the decoding process of the information encoded by the method shown. Detailed Implementation

[0097] This invention is capable of various modifications and has many different embodiments. Specific embodiments will be illustrated and described in detail below with reference to the accompanying drawings. However, the following description is not intended to limit the invention to specific implementations, but should be understood to include all modifications, equivalents, and substitutions within the scope of the invention's concept and technology. Similar reference numerals are used for similar constituent elements in the description of the various drawings.

[0098] In describing different constituent elements, terms such as "first" and "second" may be used, but the constituent elements are not limited by these terms. These terms are merely used to distinguish one constituent element from others. For example, without departing from the scope of the claims, a first constituent element can also be named a second constituent element, and similarly, a second constituent element can also be named a first constituent element. The term "and / or" includes a combination of multiple related descriptions or one of multiple related descriptions.

[0099] When a constituent element is described as being "connected" or "in contact" with other constituent elements, it should be understood that it can not only be directly connected or in contact with the aforementioned other constituent elements, but also that other constituent elements can exist between the two. Conversely, when a constituent element is described as being "directly connected" or "directly in contact" with other constituent elements, it should be understood that no other constituent elements exist between the two.

[0100] The terminology used in this application is for illustrative purposes only and is not intended to limit the invention. Singular statements also have plural meanings unless the context clearly indicates otherwise. In this application, terms such as "comprising" or "having" are used only to indicate the presence of features, numbers, steps, actions, constituent elements, components, or combinations thereof as described in the specification, and should not be construed as excluding the possibility of one or more other features, numbers, steps, actions, constituent elements, components, or combinations thereof being present or added.

[0101] The embodiments to which the present invention applies will now be described in detail with reference to the accompanying drawings. In the following description, the same reference numerals will be used for the same constituent elements in the drawings, and repeated descriptions of the same constituent elements will be omitted.

[0102] Figure 1 This is a block diagram illustrating an image encoding apparatus to which one embodiment of the present invention is applied.

[0103] See Figure 1 The image encoding apparatus 100 may include an image segmentation unit 101, an intra-frame prediction unit 102, an inter-frame prediction unit 103, a subtraction unit 104, a transform unit 105, a quantization unit 106, an entropy encoding unit 107, an inverse quantization unit 108, an inverse transform unit 109, an addition unit 110, a filtering unit 111, and a memory 112.

[0104] exist Figure 1The various components are illustrated separately to show the different special functions of the image encoding apparatus, but this does not mean that each component is composed of separate hardware or software units. That is, although the various components are listed and described for ease of explanation, it is possible to combine at least two components into one component, or to divide one component into multiple components and make them perform corresponding functions. The embodiments in which the various components are integrated and the embodiments in which they are separated, as described above, are included within the scope of the claims of this invention without departing from the essence of this invention.

[0105] Furthermore, some constituent elements may not be essential for performing the essential functions of this invention, but are merely optional elements used to improve performance. This invention can include only the constituent parts necessary for realizing the essence of the invention, excluding those merely used to improve performance, and structures that include only the essential constituent elements, excluding those optional elements used to improve performance, are also included within the scope of the claims of this invention.

[0106] The image segmentation unit 100 can segment an input image into at least one block. The input image can be of various shapes and sizes, such as an image, strip, parallel block, or fragment. A block can refer to a coding unit (CU), prediction unit (PU), or transform unit (TU). The segmentation can be performed based on at least one of a quadtree or a binary tree. A quadtree divides a parent block into four child blocks, each half the width and height of the parent block. A binary tree divides a parent block into two child blocks, each half the width or height of the parent block. By segmenting based on a binary tree, blocks can be segmented not only into squares but also into non-square shapes.

[0107] In the embodiments where the present invention is applied, the encoding unit can be used not only as the meaning of encoding execution unit, but also as the meaning of decoding execution unit.

[0108] Prediction units 102 and 103 can include an inter-frame prediction unit 103 for performing inter-frame prediction and an intra-frame prediction unit 102 for performing intra-frame prediction. After deciding whether to perform inter-frame or intra-frame prediction on a prediction unit, specific information (e.g., intra-frame prediction mode, motion vector, reference image, etc.) can be determined based on the different prediction methods. In this case, the processing unit for performing prediction can be different from the processing unit for determining the prediction method and specific content. For example, the prediction method and prediction mode can be determined at the prediction unit level, while prediction execution can be performed at the transformation unit level.

[0109] The residual value (residual block) between the generated prediction block and the original block can be input to the transform unit 105. Furthermore, information used during prediction, such as prediction mode information and motion vector information, can be encoded together with the residual value by the entropy coding unit 107 and then transmitted to the decoder. When using a specific coding mode, it is also possible to directly encode the original block and transmit it to the decoding unit without generating the prediction block through the prediction units 102 and 103.

[0110] The intra-prediction unit 102 can generate a prediction block based on pixel information within the current image, i.e., reference pixel information surrounding the current block. When the prediction mode of the blocks surrounding the current block performing intra-prediction is inter-prediction, reference pixels contained in the surrounding blocks to which inter-prediction has been applied can be replaced with reference pixels in other surrounding blocks to which intra-prediction has been applied. That is, if reference pixels are unavailable, they can be used after replacing the unavailable reference pixel information with at least one of the available reference pixels.

[0111] In intra-frame prediction, prediction modes can include directional prediction modes that use reference pixel information based on the prediction direction, and non-directional modes that do not use directional information when performing prediction. The mode used to predict luminance information can be different from the mode used to predict chrominance information. When predicting chrominance information, the intra-frame prediction mode information used in the process of predicting luminance information or the predicted luminance signal information can be used.

[0112] The intra-prediction unit 102 may include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is used to filter the reference pixels of the current block, and its application can adaptively determine whether to apply the filter based on the prediction mode of the current prediction unit. When the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0113] When the intra-prediction mode of the prediction unit is a prediction unit that performs intra-prediction based on pixel values ​​interpolated from reference pixels, the reference pixel interpolation unit of the intra-prediction unit 102 can generate reference pixels at fractional unit positions by interpolating the reference pixels. When the prediction mode of the current prediction unit is a prediction mode that generates prediction blocks without interpolating reference pixels, interpolation of reference pixels is not required. When the prediction mode of the current block is DC mode, the DC filter can generate prediction blocks by filtering.

[0114] The inter-frame prediction unit 103 generates prediction blocks using the reconstructed reference image stored in the memory 112 and motion information. The motion information may include, for example, motion vectors, reference image indexes, list 1 prediction flags, and list 0 prediction flags.

[0115] There are two representative methods for generating motion information in image encoding devices.

[0116] The first method is to generate motion information (motion vectors, reference image indexes, inter-frame prediction methods, etc.) using the motion inference process. Figure 2 This is a schematic diagram illustrating a method for generating motion information using motion inference based on existing technology. Motion inference is a method that uses a reference image that has been encoded and then re-decoded to generate motion information such as motion vectors, reference image indices, and inter-frame prediction directions for the current image region that needs to be encoded. In this case, motion inference can be performed on the entire reference image, or, to reduce complexity, only within the exploration range after setting the exploration range.

[0117] The second method for generating motion information is to utilize the motion information of the surrounding blocks of the current image block that needs to be encoded.

[0118] Figure 3 This is a schematic diagram illustrating an instance of a surrounding block that can be used in the process of generating motion information for the current block. Figure 3 These are the surrounding blocks that can be used in the process of generating motion information for the current block. Spatial candidate blocks A to E and temporal candidate block COL are illustrated. Spatial candidate blocks A to E exist within the same image as the current block, while temporal candidate block COL exists within an image different from the image to which the current block belongs.

[0119] The motion information of the current block can be selected from the motion information of the spatial candidate blocks A to E and the temporal candidate block COL surrounding the current block. At this time, an index can be defined to indicate which block's motion information is used as the current block's motion information. The index information described above also belongs to motion information. In the image encoding apparatus, motion information can be generated using the method described above, and a prediction block can be generated through motion compensation.

[0120] Furthermore, a residual block can be generated that contains residual information, namely the difference between the prediction unit generated in prediction units 102 and 103 and the original block of the prediction unit. The generated residual block can be input into the transformation unit 130 for transformation.

[0121] The inter-frame prediction unit 103 can derive a prediction block based on information from at least one of the previous or next images of the current image. Furthermore, it can also derive a prediction block of the current block based on information from a portion of the currently encoded region within the current image. The inter-frame prediction unit 103, according to one embodiment of the present invention, can include a reference image interpolation unit, a motion prediction unit, and a motion compensation unit.

[0122] The reference image interpolation unit can receive reference image information from memory 112 and generate pixel information in the reference image down to an integer pixel size. For luminance pixels, in order to generate pixel information down to an integer pixel size in 1 / 4 pixel units, an 8-tap DCT-based interpolation filter with different filtering coefficients can be used. For chrominance signals, in order to generate pixel information down to an integer pixel size in 1 / 8 pixel units, a 4-tap DCT-based interpolation filter with different filtering coefficients can be used.

[0123] The motion prediction unit performs motion prediction based on a reference image interpolated by the reference image interpolation unit. Various methods can be used to calculate motion vectors, such as FBMA (Full search-based Block Matching Algorithm), TSS (Three Step Search), and NTS (New Three-Step Search Algorithm). The motion vectors are calculated using interpolated pixels and have motion vector values ​​in 1 / 2 or 1 / 4 pixel units. The motion prediction unit can predict the prediction block of the current prediction unit using different motion prediction methods. Various motion prediction methods can be used, such as skip, merge, and AMVP (Advanced Motion Vector Prediction).

[0124] The subtraction unit 104 generates a residual block of the current block by performing a subtraction operation between the current block to be encoded and the prediction block generated in the intra-frame prediction unit 102 or the inter-frame prediction unit 103.

[0125] The transform unit 105 can transform the residual block containing residual data using transform methods such as DCT, DST, and KLT (Karhunen-Loeve Transform). In this case, the transform method can be determined based on the intra-prediction mode of the prediction unit used when generating the residual block. For example, DCT can be used in the horizontal direction and DST in the vertical direction, depending on the intra-prediction mode.

[0126] The quantization unit 106 is capable of quantizing the values ​​that have been transformed into frequency regions in the transformation unit 105. The quantization coefficients can be changed according to the importance of the block or image. The values ​​calculated in the quantization unit 106 can be provided to the inverse quantization unit 108 and the entropy coding unit 107.

[0127] The aforementioned transformation unit 105 and / or quantization unit 106 can be selectively included in the image encoding apparatus 100. That is, the image encoding apparatus 100 can perform at least one of transformation or quantization on the residual data of the residual block, or it can simultaneously skip transformation and quantization while encoding the residual block. Even if the image encoding apparatus 100 does not perform either transformation or quantization, or neither transformation nor quantization is performed, the block input to the entropy encoding unit 107 is generally referred to as a transformed block. The entropy encoding unit 107 performs entropy encoding on the input data. When performing entropy encoding, various encoding methods can be used, such as Exponential Golomb code, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding).

[0128] The entropy coding unit 107 can encode various types of information from the prediction units 102 and 103, such as residual coefficient information of the coding unit, block type information, prediction mode information, segmentation unit information, prediction unit information, transmission unit information, motion vector information, reference frame information, block interpolation information, and filtering information. In the entropy coding unit 107, the coefficients of the transform block can be encoded using various types of flags, such as those indicating non-zero coefficients, coefficients with absolute values ​​greater than 1 or 2, and the sign of the coefficients, based on partial block units within the transform block. For coefficients that cannot be encoded solely based on these flags, they can be encoded based on the absolute value of the difference between the coefficient encoded using the flags and the coefficients of the actual transform block. In the inverse quantization unit 108 and the inverse transform unit 109, the values ​​quantized in the quantization unit 106 are inverse quantized, and the values ​​transformed in the transform unit 105 are inversely transformed. A reconstructed block can be generated by merging the residual values ​​generated in the inverse quantization unit 108 and the inverse transform unit 109 with the prediction units predicted by the motion estimation unit, motion compensation unit, and intra-frame prediction unit 102 included in the prediction units 102 and 103. The adder 110 can generate a reconstructed block by performing an addition operation on the prediction block generated in the prediction units 102 and 103 and the residual block generated by the inverse transform unit 109.

[0129] The filtering unit 111 may include at least one of a deblocking filter, an offset correction unit, and an ALF (Adaptive Loop Filter).

[0130] Deblocking filters eliminate block distortion in reconstructed images caused by boundaries between blocks. To determine whether deblocking is necessary, the decision can be based on the pixels contained in several columns or rows of the block. When applying a deblocking filter to a block, a strong or weak filter can be used depending on the required deblocking intensity. Furthermore, during the application of a deblocking filter, horizontal and vertical filtering can be performed in parallel.

[0131] The offset correction unit can perform offset correction between the deblocked image and the original image on a pixel-by-pixel basis. To perform offset correction on a specific image, methods can be used such as dividing the pixels contained in the image into a certain number of regions, determining the regions where offset needs to be performed, and applying the offset to the corresponding regions, or applying the offset while considering the edge information of each pixel.

[0132] ALF (Adaptive Loop Filtering) can be performed based on a comparison between the filtered reconstructed image and the original image. It can also determine which filter to apply to the appropriate group after dividing the pixels in the image into specific groups, and perform different filtering on different groups. Information related to whether ALF is applied can be transmitted in coding units (CUs), and the shape and filtering coefficients of the ALF filters applied in each block can differ. Furthermore, it is possible to apply ALF filters of the same shape (fixed shape) regardless of the characteristics of the target block.

[0133] The memory 112 can store the reconstructed blocks or images calculated by the filtering unit 111, and the stored reconstructed blocks or images can be provided to the prediction units 102 and 103 when performing inter-frame prediction.

[0134] Next, an image decoding apparatus to which one embodiment of the present invention is applied will be described with reference to the accompanying drawings. Figure 4 This is a block diagram illustrating an image decoding apparatus 400 to which one embodiment of the present invention is applied.

[0135] See Figure 4 The image decoding device 400 may include an entropy decoding unit 401, an inverse quantization unit 402, an inverse transform unit 403, an addition unit 404, a filtering unit 405, a memory 406, and prediction units 407 and 408.

[0136] When the image bitstream generated by the image encoding device 100 is input to the image decoding device 400, the input bitstream can be decoded in the reverse order of the process performed by the image encoding device 100.

[0137] The entropy decoding unit 401 can perform entropy decoding according to the reverse steps of entropy encoding performed in the entropy encoding unit 107 of the image encoding apparatus 100. For example, it can be compatible with various methods such as Exponential Golomb, CAVLC (Context-Adaptive Variable Length Coding), and CABAC (Context-Adaptive Binary Arithmetic Coding), corresponding to the methods performed in the image encoder. In the entropy decoding unit 401, the coefficients of the transform block can be decoded using various types of flags, such as those indicating non-zero coefficients, coefficients with absolute values ​​greater than 1 or 2, and the sign of the coefficients, in partial block units within the transform block. For coefficients that cannot be represented solely by the aforementioned flags, decoding can be performed based on the sum of the coefficients represented by the flags and the signaling coefficients.

[0138] In the entropy decoding unit 401, information related to intra-frame prediction and inter-frame prediction performed in the encoder can be decoded. The inverse quantization unit 402 generates a transform block by performing inverse quantization on the quantized transform block. According to... Figure 1 The inverse quantization unit 108 in the middle works in essentially the same way.

[0139] The inverse quantization unit 403 generates a residual block by performing an inverse transform on the transform block. At this time, the transform method can be determined based on information related to the prediction method (inter-frame or intra-frame prediction), the block size and / or shape, and the intra-frame prediction mode. Figure 1 The inverse transformation unit 109 in the middle works in essentially the same way.

[0140] The addition unit 404 generates a reconstructed block by performing an addition operation on the prediction block generated in the intra-frame prediction unit 407 or the inter-frame prediction unit 408 and the residual block generated by the inverse transform unit 403. Figure 1 The addition unit 110 in the middle works in essentially the same way.

[0141] The filter unit 405 is used to reduce various types of noise that appear in the reconstructed blocks.

[0142] The filtering unit 405 may include a deblocking filter, an offset correction unit, and an ALF.

[0143] The image encoding device 100 can receive information on whether a deblocking filter is applied to a corresponding block or image, and if so, whether strong or weak filtering is applied. The image decoding device 400's deblocking filter can perform deblocking filtering on the corresponding block within the image decoding device 400 after receiving the deblocking filter information provided by the image encoding device 100.

[0144] The offset correction unit can perform offset correction on the reconstructed image based on the offset correction type and offset value information applicable to the image during encoding.

[0145] ALF can be applied to the encoding unit based on ALF applicability information and ALF coefficient information provided from the image encoding device 100. The ALF information described above can be included in a specific parameter set. The filtering unit 405 follows the same procedure as... Figure 1 The filtering section 111 in the middle works in essentially the same way.

[0146] The memory 406 is used to store the reconstructed blocks generated by the addition unit 404. According to... Figure 1 The memory 112 in the middle works in essentially the same way.

[0147] Prediction units 407 and 408 can generate prediction blocks based on prediction blocks provided by entropy decoding unit 401 and previously decoded blocks or image information provided by memory 406.

[0148] Prediction units 407 and 408 may include an intra-frame prediction unit 407 and an inter-frame prediction unit 408. Although not shown separately, prediction units 407 and 408 may also include a prediction unit determination unit. The prediction unit determination unit can receive various types of information input from the entropy decoding unit 401, such as prediction unit information, prediction mode information of the intra-frame prediction method, and motion prediction related information of the inter-frame prediction method, and distinguish prediction units from the current decoding unit, thereby determining whether the prediction unit performs inter-frame prediction or intra-frame prediction. The inter-frame prediction unit 408 can perform inter-frame prediction on the current prediction unit based on information contained in at least one of the previous or next images of the current image containing the current prediction unit, using the information required for inter-frame prediction of the current prediction unit provided by the image coding apparatus 100. Alternatively, it can perform inter-frame prediction based on information of a reconstructed portion of the current image containing the current prediction unit.

[0149] To perform inter-frame prediction, the motion prediction method for the prediction unit contained in the corresponding coding unit can be determined based on the coding unit as a reference, which method is Skip Mode, Merge Mode, or Advanced Motion Vector Prediction Mode (AMVP Mode).

[0150] The intra-frame prediction unit 407 generates a prediction block using pixels located around the block to be encoded in the current period and pixels that have already been reconstructed.

[0151] The intra-prediction unit 407 may include an AIS (Adaptive Intra Smoothing) filter, a reference pixel interpolation unit, and a DC filter. The AIS filter is used to filter the reference pixels of the current block, and its application can adaptively determine whether to apply the filter based on the prediction mode of the current prediction unit. AIS filtering can be performed on the reference pixels of the current block using the prediction mode of the prediction unit provided from the image coding apparatus 100 and the AIS filter information. When the prediction mode of the current block is a mode that does not perform AIS filtering, the AIS filter may not be applied.

[0152] When the prediction mode of the prediction unit is a prediction unit that performs intra-frame prediction based on pixel values ​​interpolated from reference pixels, the reference pixel interpolation unit of the intra-frame prediction unit 407 can generate a reference pixel at a fractional unit position by interpolating the reference pixel. The generated reference pixel at the fractional unit position can be used as the prediction pixel for pixels in the current block. When the prediction mode of the current prediction unit is a prediction mode that generates a prediction block without interpolating the reference pixel, interpolation of the reference pixel is not required. When the prediction mode of the current block is DC mode, the DC filter can generate a prediction block by filtering.

[0153] Intra-prediction unit 407 according to and Figure 1 The intra-prediction unit 102 in the middle works in essentially the same way.

[0154] The inter-frame prediction unit 408 generates inter-frame prediction blocks using reference images stored in memory 406 and motion information. The inter-frame prediction unit 408 then generates inter-frame prediction blocks according to... Figure 1 The inter-frame prediction unit 103 in the middle works in essentially the same way.

[0155] Next, various embodiments of the present invention will be described in more detail with reference to the accompanying drawings.

[0156] (First Embodiment)

[0157] Figure 5a as well as Figure 5bThis is a schematic diagram illustrating inter-frame prediction using reconstructed pixel regions according to a first embodiment of the present invention. The inter-frame prediction using reconstructed pixel regions applicable to this embodiment is particularly capable of deriving the motion vector of the current block using the reconstructed pixel regions.

[0158] exist Figure 5a In the diagram, the reconstructed pixel region C 52 is illustrated as the current block 51 that needs to be encoded or decoded, and the region adjacent to the current block 51. The current block 51 and the reconstructed pixel region C 52 are contained within the current image 50. The current image 50 can be an image, stripe, parallel block, coding tree block, coding block, or other image region. The reconstructed pixel region C 52, from an encoding perspective, can be a pixel region that was encoded before the current block 51 was encoded and then reconstructed; from a decoding perspective, it can be a region that was reconstructed before the current block 51 was decoded.

[0159] Before encoding or decoding the current block, since there is a reconstructed pixel region C 52 around the current block 51, the image encoding device 100 and the decoding device 400 can utilize the same reconstructed pixel region C 52. Therefore, the encoding device 100 does not need to encode the motion information of the current block 51, but instead uses the reconstructed pixel region C 52 to generate the motion information of the current block 51 and generate the prediction block in the same way by the image encoding device 100 and the decoding device 400.

[0160] exist Figure 5b The illustration shows an embodiment of motion inference and motion compensation using reconstructed pixel regions. Figure 5b Within the reference image 53 shown, the image will be compared with... Figure 5a The region matching the reconstructed pixel region C 52 illustrated is explored. After determining the reconstructed pixel region D most similar to the reconstructed pixel region C 52, the displacement between region 56, located at the same position as the reconstructed pixel region C 52, and the reconstructed pixel region D 54 is determined as the motion vector 57 of the reconstructed pixel region C 52. After selecting the motion vector 57 determined as described above as the motion vector of the current block 51, the predicted block 58 of the current block 51 can be derived using the aforementioned motion vector 57.

[0161] Figure 6 This is a sequence diagram for illustrating the inter-frame prediction method to which the first embodiment of the present invention applies.

[0162] Inter-frame prediction applicable to this embodiment can be performed by the inter-frame prediction unit 103 of the image encoding apparatus 100 or the inter-frame prediction unit 408 of the image decoding apparatus 400, respectively. The reference image used during the inter-frame prediction process is stored in the memory 112 of the image encoding apparatus 100 or the memory 406 of the image decoding apparatus 400. The inter-frame prediction unit 103 or the inter-frame prediction unit 408 can generate a prediction block for the current block 51 by referring to the reference image stored in the memory 112 or the memory 406.

[0163] See Figure 6 In step S61, the reconstructed pixel region 52 is first selected to be used in the motion vector derivation process of the current block that needs to be encoded or decoded. Next, in step S63, the motion vector of the reconstructed pixel region 52 is derived based on the reconstructed pixel region 52 and the reference image of the current block. (As in...) Figure 5b The explanations given above will be... Figure 5b Within the reference image 53 shown in the figure, the region that matches the reconstructed pixel region C 52 is explored. After determining the reconstructed pixel region D 54 that is most similar to the reconstructed pixel region C 52, the displacement between the region 56 located at the same position as the reconstructed pixel region C 52 and the reconstructed pixel region D 54 is determined as the motion vector 57 of the reconstructed pixel region C 52.

[0164] In step S65, the image encoding device 100 or the image decoding device 400 selects the motion vector 57 of the reconstructed pixel region C 52 determined as described above as the motion vector of the current block 51. Using the motion vector 57, a prediction block 58 of the current block 51 can be generated.

[0165] Furthermore, the reconstructed pixel region C 52 can have different shapes and / or sizes. Figures 7a to 7c This is a schematic diagram illustrating an embodiment of the reconstructed pixel region. Figures 7a to 7c In the diagram, M, N, O, and P represent pixel intervals, respectively. O and P can also be negative if the absolute values ​​of O and P are less than the horizontal or vertical length of the current block.

[0166] Furthermore, the reconstructed pixel regions on the upper and left sides of the current block can be used as reconstructed pixel regions C 52 respectively, or the two regions can be merged and used as a single reconstructed pixel region C 52. Additionally, the reconstructed pixel region C 52 can be used after secondary sampling. Because the motion information is derived using only the decoded information surrounding the current block as described above, it is not necessary to transmit motion information from the encoding device 100 to the decoding device 400.

[0167] In embodiments to which the present invention is applied, since the decoding device 400 also performs motion inference, performing motion inference over the entire reference image can lead to an extreme increase in its complexity. Therefore, the computational complexity of the decoding device 400 can be reduced by transmitting the exploration range in block units or upper-level headers or by fixing the same exploration range in both the encoding device 100 and the decoding device 400.

[0168] Figure 8 This is a sequence diagram illustrating the decision process of an inter-frame prediction method applicable to one embodiment of the present invention. In inter-frame prediction using reconstructed pixel regions applicable to the present invention, and in existing inter-frame prediction methods, the optimal method can be determined through rate-distortion optimization (RDO). Figure 8 The process illustrated can be executed by the image encoding device 100.

[0169] See Figure 9 In step S81, cost_A is first calculated by performing inter-frame prediction in the existing manner, and in step S82, cost_B is calculated by performing inter-frame prediction using the reconstructed pixel region applicable to the present invention in the manner described above.

[0170] Next, in step S83, cost_A and cost_B are compared to determine which method is more advantageous. If cost_A is lower, inter-frame prediction using the existing method is performed in step S84; otherwise, inter-frame prediction using the reconstructed pixel region is performed in step S85.

[0171] Figure 9 For indicating passage Figure 8 The diagram illustrates the encoding process of information related to the inter-frame prediction method, as determined by the illustrated process. Next, we will proceed according to... Figure 8 The information used to indicate the inter-frame prediction method determined by the illustrated process is called Decoder-side Motion Vector Derivation (DMVD) indication information. DMVD indication information can be used to indicate whether to perform inter-frame prediction using the existing method or to perform inter-frame prediction using reconstructed pixel regions.

[0172] See Figure 9 In step S901, the method used to indicate passage will be... Figure 8The process illustrated involves encoding the decoder-side motion vector derivation (DMVD) indication information for the inter-frame prediction method. The DMVD indication information can be a 1-bit flag or one of multiple indices. Next, in step S902, the motion information is encoded, and the corresponding algorithm terminates.

[0173] Alternatively, information indicating whether or not inter-frame prediction using reconstructed pixel regions is used in the embodiments of the present invention can be generated first in the upper-level header and then encoded. That is, when the information indicating whether or not inter-frame prediction using reconstructed pixel regions is used is true, the decoder-side motion vector derivation (DMVD) indication information is encoded. When the information indicating whether or not inter-frame prediction using reconstructed pixel regions is used is false, the bitstream does not contain decoder-side motion vector derivation (DMVD) indication information, and the current block will be predicted using existing inter-frame prediction.

[0174] In addition, information in the higher-level header indicating whether or not inter-frame prediction using reconstructed pixel regions is used can be included in block headers, strip headers, parallel block headers, image headers, or sequence headers for transmission.

[0175] Figure 10 It is used to follow such Figure 9 The diagram illustrates the decoding process of decoder-side motion vector derivation (DMVD) indication information encoded in the manner shown.

[0176] In step S1001, the decoding device 400 decodes the decoder-side motion vector derivation (DMVD) indication information. Next, in step S1002, the motion information is decoded, and then the corresponding algorithm ends.

[0177] If the upper-level header of the bitstream contains information indicating whether or not inter-frame prediction using reconstructed pixel regions is used, and this information is true, the bitstream will contain decoder-side motion vector derivation (DMVD) indication information. Conversely, if this information is false, the bitstream will not contain DMVD indication information, and existing inter-frame prediction will be used to predict the current block.

[0178] A higher-level header containing information indicating whether or not inter-frame prediction using reconstructed pixel regions is used can be included in block headers, strip headers, parallel block headers, image headers, or sequence headers for transmission.

[0179] (Second Embodiment)

[0180] Next, a second embodiment to which the present invention is applied will be described with reference to the accompanying drawings.

[0181] In the second embodiment, the inter-frame prediction using reconstructed pixel regions as described in the first embodiment is applied to inter-frame prediction using affine transformation. Specifically, in order to derive the motion vectors of the control points used in the inter-frame prediction process using affine transformation, the motion vector derivation method using reconstructed pixel regions as described above is applied. For ease of explanation, the inter-frame prediction using affine transformation as described in the second embodiment of the present invention will be simply referred to as affine inter-frame prediction.

[0182] Figure 11 This is a schematic diagram used to illustrate affine inter-frame prediction.

[0183] Affine inter-frame prediction generates a prediction block by obtaining motion vectors from the four corners of the current block that needs to be encoded or decoded. The four corners of the current block can be considered as control points.

[0184] See Figure 11 The block identified by motion vectors 11-2, 11-3, 11-4 and 11-5 at the four corners (i.e., control points) of the current block (not shown) in the current image can be the predicted block 11-6 of the current block.

[0185] By using affine inter-frame prediction as described above, it is possible to predict blocks or image regions that are rotated, enlarged / shrunk, moved, reflected, and sheared.

[0186] Formula 1 below is the general determinant of an affine transformation.

[0187] [Formula 1]

[0188]

[0189] Formula 1 is a formula for two-dimensional coordinate transformation, where (x, y) are the original coordinates, (x', y') are the target coordinates, and a, b, c, d, e, and f are transformation parameters.

[0190] To apply the affine transformation described above to a video codec, the transformation parameters need to be transmitted to the image decoding device, which leads to a significant increase in overhead. Therefore, existing video codecs simply apply the affine transformation using the surrounding reconstructed N control points.

[0191] Formula 2 below is a method for deriving the motion vector of any sub-block within the current block using two control points at the top left and top right edges of the current block.

[0192] [Formula 2]

[0193]

[0194] In Formula 2, (x, y) represents the position of any sub-block within the current block, W represents the horizontal length of the current block, and (MV x MV y ) is the motion vector of the sub-block, (MV 0x MV 0y ) is the motion vector of the upper left control point, while (MV) 1x MV 1y ) is the motion vector of the upper right control point.

[0195] Next, Formula 3 below is a method for deriving the motion vector of any sub-block within the current block using three control points at the top left, top right, and bottom left edges of the current block.

[0196] [Formula 3]

[0197]

[0198] In Formula 3, (x, y) represents the position of any sub-block, W and H are the horizontal and vertical lengths of the current block, respectively, (MV x MV y ) represents the motion vector of the sub-blocks within the current block, (MV 0x MV 0y ) is the motion vector of the upper left control point, (MV 1x MV 1y ) is the motion vector of the upper right control point, while (MV) 2x MV 2y ) is the motion vector of the lower left control point.

[0199] In the second embodiment to which the present invention is applied, in order to derive the motion vectors of the control points used in the affine inter-frame prediction process, the motion vector derivation method of the first embodiment described above using reconstructed pixel regions is applied. Therefore, it is not necessary for the image encoding device 100 to transmit the motion vector information of multiple control points to the image decoding device 400.

[0200] Figure 12a as well as Figure 12bThis is a schematic diagram illustrating the derivation of motion vectors for control keys using reconstructed pixel regions according to the second embodiment of the present invention.

[0201] See Figure 12a The current block 12a-2, which requires encoding or decoding, is contained within the current image 12a-1. The four control points used to perform affine inter-frame prediction are circularly marked at the four corners of the current block 12a-2. Figure 12a In addition, in the above figures, the reconstructed pixel regions a 12a-3, b 12a-4, and c 12a-5 are illustrated as areas adjacent to the three control points at the upper left, upper right, and lower left.

[0202] In this embodiment, the reconstructed pixel regions a 12a-3, b 12a-4, and c 12a-5 are used according to... Figure 12b The motion vectors of the three control points at the top left, top right, and bottom left are derived as shown in the diagram. However, for the control point at the bottom right of the current block, there may be no reconstructed pixel area around it. In the case described above, the motion vector of the sub-block d12a-6, which has an arbitrary size, obtained through existing inter-frame prediction methods, can be used as the motion vector of the bottom right control point of the current block.

[0203] See Figure 12b ,exist Figure 12b Within the reference image 12b-1 shown, the image will be compared with... Figure 12a The regions d12b-10, e12b-11, and f12b-12 that match the reconstructed pixel regions a12a-3, b12a-4, and c12a-5, respectively, are explored. The displacements between regions 12b-6, 12b-7, and 12b-8, which are located at the same positions as the reconstructed pixel regions a12a-3, b12a-4, and c12a-5, are determined as motion vectors 12b-2, 12b-3, and 12b-4 of the reconstructed pixel regions a12a-3, b12a-4, and c12a-5, respectively. The motion vectors 12b-2, 12b-3, and 12b-4, determined as described above, are selected as the motion vectors for the three control points at the upper left, upper right, and lower left ends of the current block 12a-2. Furthermore, the motion vector for the lower right control point can be obtained using the motion vector of sub-block d12a-6 obtained using existing inter-frame prediction methods.

[0204] According to Formula 4 below, the motion vector of any sub-block within the current block can be derived using the motion vectors of the four control points derived as described above.

[0205] [Formula 4]

[0206]

[0207] In Formula 4, (x, y) represents the position of any sub-block within the current block, and W and H are the horizontal and vertical lengths of the current block, respectively. x MV y ) represents the motion vector of the sub-blocks within the current block, (MV 0x MV 0y ) is the motion vector of the upper left control point, (MV 1x MV 1y ) is the motion vector of the upper right control point, (MV 2x MV 2y ) is the motion vector of the lower left control point, (MV 3x MV 3y ) is the motion vector of the lower right control point.

[0208] Furthermore, the size and / or shape of the reconstructed pixel regions a 12a-3, b 12a-4, and c 12a-5 can have a reference. Figures 7a to 7c The various shapes and / or sizes described above. The size and / or shape of sub-block d 12a-6 can also use the same size and / or shape preset in the encoding device 100 and the decoding device 400. Furthermore, the horizontal and / or vertical information of sub-block d 12a-6 can be transmitted via block units or upper-level headers, or the size information can be transmitted in powers of 2 units.

[0209] As described above, after deriving the motion vectors at the four control points, the motion vector of the current block 12a-2 or the motion vector of any sub-block within the current block 12a-2 can be derived using these vectors. Furthermore, using the derived motion vectors, the prediction block of the current block 12a-2 or the prediction block of any sub-block within the current block 12a-2 can be derived. Specifically, referring to Formula 4 above, since the position of the current block 12a-2 is (0, 0), the motion vector of the current block 12a-2 will be the motion vector (MV) of the upper left control point. 0x MV 0y Therefore, the predicted block of the current block 12a-2 can be obtained using the motion vector of the upper left control point mentioned above. If the current block is an 8×8 block and is divided into 4 4×4 sub-blocks, the motion vector of the sub-block at position (3,0) in the current block can be obtained by substituting 3 for x, 0 for y, and 9 for W and H in the variables of Formula 4 above.

[0210] Next, we will combine Figure 13 The inter-frame prediction method applicable to the second embodiment of the present invention will be described. Figure 13 This is a sequence diagram for illustrating the inter-frame prediction method applicable to the second embodiment of the present invention.

[0211] Inter-frame prediction applicable to this embodiment can be performed by the inter-frame prediction unit 103 of the image encoding apparatus 100 or the inter-frame prediction unit 408 of the image decoding apparatus 400, respectively. The reference image used during the inter-frame prediction process is stored in the memory 112 of the image encoding apparatus 100 or the memory 406 of the image decoding apparatus 400. The inter-frame prediction unit 103 or the inter-frame prediction unit 408 can generate a prediction block for the current block 51 by referring to the reference image stored in the memory 112 or the memory 406.

[0212] See Figure 13 In step S131, at least one reconstructed pixel region is first selected to be used in the motion vector derivation process of at least one control point of the current block that needs to be encoded or decoded. Figure 12a as well as Figure 12b In the illustrated embodiment, three reconstructed pixel regions, a 12a-3, b 12a-4, and c 12a-5, were selected to derive the motion vectors of the three control points at the upper left, upper right, and lower left ends of the current block 12a-2. However, this is not a limitation; one or two reconstructed pixel regions can also be selected to derive the motion vectors of one or two of the three control points.

[0213] Next, in step S133, motion vectors for at least one reconstructed pixel region are derived based on the at least one reconstructed pixel region selected in step S131 and the reference image of the current block. In step S135, the image encoding device 100 or the image decoding device 400 selects each motion vector of the reconstructed pixel region C 52 determined as described above as the motion vector of at least one control point of the current block. Using the at least one motion vector selected as described above, a predicted block of the current block can be generated.

[0214] Figure 14 This is a sequence diagram illustrating the decision process of the inter-frame prediction method applicable to the second embodiment of the present invention. In affine inter-frame prediction applicable to the second embodiment of the present invention and in existing inter-frame prediction methods, the optimal method can be determined through rate-distortion optimization (RDO). Figure 14 The process illustrated can be executed by the image encoding device 100.

[0215] See Figure 14In step S14, cost_A is calculated by performing inter-frame prediction in the conventional manner, and in step S142, cost_B is calculated by performing affine inter-frame prediction applicable to the second embodiment of the present invention in the manner described above.

[0216] Next, in step S143, cost_A and cost_B are compared to determine which method is more advantageous. When cost_A is lower, inter-frame prediction using the existing method is performed in step S144; otherwise, affine inter-frame prediction according to the second embodiment of the present invention is performed in step S145.

[0217] Figure 15 For indicating passage Figure 14 The diagram illustrates the encoding process of information related to the inter-frame prediction method, as determined by the process shown. Next, we will proceed according to... Figure 14 The information used to indicate the inter-frame prediction method determined by the illustrated process is called Decoder-side Control Point Motion Vector Derivation (DCMVD) indication information. The DCMVD indication information can be used to indicate whether to perform inter-frame prediction using existing methods or to perform affine inter-frame prediction applicable to the second embodiment of this invention.

[0218] See Figure 15 In step S151, the method used to indicate passage will be... Figure 14 The process illustrated involves encoding the decoder-side control point motion vector derivation (DCMVD) indication information for the inter-frame prediction method. The DCMVD indication information can be a 1-bit flag or one of multiple indices. Next, in step S152, the motion information is encoded, and then the corresponding algorithm terminates.

[0219] Furthermore, information indicating whether or not affine inter-frame prediction is used according to the second embodiment of the present invention can be generated first in the upper-level header before encoding it. That is, when the information indicating whether or not affine inter-frame prediction is used according to the second embodiment of the present invention is true, the decoder-side control point motion vector derivation (DCMVD) indication information is encoded. When the information indicating whether or not affine inter-frame prediction is used according to the second embodiment of the present invention is false, the bitstream does not contain decoder-side control point motion vector derivation (DCMVD) indication information, and the current block will be predicted using existing inter-frame prediction.

[0220] Furthermore, information in the upper-level header indicating whether or not affine inter-frame prediction is used in accordance with the present invention can be included in the block header, strip header, parallel block header, image header, or sequence header for transmission.

[0221] Figure 16 It is used to follow such Figure 15 The diagram illustrates the decoding process of decoder-side control point motion vector derivation (DCMVD) indication information encoded in the manner shown.

[0222] In step S161, the decoding device 400 decodes the decoder-side control point motion vector derivation (DCMVD) indication information. Next, in step S162, the motion information is decoded, and then the corresponding algorithm ends.

[0223] If the upper-level header of the bitstream contains information indicating whether or not affine inter-frame prediction is used according to the second embodiment of the present invention, and the information indicating whether or not inter-frame prediction using reconstructed pixel regions is true, the bitstream will contain decoder-side control point motion vector derivation (DCMVD) indication information. When the information indicating whether or not affine inter-frame prediction is used according to the second embodiment of the present invention is false, the bitstream will not contain decoder-side control point motion vector derivation (DCMVD) indication information, and in this case, existing inter-frame prediction will be used to predict the current block.

[0224] Information in the upper-level header indicating whether or not affine inter-frame prediction is used in accordance with the second embodiment of the present invention can be included in the block header, strip header, parallel block header, image header, or sequence header for transmission.

[0225] Figure 17 This is a sequence diagram illustrating an embodiment of an image decoding method that uses reconstructed pixel regions to derive motion vectors of three control points and generate a predicted block for the current block. Figure 17 The process shown in the diagram is the same as... Figure 12a as well as Figure 12b The illustrated embodiments are relevant.

[0226] To derive the motion vectors of the three control points at the top left, top right, and bottom left of the current block 12a-2, three reconstructed pixel regions, a 12a-3, b 12a-4, and c 12a-5, were selected. However, this is not a limitation; one or two reconstructed pixel regions can also be selected to derive the motion vectors of one or two of the three control points.

[0227] The image decoding device 400 can determine which type of inter-frame prediction to perform based on decoder-side control point motion vector derivation (DCMVD) indication information. In step S171, when the decoder-side control point motion vector derivation (DCMVD) indication information indicates the use of affine inter-frame prediction applicable to the present invention, in step S172, the motion vectors of the control points at the upper left, upper right, and lower left ends of the current block are inferred and selected respectively using the reconstructed pixel region.

[0228] Next, in step S173, the motion vector obtained by decoding the motion information transmitted in the bitstream is set as the motion vector of the lower right control point. In step S174, the inter-frame prediction block of the current block is generated by using the affine transformation of the motion vectors of the four control points derived in steps S172 and S173. When affine inter-frame prediction is not used, in step S175, the motion information is decoded according to the existing inter-frame prediction, and the prediction block of the current block is generated using the decoded motion information.

[0229] (Third Embodiment)

[0230] Figure 18 This is a schematic diagram illustrating a current block that is divided into multiple regions for inter-frame prediction according to the third embodiment of the present invention.

[0231] exist Figure 18 In the diagram, the reconstructed pixel region C 503 is illustrated, representing the current block 500 that needs to be encoded or decoded, and the regions adjacent to the current block 500. The current block 500 is divided into region A 500-a and region B 500-b.

[0232] Based on the correlation between pixels, pixels in the reconstructed pixel region C 503 are highly likely to be similar to pixels contained in region A 500-a, but may be dissimilar to pixels contained in region B 500-b. Therefore, motion inference and motion compensation using the reconstructed pixel region C 503 are performed during the inter-frame prediction process in region A 500-a to find the correct motion while preventing increased overhead. The inter-frame prediction method for region B 500-b is applicable to existing inter-frame prediction methods.

[0233] Figure 19 This is a sequence diagram for illustrating the inter-frame prediction method to which the third embodiment of the present invention applies.

[0234] Inter-frame prediction applicable to this embodiment can be performed by the inter-frame prediction unit 103 of the image encoding apparatus 100 or the inter-frame prediction unit 408 of the image decoding apparatus 400, respectively. The reference image used during the inter-frame prediction process is stored in the memory 112 of the image encoding apparatus 100 or the memory 406 of the image decoding apparatus 400. The inter-frame prediction unit 103 or the inter-frame prediction unit 408 can generate prediction blocks for region A 500-a and region B 500-b within the current block by referring to the reference images stored in the memory 112 or the memory 406, respectively.

[0235] First, such as Figure 18 As shown, in step S51, the current block that needs to be encoded or decoded will be divided into multiple regions including a first region and a second region. The first region and the second region can correspond to region A 500-a and region B 500-b as illustrated in Figure 5, respectively. Figure 18 The current block 500 shown in the figure is divided into two regions, A region 500-a and B region 500b, but it can also be divided into three or more regions, and can also be divided into regions of different sizes and / or shapes.

[0236] Next, in step S53, prediction blocks for the first region and the second region are obtained using different inter-frame prediction methods. The inter-frame prediction method for region A 500-a can be the method described above, which uses the reconstructed pixel region C 503 to perform motion inference and motion compensation. The inter-frame prediction method for region B 500-b can be any existing inter-frame prediction method.

[0237] The method described in this embodiment, which divides the current block into multiple regions and uses different inter-frame prediction methods to derive the predicted blocks for each region, is called mixed inter-prediction.

[0238] Figure 20 This is a schematic diagram illustrating an embodiment of motion inference and motion compensation using reconstructed pixel regions.

[0239] See Figure 20 Within reference image 600, the comparison with Figure 18 The reconstructed pixel region C 503 shown in the figure is explored. For example... Figure 20 As shown, after determining the reconstructed pixel region D 603 that is most similar to the reconstructed pixel region C 503, the displacement between region 601, which is located at the same position as the reconstructed pixel region C 503, and the reconstructed pixel region D 603 is determined as the motion vector 605 of region A 500-a.

[0240] That is, the motion vector 605 inferred from the reconstructed pixel region C 503 is selected as the motion vector of region A 500-a in the current block. Then, the predicted block of region A 500-a is generated using the motion vector 605.

[0241] In addition, such as Figures 7a to 7c As shown, the reconstructed pixel region C 503 can have different shapes and / or sizes. Furthermore, the reconstructed pixel regions on the upper and left sides can be used separately. Additionally, they can be used after resampling the reconstructed pixel region. Because the motion information is derived using only the decoded information surrounding the current block as described above, it is not necessary to transmit motion information from the encoding device 100 to the decoding device 400.

[0242] In embodiments to which the present invention is applied, since the decoding device 400 also performs motion inference, performing motion inference over the entire reference image can lead to an extreme increase in its complexity. Therefore, the computational complexity of the decoding device 400 can be reduced by transmitting the exploration range in block units or upper-level headers or by fixing the same exploration range in both the encoding device 100 and the decoding device 400.

[0243] In addition, in the context of Figure 18 When speculating and encoding the motion vector of region B 500-b as shown in the figure, the motion vector of region B 500-b can be predicted using the motion information of the decoded block in the reconstructed pixel region C 503, and then the residual vector corresponding to the difference between the motion vector of region B 500-b and its predicted motion vector can be encoded.

[0244] Alternatively, the residual vector can be encoded after predicting the motion vector of region B500-b using motion vector 605 derived from the motion vector of region A500-a.

[0245] Alternatively, the motion vector of region B 500-b can be predicted and the residual vector encoded after the motion vector of the decoded block in the reconstructed pixel region C 503 and the inferred motion vector 605 of region A500-a are used to form a motion vector prediction set.

[0246] Alternatively, block merging can be performed after obtaining motion information from pre-defined positions in blocks adjacent to the current block. Block merging refers to directly applying surrounding motion information to the block that needs to be encoded. In this case, multiple pre-defined positions can be set, and an index can be used to indicate which position to perform the block merging at.

[0247] Furthermore, the size of region B 500-b can be encoded in the encoding device 100 and then transmitted to the decoder via block units or upper headers, or the same value or ratio preset in the encoding device 100 and the decoding device 400 can be used.

[0248] Figure 21 This is a sequence diagram illustrating the decision process of an inter-frame prediction method applicable to one embodiment of the present invention. In the hybrid inter-frame prediction method applicable to the present invention and in existing inter-frame prediction methods, the optimal method can be determined through rate-distortion optimization (RDO). Figure 21 The process illustrated can be executed by the image encoding device 100.

[0249] See Figure 21 In step S801, cost_A is calculated by performing inter-frame prediction in the conventional manner, and in step S802, cost_B is calculated by performing inter-frame prediction in a hybrid manner applicable to the present invention, which divides the current block into at least two regions and performs inter-frame prediction separately, as described above.

[0250] Next, in step S803, cost_A and cost_B are calculated to determine which method is more advantageous. When cost_A is low, inter-frame prediction using the existing method is performed in step S804; otherwise, inter-frame prediction using a hybrid method is performed in step S805.

[0251] Figure 22 It is for indicating passage Figure 21 The diagram illustrates the process by which information on the inter-frame prediction method, determined by the process shown, is transmitted to the decoding device 400.

[0252] In step S901, information indicating which inter-frame prediction method was used for the block currently requiring encoding is encoded. This information can be a 1-bit flag or one of multiple indices. Next, in step S902, motion information is encoded, and then the corresponding algorithm ends.

[0253] Alternatively, information indicating whether or not inter-frame prediction using the hybrid method applicable to embodiments of the present invention is used can be generated first in the upper-level header before encoding. That is, when the information in the upper-level header indicating whether or not inter-frame prediction using the hybrid method is used is true, the information indicating which method of inter-frame prediction is used in the current block to be encoded will be encoded. When the information indicating whether or not inter-frame prediction using the hybrid method is used is false, the bitstream does not contain information indicating which method of inter-frame prediction is used. In this case, it is not necessary to divide the current block into multiple regions; instead, existing inter-frame prediction is used to predict the current block.

[0254] In addition, information in the upper-level header indicating whether inter-frame prediction is used in the blending mode can be included in the block header, strip header, parallel block header, image header, or sequence header for transmission.

[0255] Figure 23 It is through Figure 21 The diagram illustrates the decoding process of information encoded in a specific manner, which indicates which method of inter-frame prediction was used to indicate the block currently being encoded.

[0256] In step S1001, the decoding device 400 decodes the information indicating which type of inter-frame prediction was used for the block that needs to be encoded. Next, in step S1002, the motion information is decoded, and then the corresponding algorithm ends.

[0257] If the upper-level header of the bitstream contains information indicating whether inter-frame prediction is used in a hybrid mode, and this information is true, the bitstream will contain information indicating which inter-frame prediction mode is used for the block currently being encoded. Conversely, if the information indicating whether inter-frame prediction is used is false, the bitstream will not contain information indicating which inter-frame prediction mode is used. In this case, it is not necessary to divide the current block into multiple regions; instead, existing inter-frame prediction is used to predict the current block.

[0258] The upper-level header containing information indicating whether inter-frame prediction is used in the blending mode can be a block header, stripe header, parallel block header, image header, or sequence header. The information in the upper-level header indicating whether inter-frame prediction is used in the blending mode can be included in the block header, stripe header, parallel block header, image header, or sequence header for transmission.

[0259] Figure 24This is a schematic diagram illustrating the sequence of methods for generating inter-frame prediction blocks using information indicating the inter-frame prediction method used. Figure 24 The method illustrated can be executed by the decoding device 400.

[0260] First, in step S1101, it is determined whether the information indicating which method of inter-frame prediction was used indicates the use of a hybrid inter-frame prediction method. Next, if it is determined that a hybrid inter-frame prediction method is used for the current block that needs to be decoded, in step S1102, the current block is divided into multiple regions. For example, the current block can be divided into region A 500-a and region B 500-b as shown in Figure 5.

[0261] At this time, the size of each segmented region can be signaled from the encoding device 100 to the decoding device 400 through block units or upper-level headers, or it can be set to a preset value.

[0262] Next, in step S1103, according to Figure 20 The method illustrated in the figure infers the motion vector of the first region, such as region A 00-a, and generates a prediction block.

[0263] Next, in step S1104, the predicted block for the second region, such as region B 500-b, will be generated using the decoded motion vectors, and then the algorithm will end.

[0264] If the information indicating which inter-frame prediction method was used indicates that a mixed inter-frame prediction method is not used, or if the information in the upper-level header indicating whether a mixed inter-frame prediction method is used is false, the existing inter-frame prediction method will be applied as the prediction method for the current block 500. That is, in step S1105, a prediction block for the current block 500 will be generated using the decoded motion information, and then the algorithm will end. The size of the prediction block is the same as the size of the current block 500 that needs to be decoded.

[0265] (Example 4)

[0266] Next, a fourth embodiment to which the present invention is applicable will be described with reference to the accompanying drawings. The fourth embodiment relates to a method for reducing blocking artifacts that may occur at block boundaries when performing inter-frame prediction in a hybrid manner applicable to the third embodiment as described above.

[0267] Figure 25 This is a schematic diagram illustrating a method for reducing blocking artifacts that may occur when performing inter-frame prediction using a hybrid approach to which the present invention is applicable. Figure 25The predicted block 1 and predicted block 2 shown in the figure can respectively correspond to Figure 18 The diagram shows the prediction blocks for region A 500-a and region B 500-b.

[0268] To summarize the fourth embodiment to which the present invention applies, firstly, the region located at the boundary of the prediction block is divided into sub-blocks of a specific size. Next, a new prediction block is generated by applying motion information from the sub-blocks surrounding the existing sub-blocks of the prediction block to the existing sub-blocks. Then, a final set of sub-blocks of the prediction block is generated by weighted summing of the new prediction block and the existing sub-blocks of the prediction block. This can be referred to as Overlapped Block Motion Compensation (OBMC).

[0269] See Figure 25 For those existing Figure 25 Sub-block P2 in prediction block 1 can generate a new prediction block P2 by applying the motion information of the surrounding sub-block A2 to P2, and then proceed according to... Figure 26 The illustrated method uses weighted summation to generate the final sub-predicted blocks.

[0270] For the sake of clarity, assume Figure 25 The size of each sub-block shown in the figure is 4×4. It is assumed that there are 8 sub-blocks A1 to A8 adjacent to the top of the predicted block 1 and 8 sub-blocks B1 to B8 adjacent to the left. It is also assumed that there are 4 sub-blocks C1 to C4 adjacent to the top of the predicted block 2 and 4 sub-blocks D1 to D8 adjacent to the left.

[0271] Although for the sake of explanation, it is assumed that the horizontal and vertical lengths of each sub-block are 4, it is also possible to signal the various values ​​to the decoding device 400 after encoding them through block units or higher-level headers. Therefore, the size of the sub-blocks can be set to the same value in both the encoding device 100 and the decoding device 400. Alternatively, the encoding device 100 and the decoding device 400 can also use sub-blocks of the same preset size.

[0272] Figure 26 This is a schematic diagram illustrating the applicable method for weighted summation of sub-blocks within a prediction block and the sub-blocks adjacent to the top.

[0273] See Figure 26 The final predicted pixel c will be generated using the formula described below.

[0274] [Formula 5]

[0275] c = W1 × a + (1 - W1) × b

[0276] Besides predicting pixel c, the remaining 15 pixels can also be calculated using a similar method as described above. Figure 13 P2 to P8 in the middle will be replaced by passing Figure 26 The process illustrated applies to the new predicted pixels using a weighted summation.

[0277] Figure 27 This is a diagram illustrating the applicable method for weighted summation of sub-blocks within a prediction block and the sub-blocks adjacent to its left. Sub-blocks P9 to P15 will be calculated as follows: Figure 27 The method shown in the figure applies a weighted summation to replace the new predicted pixel value.

[0278] See Figure 26 The same weighting value will be applied to pixels on a row-by-row basis. See [link / reference] Figure 27 The same weighting will be applied to pixels on a column-by-column basis.

[0279] for Figure 25 Sub-block P1 can also be accessed by following... Figure 26 The method shown in the diagram is to perform a weighted summation of the pixel values ​​with those in the surrounding sub-block A1, and then... Figure 27 The method illustrated is used to perform a weighted summation of the pixel values ​​within sub-block B1 to obtain the final predicted value.

[0280] For those that exist Figure 25 The sub-blocks P16 to P22 in the predicted block 2 shown in the figure can also be obtained by using... Figure 26 or Figure 27 The weighted summation method shown in the diagram is used to obtain the final predicted value. At this time, the surrounding sub-blocks used in the weighted summation process can be C1~C4 to D1~D4.

[0281] Furthermore, a weighted summation method can be used to replace the pixel values ​​of surrounding sub-blocks C1 to C4 and D1 to D4 with new values, rather than simply replacing the pixel values ​​of sub-blocks P16 to P22. Taking sub-block C2 as an example, after generating a predicted sub-block by applying the motion information of sub-block P17 to sub-block C2, a weighted summation can be performed on the pixel values ​​within the predicted sub-block and the pixel values ​​of sub-block C2 to generate the pixel values ​​of C2 subject to weighted summation.

[0282] Figure 28 This is a sequence diagram illustrating the process of determining whether a weighted summation is applicable between sub-blocks at block boundaries when performing inter-frame prediction using the hybrid method applicable to this invention.

[0283] In step S1501, the variable BEST_COST, used to store the optimal cost, is initialized to its maximum value; COMBINE_MODE, used to store whether inter-frame prediction using the hybrid method is used, is initialized to false; and WEIGHTED_SUM, used to store whether weighted summation between sub-blocks is used, is initialized to false. Next, in step S1502, cost_A is calculated after performing inter-frame prediction using the existing method. In step S1503, cost_B is calculated after performing inter-frame prediction using the hybrid method. In step S1504, the two costs are compared. When the value of cost_A is smaller, in step S1505, by setting COMBINE_MODE to false, it indicates that inter-frame prediction using the hybrid method is not used, and cost_A is stored in BEST_COST.

[0284] In step S1504, the two costs are compared. When the value of cost_A is smaller, in step S1505, by setting COMBINE_MODE to false, it indicates that inter-frame prediction using the hybrid method is not used, and cost_A is stored in BEST_COST. Next, in step S1507, a weighted summation is applied between sub-blocks, and cost_C is calculated. In step S1508, BEST_COST is compared with cost_C. When BEST_COST is less than cost_C, in step S1509, by setting the WEIGHTED_SUM variable to false, it indicates that a weighted summation is not applied between sub-blocks; otherwise, in step S1510, by setting the WEIGHTED_SUM variable to true, it indicates that a weighted summation is applied between sub-blocks, and then the corresponding algorithm ends.

[0285] Figure 29 It is through Figure 27 The diagram illustrates the sequence of the encoding process for the information determined by the method shown, which is used to indicate whether the weighted summation between sub-blocks is applicable. Figure 29 The illustrated process can be executed by the image encoding device 100. First, in step S1601, the encoding device 100 encodes information indicating which method of inter-frame prediction was used. Next, in step S1602, motion information is encoded. Then, in step S1603, information indicating whether weighted summation between sub-blocks is applicable is encoded.

[0286] If the upper-level header of the bitstream contains information indicating whether inter-frame prediction is used in the blending mode, and this information is true, it will be included in the bitstream after encoding the information indicating whether weighted summation is applicable between the aforementioned sub-blocks. However, if the information in the upper-level header indicating whether inter-frame prediction is used in the blending mode is false, the bitstream will not contain information indicating whether weighted summation is applicable between sub-blocks.

[0287] Figure 30 This is a sequence diagram illustrating the decoding process of information used to indicate whether a weighted summation between sub-blocks is applicable. Figure 30 The illustrated process can be executed by the image decoding device 400. First, in step S1701, the decoding device 400 decodes the information indicating which method of inter-frame prediction was used. Next, in step S1702, the motion information is decoded. Next, in step S1703, the information indicating whether the weighted summation between sub-blocks is applicable is decoded.

[0288] If the upper-level header of the bitstream contains information indicating whether inter-frame prediction is used in the mixing mode and the information indicating whether inter-frame prediction is used in the mixing mode is true, it can be included in the bitstream after encoding the information indicating whether weighted summation is applicable between the aforementioned sub-blocks.

[0289] However, if the information in the upper-level packet header indicating whether inter-frame prediction is used to indicate the mixing method is false, the bitstream may not contain information indicating whether weighted summation is applicable between sub-blocks. In the case described above, it can be inferred that the information indicating whether weighted summation is applicable between sub-blocks means that weighted summation is not applicable between sub-blocks.

[0290] (5th embodiment)

[0291] Figure 31a as well as Figure 31b This is a schematic diagram illustrating inter-frame prediction using reconstructed pixel regions according to the fifth embodiment of the present invention. The inter-frame prediction using reconstructed pixel regions applicable to this embodiment is particularly capable of deriving the motion vector of the current block using the reconstructed pixel regions.

[0292] exist Figure 31aIn the diagram, the reconstructed pixel region C 251 is illustrated, representing the current block 252 that needs to be encoded or decoded, and the regions adjacent to the current block 252. The reconstructed pixel region C 251 includes two regions: the left-hand region and the upper-hand region of the current block 252. The current block 252 and the reconstructed pixel region C 251 are contained within the current image 250. The current image 250 can be an image, a strip, a parallel block, a coding tree block, a coding block, or other image region. From an encoding perspective, the reconstructed pixel region C 251 can be a pixel region that was encoded before the current block 252 was encoded and then reconstructed; from a decoding perspective, it can be a region that was reconstructed before the current block 252 was decoded.

[0293] Before encoding or decoding the current block, since a reconstructed pixel region C 251 exists around the current block 252, the image encoding device 100 and the decoding device 400 can utilize the same reconstructed pixel region C 251. Therefore, the encoding device 100 does not need to encode the motion information of the current block 252, but instead uses the reconstructed pixel region C 251 to generate the motion information of the current block 252 and generate the prediction block in the same way by the image encoding device 100 and the decoding device 400.

[0294] exist Figure 31b The illustration shows an embodiment of motion inference and motion compensation using reconstructed pixel regions. Figure 31b Within the reference image 253 shown, the image will be compared with... Figure 31a The region matching the reconstructed pixel region C 251 illustrated is explored. After determining the reconstructed pixel region D 256 most similar to the reconstructed pixel region C 52, the displacement between region 254, located at the same position as the reconstructed pixel region C 251, and the reconstructed pixel region D 256 is determined as the motion vector 257 of the reconstructed pixel region C 251. After selecting the motion vector 257 determined as described above as the motion vector of the current block 252, the predicted block of the current block 252 can be derived using the aforementioned motion vector 257.

[0295] Figure 32 According to Figure 31b The diagram illustrates the scenario where the motion vector 257, which is inferred in the manner shown, is set as the initial motion vector, and the current block 252 is divided into multiple sub-blocks A to D. The motion inference is then performed in sub-block units.

[0296] Subblocks A through D can have any size. Figure 32 The diagram shows MV_A to MV_D as the initial motion vectors of sub-blocks A to D, respectively. Figure 31b The motion vector 257 shown in the figure is the same.

[0297] The size of each sub-block can be transmitted to the decoding device 400 after being encoded by block unit or upper-level header, or the same sub-block size value can be used in the encoding device 100 and the decoding device 400.

[0298] In addition, such as Figures 7a to 7c As shown, the reconstructed pixel region C 251 can have different shapes and / or sizes. Furthermore, the reconstructed pixel regions on the upper and left sides of the current block can also be used as reconstructed pixel regions C, or as follows: Figures 31a to 31b The two regions described above are merged as shown in the diagram and then used as a single reconstructed pixel region C251. Furthermore, it can also be used after resampling the reconstructed pixel region C251.

[0299] For the sake of clarity, let's assume that... Figure 31a as well as Figure 31b The reconstructed pixel region C 251 shown in the figure will be explained as a reconstructed pixel region.

[0300] Figure 33 This is a schematic diagram illustrating an example of dividing the reconstructed pixel region C251 and the current block into sub-block units. See also... Figure 33 The reconstructed pixel region C 251 is divided into sub-blocks a 285, b 286, c 287 and d 288, and the current block is divided into sub-blocks A 281, B 282, C 283 and D 284.

[0301] The reconstructed pixel region of sub-block A 281 can use sub-blocks a 285 and c 287; the reconstructed pixel region of sub-block B 282 can use sub-blocks b 286 and c 287; the reconstructed pixel region of sub-block C 283 can use sub-blocks a 285 and d 288; and finally, the reconstructed pixel region of sub-block D 284 can use sub-blocks b 286 and d 288.

[0302] Figure 34 This is a sequence diagram illustrating an example of an inter-frame prediction method that utilizes reconstructed pixel regions. See also... Figure 31a as well as Figure 31b In step S291, the reconstructed pixel region 251 of the current block 252 is set. Next, in step S292, motion inference is performed on the reference image 253 using the reconstructed pixel region 251. Based on the result of the motion inference, the motion vector 257 of the reconstructed pixel region 251 can be calculated. Next, in step S293, according to... Figure 33 The illustrated method sets the reconstructed pixel region in sub-block units. Next, in step S294, after setting the motion vector 257 inferred in step S292 as the starting point, motion inference is performed in sub-block units of the current block.

[0303] Figure 35 This is a schematic diagram illustrating an example of how the present invention divides a reconstructed pixel region into sub-blocks using reconstructed blocks existing around the current block.

[0304] In one embodiment of the invention, the surrounding reconstructed pixel regions used in the prediction process of the current block can be segmented based on the segmentation structure of the surrounding reconstructed blocks. In other words, the reconstructed pixel regions can be segmented based on at least one of the following: the number of surrounding reconstructed blocks, the size of the surrounding reconstructed blocks, the shape of the surrounding reconstructed blocks, or the boundaries between the surrounding reconstructed blocks.

[0305] See Figure 35 There are reconstruction blocks 12101 to 52105 surrounding the current block 2100 that needs to be encoded or decoded. If we follow... Figure 5a The illustrated method sets the reconstructed pixel region because the sharp differences in pixel values ​​that may exist at the boundaries of reconstructed block 12101 to reconstructed block 52105 could lead to a decrease in efficiency during motion estimation. Therefore, according to... Figure 35 The method illustrated may help improve efficiency when using sub-blocks from a to e to divide the reconstructed pixel region. Figure 35 The reconstructed pixel region shown in the figure can also be segmented according to the segmentation method of the reconstructed blocks surrounding the current block 2100.

[0306] Specifically, when segmenting the reconstructed pixel region, the number of surrounding already reconstructed blocks can be considered. See also... Figure 35 Above the current block 2100, there are two reconstruction blocks: reconstruction block 1 (2101) and reconstruction block 2 (2102). To the left of the current block 2100, there are three reconstruction blocks: reconstruction blocks 3 (2103) to 5 (2105). Considering the above, the reconstructed pixel area above the current block 2100 can be divided into two sub-blocks, a and b, while the reconstructed pixel area to the left of the current block 2100 can be divided into three sub-blocks, c to e.

[0307] Alternatively, the size of surrounding reconstructed blocks can be considered when segmenting the reconstructed pixel region. For example, the height of sub-block c of the reconstructed pixel region to the left of the current block 2100 is the same as that of reconstructed block 3 2103, the height of sub-block d is the same as that of reconstructed block 4 2104, and the height of sub-block e is the same as the remaining value after subtracting the heights of sub-block c and sub-block d from the height of the current block 2100.

[0308] Alternatively, when segmenting the reconstructed pixel region, the boundaries between surrounding reconstructed blocks can be considered. Taking into account the boundaries between reconstructed blocks 1 (2101) and 2 (2102) above the current block 2100, the reconstructed pixel region above the current block 2100 can be segmented into two sub-blocks, a and b. Similarly, taking into account the boundaries between reconstructed blocks 3 (2103) and 4 (2104) to the left of the current block 2100, and the boundaries between reconstructed blocks 4 (2104) and 5 (2105), the reconstructed pixel region to the left of the current block 2100 can be segmented into three sub-blocks, c through e.

[0309] Furthermore, there may be several different conditions regarding which region among sub-blocks a to e should be used for motion inference. For example, motion inference can be performed using only the largest reconstructed pixel region, or m regions can be selected from the top and n regions from the left in priority order for motion inference. Alternatively, a single reconstructed pixel region 251, as illustrated in Figure 5, can be used after mitigating sharp differences in pixel values ​​by applying a filter such as a low-pass filter between sub-blocks a to e.

[0310] Figure 36 This is a schematic diagram illustrating an example of how the present invention divides the current block into multiple sub-blocks using reconstruction blocks existing around the current block.

[0311] Will Figure 36 The illustrated method of dividing the current block into multiple sub-blocks is similar to that of... Figure 35 The method for segmenting the reconstructed pixel region illustrated is similar. That is, the current block that needs to be encoded or decoded can be segmented based on the segmentation structure of the surrounding reconstructed blocks. In other words, the current block can be segmented based on at least one of the following: the number of surrounding reconstructed blocks, the size of surrounding reconstructed blocks, the shape of surrounding reconstructed blocks, or the boundaries between surrounding reconstructed blocks.

[0312] Figure 36The current block shown in the diagram is divided into multiple sub-blocks A to F. Inter-frame prediction can be performed on each sub-block unit obtained through the segmentation process described above. At this time, sub-block A can use... Figure 10 Reconstruction regions a and c, sub-block B can use reconstruction regions b and c, sub-block C can use reconstruction regions a and d, sub-block D can use reconstruction regions b and d, sub-block E can use reconstruction regions a and e, and sub-block F can use reconstruction regions b and e to perform inter-frame prediction respectively.

[0313] Alternatively, priority can be set based on the size of sub-blocks and the reconstructed pixel region. For example, because... Figure 36 The height of sub-block A shown in the diagram is greater than its length, thus allowing reconstruction region a to be given a higher priority than reconstruction region c, enabling inter-frame prediction to be performed using only reconstruction region a. Alternatively, it is possible to give reconstruction region c a higher priority based on image characteristics and other factors.

[0314] Figure 37 This is a sequence diagram illustrating a method for dividing a current block into multiple sub-blocks according to one embodiment of the present invention.

[0315] See Figure 37 In step S2201, the current block is first divided into multiple sub-blocks based on the blocks surrounding the current block that need to be encoded or decoded. The blocks surrounding the current block are as follows: Figure 36 The diagram shows the reconstructed blocks. (As shown in the image...) Figure 36 The above description indicates that the current block requiring encoding or decoding can be segmented based on the segmentation structure of surrounding reconstructed blocks. That is, the current block can be segmented based on at least one of the following: the number of surrounding reconstructed blocks, the size of surrounding reconstructed blocks, the shape of surrounding reconstructed blocks, or the boundaries between surrounding reconstructed blocks.

[0316] Next, in step S2203, multiple sub-blocks within the current block are encoded or decoded. In one embodiment of the present invention, as described above, inter-frame prediction can be used to... Figure 36 The sub-blocks A through F of the current block shown in the diagram are encoded or decoded respectively. At this time, sub-block A can use... Figure 35Reconstructed regions a and c, sub-block B can use reconstructed regions b and c, sub-block C can use reconstructed regions a and d, sub-block D can use reconstructed regions b and d, sub-block E can use reconstructed regions a and e, and sub-block F can use reconstructed regions b and e respectively to perform inter-frame prediction. Furthermore, inter-frame prediction related information, such as sub-block information or motion information used to indicate whether sub-blocks are divided into sub-blocks by performing inter-frame prediction on sub-blocks A to F respectively, can be encoded or decoded.

[0317] Figure 37 The illustrated method can be executed by the inter-frame prediction unit 103 of the image encoding apparatus 100 or the inter-frame prediction unit 408 of the image decoding apparatus 400, respectively. The reference image used during the inter-frame prediction process is stored in the memory 112 of the image encoding apparatus 100 or the memory 406 of the image decoding apparatus 400. The inter-frame prediction unit 103 or the inter-frame prediction unit 408 can generate a prediction block for the current block 51 by referring to the reference image stored in the memory 112 or the memory 406.

[0318] Figure 38 This is a sequence diagram illustrating a method for dividing a reconstructed region into multiple sub-blocks during the encoding or decoding process of the current block, according to one embodiment of the present invention.

[0319] See Figure 13 In step S2211, the reconstructed pixel region is first divided into multiple sub-blocks based on the blocks surrounding the current block that needs to be encoded or decoded. (For example, combining...) Figure 35 and / or Figure 36 The above description describes the segmentation of surrounding reconstructed pixel regions used in the prediction process of the current block based on the segmentation structure of surrounding reconstructed blocks. In other words, reconstructed pixel regions can be segmented based on at least one of the following: the number of surrounding reconstructed blocks, the size of surrounding reconstructed blocks, the shape of surrounding reconstructed blocks, or the boundaries between surrounding reconstructed blocks.

[0320] Next, in step S2213, at least one of the multiple sub-blocks within the current block is encoded or decoded using at least one sub-block contained in the reconstructed pixel region. For example, as in combination Figure 36 As explained above, sub-block A can use Figure 35Reconstructed regions a and c, sub-block B can use reconstructed regions b and c, sub-block C can use reconstructed regions a and d, sub-block D can use reconstructed regions b and d, sub-block E can use reconstructed regions a and e, and sub-block F can use reconstructed regions b and e respectively to perform inter-frame prediction. Furthermore, inter-frame prediction related information, such as sub-block information or motion information used to indicate whether sub-blocks are divided into sub-blocks by performing inter-frame prediction on sub-blocks A to F respectively, can be encoded or decoded.

[0321] Figure 38 The illustrated method can be executed by the inter-frame prediction unit 103 of the image encoding apparatus 100 or the inter-frame prediction unit 408 of the image decoding apparatus 400, respectively. The reference image used during the inter-frame prediction process is stored in the memory 112 of the image encoding apparatus 100 or the memory 406 of the image decoding apparatus 400. The inter-frame prediction unit 103 or the inter-frame prediction unit 408 can generate a prediction block for the current block 51 by referring to the reference image stored in the memory 112 or the memory 406.

[0322] Figure 39 It is to utilize according to such Figure 36 The illustrated sequence diagram shows an embodiment of an inter-frame prediction method for sub-blocks of the current block segmented in the manner shown. Figure 39 The method illustrated can be executed by the inter-frame prediction unit 103 of the image coding apparatus 100.

[0323] First, the two variables used in this method, namely the decoder-side motion vector derivation (DMVD) indication information and SUB_BLOCK, will be explained. The decoder-side motion vector derivation (DMVD) indication information is used to indicate whether to perform inter-frame prediction using existing methods or the inter-frame prediction using reconstructed pixel regions applicable to the present invention, as described above. When the decoder-side motion vector derivation (DMVD) indication information is false, it indicates that inter-frame prediction using existing methods is performed. When the decoder-side motion vector derivation (DMVD) indication information is true, it indicates that inter-frame prediction using reconstructed pixel regions applicable to the present invention is performed.

[0324] The variable SUB_BLOCK indicates whether the current block should be split into sub-blocks. When SUB_BLOCK is false, it means the current block will not be split into sub-blocks. Conversely, when SUB_BLOCK is true, it means the current block will be split into sub-blocks.

[0325] See Figure 39In step S2301, the variable Decoder-Side Motion Vector Derivation (DMVD) indication information, which indicates whether or not inter-frame prediction using the reconstructed pixel region is performed, is first set to false. The variable SUB_BLOCK, which indicates whether or not the block is segmented into sub-blocks, is set to false. Then, cost_1 is calculated after inter-frame prediction of the current block is performed.

[0326] Next, in step S2302, SUB_BLOCK is set to true, and cost_2 is calculated after inter-frame prediction is performed. Next, in step S2303, the decoder-side motion vector derivation (DMVD) indication information is set to true, and SUB_BLOCK is set to false. Inter-frame prediction is then performed, and cost_3 is calculated. Next, in step S2304, both the decoder-side motion vector derivation (DMVD) indication information and SUB_BLOCK are set to true, and inter-frame prediction is performed, and cost_4 is calculated. The optimal inter-frame prediction method is determined after comparing the calculated costs_1 to cost_4. The decoder-side motion vector derivation (DMVD) indication information and SUB_BLOCK information related to the determined optimal inter-frame prediction method are saved, and then the corresponding algorithm ends.

[0327] Figure 40 According to Figure 39 The diagram illustrates the encoding method for information determined by inter-frame prediction. Figure 40 The encoding method illustrated can be executed by the image encoding device 100.

[0328] In step S2401, because Figure 36 The total number of sub-blocks in the current block illustrated is set to 6. Therefore, the variable BLOCK_NUM, which indicates the total number of sub-blocks to be encoded, is initialized to 6, and the variable BLOCK_INDEX, which indicates the index of the sub-blocks to be encoded, is initialized to 0. Since the current block is divided into sub-blocks based on the reconstructed blocks surrounding it, it is not necessary to encode the number of sub-blocks separately. Because the image decoding device 400 also divides the current block into sub-blocks in the same way as the image encoding device 100, the image decoding device 400 can confirm the number of possible sub-blocks within the current block.

[0329] Following step S2401, in step S2402, the information SUB_BLOCK, used to indicate whether to divide the current block into sub-blocks, is encoded. In step S2403, it is confirmed whether to divide the current block into sub-blocks. If it is not divided into sub-blocks, in step S2404, the value of the BLOCK_NUM variable is changed to 1.

[0330] Next, in step S2405, the decoder-side motion vector derivation (DMVD) indication information, which indicates whether or not inter-frame prediction using the reconstructed pixel region is used, is encoded. In step S2406, it is confirmed whether inter-frame prediction using the reconstructed pixel region is used. If inter-frame prediction using the reconstructed pixel region is not used, motion information is encoded in step S2407. Conversely, if it is used, in step S2408, the value of BLOCK_INDEX is incremented, and in S2409, it is compared with the BLOCK_NUM variable. When the value of BLOCK_INDEX is the same as the value of BLOCK_NUM, it indicates that there are no more sub-blocks to be encoded within the current block, and therefore the corresponding algorithm ends. When the two values ​​are different, the process repeats from step S2406 after moving to the next sub-block within the current block that needs to be encoded.

[0331] Figure 41 According to Figure 40 The illustrated sequence diagram illustrates an example of a decoding method for information encoded by the encoding method. In step S2501, because... Figure 36 The total number of sub-blocks of the current block illustrated is set to 6. Therefore, the variable BLOCK_NUM, which indicates the total number of sub-blocks to be decoded, is initialized to 6, and the variable BLOCK_INDEX, which indicates the index of the sub-blocks to be encoded, is initialized to 0. As described above, since both the image decoding device 400 and the image encoding device 100 divide the current block into sub-blocks in the same way based on the reconstructed blocks surrounding the current block, it is not necessary to provide the image decoding device 400 with information indicating the number of sub-blocks separately. The image decoding device 400 can determine the number of possible sub-blocks within the current block based on the reconstructed blocks surrounding the current block.

[0332] Following step S2501, in step S2502, the information SUB_BLOCK, which indicates whether to divide the current block into sub-blocks, is decoded. In step S2403, it is confirmed whether to divide the current block into sub-blocks. If it is not divided into sub-blocks, in step S2404, the value of the BLOCK_NUM variable is changed to 1.

[0333] Next, in step S2505, the decoder-side motion vector derivation (DMVD) indication information, which indicates whether or not inter-frame prediction using the reconstructed pixel region is used, is decoded. In step S2506, it is confirmed whether inter-frame prediction using the reconstructed pixel region is used. If inter-frame prediction using the reconstructed pixel region is not used, in step S2507, motion information is decoded. Conversely, if it is used, in step S2508, the value of BLOCK_INDEX is incremented, and in S2509, it is compared with the BLOCK_NUM variable. When the value of BLOCK_INDEX is the same as the value of BLOCK_NUM, it indicates that there are no sub-blocks that need to be decoded in the current block, and therefore the corresponding algorithm ends. When the two values ​​are different, the process repeats from step S2506 after moving to the next sub-block that needs to be decoded in the current block.

[0334] (Sixth Embodiment)

[0335] Next, a sixth embodiment to which the present invention is applied will be described with reference to the accompanying drawings.

[0336] Figure 42a And 42b is a schematic diagram for illustrating the sixth embodiment to which the present invention is applicable.

[0337] like Figure 42a as well as Figure 42b As shown, assuming that there are reconstruction blocks 1 to 6 around the current block 2600, it is possible to proceed as follows: Figure 36 The diagram illustrates how the reconstructed pixel region is divided into sub-blocks a to f. Following the... Figure 37 The method shown in the diagram can divide the current block 2600 into sub-blocks A to I.

[0338] Because sub-blocks F, G, H, and I are located in separate positions that are not adjacent to the reconstructed pixel region, inter-frame prediction using the reconstructed pixel region may be incorrect. Therefore, existing inter-frame prediction can be performed in sub-blocks F, G, H, and I, while inter-frame prediction using the reconstructed pixel region is only used in sub-blocks A to E.

[0339] When performing inter-frame prediction on sub-blocks A through E using reconstructed pixel regions, inter-frame prediction can be performed using reconstructed pixel regions adjacent to each sub-block. For example, sub-block B can use reconstructed pixel region b, sub-block C can use reconstructed pixel region c, sub-block D can use reconstructed pixel region e, and sub-block E can use reconstructed pixel region f. For sub-block A, inter-frame prediction can be performed using one or both reconstructed pixel regions a and d, or simultaneously, according to a pre-defined priority order.

[0340] Alternatively, when performing inter-frame prediction on sub-blocks A through E using reconstructed pixel regions, the index used to indicate which reconstructed pixel region is used for each sub-block can be encoded. For example, inter-frame prediction for sub-block A can also be performed using reconstructed pixel region b from reconstructed pixel regions a through f. Sub-block E can also be performed using reconstructed pixel region c. In the case described above, the priority order and index can be determined and assigned based on the horizontal or vertical size of each reconstructed pixel region a through f, the number of pixels in each region, their positions, etc.

[0341] For sub-blocks F to I, encoding or decoding can be performed by executing existing inter-frame prediction. As one approach, such as... Figure 42b As shown, it is possible to encode or decode using existing inter-frame prediction after merging sub-blocks F to I into one.

[0342] Figure 43 For reference Figure 42a as well as Figure 42b A sequence diagram illustrating an example of the inter-frame prediction mode determination method applicable to the sixth embodiment of the present invention is provided. For ease of explanation, this example assumes that the peripheral reconstruction blocks include reconstruction blocks 1 2601 to 6 2606 as illustrated in FIG. 42, and that the reconstructed pixel region is divided into reconstructed pixel regions a to f. Furthermore, it is assumed that the current block 2600 is divided into sub-blocks A to F. In this case, sub-block F is Figure 42a Sub-blocks F to I in the process are merged into one result.

[0343] Furthermore, the case of encoding an index indicating which sub-reconstruction region to use when performing inter-frame prediction utilizing reconstructed pixel regions will be illustrated as an example. Sub-block F will be illustrated as an example of encoding or decoding by performing existing inter-frame prediction. Next, it is assumed that sub-block F in the sub-blocks within the current block was the last to be encoded or decoded.

[0344] See Figure 43 In step S2701, firstly, after performing inter-frame prediction without dividing the current block into sub-blocks, cost_1 is calculated. Next, in step S2702, after performing inter-frame prediction on sub-blocks A to F respectively, cost_A to cost_F are calculated, and then cost_2 is calculated by adding them. In step S2703, the calculated cost_1 and cost_2 are compared. If cost_1 is smaller, in step S2704, it is decided not to divide the block into sub-blocks; otherwise, in step S2705, after deciding to divide the block into sub-blocks, inter-frame prediction is performed, and then the algorithm ends.

[0345] Figure 44 According to Figure 43 The illustrated method is a schematic diagram illustrating the information encoding process. In step S2801, because... Figure 42b The total number of sub-blocks in the current block illustrated is set to 6. Therefore, the variable BLOCK_NUM, which indicates the total number of sub-blocks to be encoded, is initialized to 6, and the variable BLOCK_INDEX, which indicates the index of the sub-blocks to be encoded, is initialized to 0. Since the current block is divided into sub-blocks based on the reconstructed blocks surrounding it, it is not necessary to encode the number of sub-blocks separately. Because the image decoding device 400 also divides the current block into sub-blocks in the same way as the image encoding device 100, the image decoding device 400 can confirm the number of possible sub-blocks within the current block.

[0346] Following step S2801, in step S2802, the information SUB_BLOCK, used to indicate whether to divide the current block into sub-blocks, is encoded. In step S2803, confirmation is made regarding whether to divide the current block into sub-blocks. If it is not divided into sub-blocks, in step S2804, the value of the BLOCK_NUM variable is changed to 1.

[0347] In step S2805, the values ​​of BLOCK_INDEX and BLOCK_NUM-1 are compared to determine if the current block is one that requires existing inter-frame prediction rather than inter-frame prediction using reconstructed pixel regions. If the two values ​​are the same, it indicates that this is the last block using existing inter-frame prediction, and therefore, in step S2806, motion information is encoded. Otherwise, it indicates that this is a sub-block using inter-frame prediction using reconstructed pixel regions, and therefore, in step S2807, the index indicating which sub-reconstructed region to use is encoded. Alternatively, this step can be omitted, and the same reconstructed region agreed upon in the encoding and decoding devices can be used.

[0348] Next, in step S2808, the index of the sub-block is increased. In step S2809, BLOCK_NUM and BLOCK_INDEX are compared to determine if they are the same, thereby confirming whether the encoding of all sub-blocks existing in the current block has been completed. Otherwise, proceed to step S2805 and continue executing the corresponding algorithm.

[0349] Figure 45 According to Figure 44 The diagram illustrates the decoding process of the information encoded by the method. In step S2901, because... Figure 42bThe total number of sub-blocks in the current block illustrated is set to 6. Therefore, the variable BLOCK_NUM, which indicates the total number of sub-blocks to be encoded, is initialized to 6, and the variable BLOCK_INDEX, which indicates the index of the sub-blocks to be encoded, is initialized to 0. Since the current block is divided into sub-blocks based on the reconstructed blocks surrounding it, it is not necessary to encode the number of sub-blocks separately. Because the image decoding device 400 also divides the current block into sub-blocks in the same way as the image encoding device 100, the image decoding device 400 can confirm the number of possible sub-blocks within the current block.

[0350] Following step S2901, in step S2902, the information SUB_BLOCK, which indicates whether to divide the current block into sub-blocks, is decoded. In step S2903, it is confirmed whether to divide the current block into sub-blocks. If it is not divided into sub-blocks, in step S2904, the value of the BLOCK_NUM variable is changed to 1.

[0351] In step S2905, the values ​​of BLOCK_INDEX and BLOCK_NUM-1 are compared to determine if the current block is one that requires existing inter-frame prediction rather than inter-frame prediction using reconstructed pixel regions. If the two values ​​are the same, it indicates that this is the last block using existing inter-frame prediction, and therefore, in step S2906, motion information is encoded. Otherwise, it indicates that this is a sub-block using inter-frame prediction using reconstructed pixel regions, and therefore, in step S2907, the index indicating which sub-reconstructed region to use is decoded. Alternatively, this step can be omitted, and the same reconstructed region agreed upon in the encoding and decoding devices can be used. Next, in step S2908, the sub-block index is incremented. In step S2909, BLOCK_NUM and BLOCK_INDEX are compared to determine if encoding of all sub-blocks existing within the current block is complete. Otherwise, the process moves to step S2905 and continues with the corresponding algorithm.

[0352] The exemplary methods in this disclosure are described as a sequence of actions for clarity of explanation, but this is not intended to limit the order in which the steps are executed. The steps can be executed simultaneously or in different orders if necessary. To implement the methods in this disclosure, additional steps can be added to the example steps, or only the remaining steps (excluding some steps) can be included, or additional steps can be added after excluding some steps.

[0353] The various embodiments described herein are not a list of all possible combinations, but are merely illustrative of representative forms of the disclosure. The matters described in the various embodiments may apply independently or in combination of two or more.

[0354] Furthermore, the various embodiments described in this disclosure can be implemented using hardware, firmware, software, or a combination thereof. When implemented in hardware, they can be implemented using one or more ACICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general-purpose processors, controllers, microcontrollers, microprocessors, etc.

[0355] The scope of this disclosure includes software, device-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that enable actions in methods of various embodiments to be executed on an apparatus or computer, and a device- or computer-executable non-transitory computer-readable medium storing the aforementioned software or instructions.

[0356] Industrial applicability: This invention is applicable to the field of encoding or decoding image signals.

Claims

1. An image decoding method, executed by an image decoding device, the image decoding method comprising the following steps: Motion vectors for the current block in the current image are derived based on a first reconstructed pixel region, wherein the first reconstructed pixel region was reconstructed before and adjacent to the current block, and the current block is divided into multiple sub-blocks; and Predict the sub-block based on the motion vector derived from the current block. in, The steps for predicting the sub-blocks of the current block include the following: By setting the motion vector of the current block as the starting point, the second reconstructed pixel region of the sub-block within the reference image of the current block is obtained through motion search; The sub-block motion vector of the sub-block is derived based on the displacement between the second reconstructed pixel region and the temporary region of the sub-block in the reference image of the current block, wherein the temporary region is obtained based on the third reconstructed pixel region of the sub-block in the current image. as well as Predict the sub-block based on the sub-block motion vector. The third reconstructed pixel region is a portion of the first reconstructed pixel region, which is divided based on the sub-block partitioning. The motion search is performed only within a predefined search range defined in the image decoding device. The search area is a portion of the reference image. The transform coefficients of the current block are decoded by entropy decoding of the bitstream. The transformed block of the current block is generated by performing inverse quantization and inverse transform on the transform coefficients of the current block. The inverse transform includes a vertical inverse transform and a horizontal inverse transform. The vertical inverse transform is performed using a DST-based transform, and the horizontal inverse transform is performed using a DCT-based transform. The current block is reconstructed based on the transformed block.

2. The image decoding method according to claim 1, wherein, Based on the first flag, it is determined whether the motion vector of the block is derived through motion search based on the reconstructed pixel region. The first flag is defined on a block-by-block basis and indicates whether the decoder-side motion derivation is used for the block.

3. The image decoding method according to claim 2, wherein, The first flag depends on a second flag indicating whether decoder-side motion derivation is permitted for an image including the block.

4. The image decoding method according to claim 3, wherein, In response to the second flag being true, the decoder-side motion derivation is adaptively applied to the block based on the first flag. In response to the second flag being false, the decoder-side motion derivation is not used for the block.

5. The image decoding method according to claim 4, wherein, The second flag is defined in units of images and is signaled from the image header of the bitstream.

6. An image encoding method, executed by an image encoding device, the image encoding method comprising the following steps: Motion vectors for the current block in the current image are derived based on a first reconstructed pixel region, wherein the first reconstructed pixel region was reconstructed before and adjacent to the current block, and the current block is divided into multiple sub-blocks; and Predict the sub-block based on the motion vector derived from the current block. in, The steps for predicting the sub-blocks of the current block include the following: By setting the motion vector of the current block as the starting point, the second reconstructed pixel region of the sub-block within the reference image of the current block is obtained through motion search; The sub-block motion vector of the sub-block is derived based on the displacement between the second reconstructed pixel region and the temporary region of the sub-block in the reference image of the current block, wherein the temporary region is obtained based on the third reconstructed pixel region of the sub-block in the current image. as well as Predict the sub-block based on the sub-block motion vector. The third reconstructed pixel region is a portion of the first reconstructed pixel region, which is divided based on the sub-block partitioning. The motion search is performed only within a predefined search range defined in the image encoding device. The search area is a portion of the reference image. The transformed block of the current block is generated based on the original block of the current block. The transform coefficients of the current block are obtained by performing a transform and quantization on the transformed block. The transform includes a vertical transform and a horizontal transform. The vertical transform is performed using a DST-based transform, and the horizontal transform is performed using a DCT-based transform. A bit stream is generated by encoding the transform coefficients of the current block.

7. A non-transitory computer-readable medium storing a computer program that, when executed in conjunction with a computer as hardware, generates a bitstream by performing an encoding method, wherein... The encoding method includes the following steps: Motion vectors for the current block in the current image are derived based on a first reconstructed pixel region, wherein the first reconstructed pixel region was reconstructed before and adjacent to the current block, and the current block is divided into multiple sub-blocks; and Predict the sub-block based on the motion vector derived from the current block. The step of predicting the sub-blocks of the current block includes the following steps: By setting the motion vector of the current block as the starting point, the second reconstructed pixel region of the sub-block within the reference image of the current block is obtained through motion search; The sub-block motion vector of the sub-block is derived based on the displacement between the second reconstructed pixel region and a temporary region of the sub-block within the reference image of the current block, wherein the temporary region is obtained based on a third reconstructed pixel region of the sub-block within the current image; and Predict the sub-block based on the sub-block motion vector. The third reconstructed pixel region is a portion of the first reconstructed pixel region, which is divided based on the sub-block partitioning. The motion search is performed only within a predefined search range in the image encoding device. The search area is a portion of the reference image. The transformed block of the current block is generated based on the original block of the current block. The transform coefficients of the current block are obtained by performing a transform and quantization on the transformed block. The transform includes a vertical transform and a horizontal transform. The vertical transform is performed using a DST-based transform, and the horizontal transform is performed using a DCT-based transform. The bit stream is generated by encoding the transform coefficients of the current block.

8. A method for transmitting a bit stream, comprising the following steps: The bitstream is generated by performing an encoding method; and Transmit the bit stream, in, The encoding method includes the following steps: Motion vectors for the current block in the current image are derived based on a first reconstructed pixel region, wherein the first reconstructed pixel region was reconstructed before and adjacent to the current block, and the current block is divided into multiple sub-blocks; and Predict the sub-block based on the motion vector derived from the current block; The step of predicting the sub-blocks of the current block includes the following steps: By setting the motion vector of the current block as the starting point, the second reconstructed pixel region of the sub-block within the reference image of the current block is obtained through motion search; The sub-block motion vector of the sub-block is derived based on the displacement between the second reconstructed pixel region and a temporary region of the sub-block within the reference image of the current block, wherein the temporary region is obtained based on a third reconstructed pixel region of the sub-block within the current image; and Predict the sub-block based on the sub-block motion vector. The third reconstructed pixel region is a portion of the first reconstructed pixel region, which is divided based on the sub-block partitioning. The motion search is performed only within a predefined search range in the image encoding device. The search area is a portion of the reference image. The transformed block of the current block is generated based on the original block of the current block. The transform coefficients of the current block are obtained by performing a transform and quantization on the transformed block. The transform includes a vertical transform and a horizontal transform. The vertical transform is performed using a DST-based transform, and the horizontal transform is performed using a DCT-based transform. The bit stream is generated by encoding the transform coefficients of the current block.