Video encoding and decoding method and equipment thereof
By generating motion information of specific blocks of size and size before high-resolution image encoding and decoding, and using this information during the encoding and decoding process, the problem of image quality reduction and high computational complexity caused by square codec units in the prior art is solved, and more efficient codec performance is achieved.
Patent Information
- Application Number
- CN202380079875.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-04
- Filing Date
- 2023-10-10
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art uses square-uniform codec units in high-resolution image encoding and decoding, resulting in a decrease in image quality of the reconstructed image and a high computational complexity.
Motion information in units of a specific size is generated based on pixel data or motion information of the reconstructed picture before encoding and decoding, and the motion information is used during the encoding and decoding of each block.
It improves the accuracy and prediction performance of the motion information of the current block, reduces the computational complexity, and improves the encoding and codec performance and efficiency.
Smart Images

Figure CN120202670A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video coding, and more particularly, to an image decoding method, an image decoding apparatus, an image encoding method, an image encoding apparatus, and a storage medium storing a bitstream generated by the image encoding method. Background Art
[0002] In the case of a conventional compression method, during the process of determining the size of a coding / decoding unit included in a picture, it is determined whether to perform segmentation, and then a square coding / decoding unit is determined through a recursive segmentation process of uniformly dividing it into four coding / decoding units of the same size. However, recently, due to using a square uniform-shaped coding / decoding unit for high-resolution images, the deterioration of the image quality of the reconstructed image has become a problem. Accordingly, methods and apparatuses for dividing high-resolution images into various types of coding / decoding units have been proposed. Summary of the Invention
[0003] Technical Problem
[0004] A method and an apparatus are provided that can improve the accuracy or prediction performance of motion information of a current block and improve coding / decoding performance or efficiency while reducing computational complexity.
[0005] More specifically, a method and an apparatus are provided that can improve the accuracy or prediction performance of motion information of blocks in a current picture and improve coding / decoding performance or efficiency by: before starting to code / decoding a current picture, generating motion information of the current picture in units of blocks of a specific size based on at least one of pixel data or motion information of at least one reconstructed picture; and using the generated motion information when coding / decoding each block of the current picture.
[0006] Solution to the Problem
[0007] According to a first aspect of the present disclosure, an image decoding method executed by an apparatus is provided, the method including: generating reference motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; obtaining motion information of a current block included in the current picture based on the generated reference motion information and at least one syntax information obtained from a bitstream; and reconstructing the current block based on the obtained motion information.
[0008] According to a second aspect of the present disclosure, an image decoding apparatus is provided. The image decoding apparatus includes at least one processor configured to perform operations, and the operations may include: generating reference motion information for a block included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; obtaining motion information for a current block included in the current picture based on the generated reference motion information and at least one syntax information obtained from a bitstream; and reconstructing the current block based on the obtained motion information.
[0009] Preferably, generating reference motion information for a block included in a current picture may include: inputting at least one of pixel data or motion information of at least one reconstructed picture into an artificial neural network to generate reference motion information for a block included in the current picture.
[0010] Preferably, the motion information may include at least one of first motion vector information, second motion vector information, first prediction list utilization information, second prediction list utilization information, and prediction mode information.
[0011] Preferably, the method or the operation may further include at least one of an operation of aggregating the generated reference motion information, an operation of scaling the generated reference motion information, or an operation of transforming the format of the generated reference motion information.
[0012] Preferably, generating reference motion information for a block included in a current picture may include: downsampling at least one reconstructed picture; and generating reference motion information for a block included in the current picture based on at least one of pixel data or motion information of the downsampled reconstructed picture.
[0013] Preferably, obtaining motion information for the current block may include: excluding the generated reference motion information from candidates for the motion information of the current block when a time interval between the reconstructed picture and the current picture is equal to or greater than a specific value.
[0014] Preferably, generating reference motion information for a block included in a current picture may include: downsampling a reconstructed picture based on a low-resolution reference picture when applying reference picture resampling.
[0015] Preferably, the method or the operation may further include: generating motion information for another image by performing interpolation or extrapolation based on motion information of at least one reconstructed picture.
[0016] Preferably, at least one syntax information may include information indicating whether the generated reference motion information is used to obtain the motion information of the current block.
[0017] Preferably, obtaining motion information of a current block may include: constructing a merge candidate list including generated reference motion information; and obtaining the motion information of the current block based on the constructed merge candidate list.
[0018] Preferably, the generated reference motion information may include motion information at a position corresponding to an x coordinate and a y coordinate, where the x coordinate is obtained in a reference motion information map by adding half the width of the current block to the upper left sample position of the generated current block, and the y coordinate is obtained by adding half the height of the current block to the upper left sample position of the current block.
[0019] Preferably, obtaining motion information of a current block may include dividing the current block into a plurality of sub - blocks and assigning the generated reference motion information to the plurality of sub - blocks, and reconstructing the current block may include performing motion compensation on the sub - blocks based on the generated reference motion information of the sub - blocks.
[0020] Preferably, obtaining motion information of a current block may include obtaining a plurality of control point motion vectors of the current block based on the generated reference motion information, and reconstructing the current block may include performing motion compensation on sub - blocks of the current block by using the obtained control point motion vectors.
[0021] According to a third aspect of the present disclosure, there is provided an image encoding method performed by a device, including: generating reference motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; determining motion information of a current block included in the current picture; and encoding the motion information of the current block into a bitstream based on the generated reference motion information.
[0022] According to a fourth aspect of the present disclosure, there is provided an image encoding device, the image encoding device including at least one processor configured to perform operations, and the operations may include: generating reference motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; determining motion information of a current block included in the current picture; and encoding the motion information of the current block into a bitstream based on the generated reference motion information.
[0023] According to another aspect of the present disclosure, there is provided a storage medium storing a bitstream generated by an image encoding method or an image encoding device.
[0024] Advantageous effects of the invention
[0025] According to the present disclosure, the accuracy or prediction performance of the motion information of the current block can be improved, and the encoding / decoding performance or efficiency can be improved while reducing the computational complexity.
[0026] More specifically, according to the present disclosure, the accuracy or prediction performance of the motion information of the blocks of the current picture and the coding / decoding performance or efficiency can be improved by the following operations: before starting to code / decodethe current picture, generating the motion information of the current picture in units of blocks of a specific size based on at least one of the pixel data or motion information of at least one reconstructed picture; and using the generated motion information when coding / decoding each block of the current picture. Description of the Drawings
[0027] To better understand the drawings cited herein, a brief description of each drawing is provided.
[0028] Figure 1 is a schematic block diagram of an image decoding apparatus according to an embodiment.
[0029] Figure 2 is a flowchart of an image decoding method according to an embodiment.
[0030] Figure 3 illustrates a process of determining at least one coding / decoding unit by dividing a current coding / decoding unit performed by an image decoding apparatus according to an embodiment.
[0031] Figure 4 illustrates a process of determining at least one coding / decoding unit by dividing a non-square coding / decoding unit performed by an image decoding apparatus according to an embodiment.
[0032] Figure 5 illustrates a process of dividing a coding / decoding unit based on at least one of block shape information or segmentation shape mode information performed by an image decoding apparatus according to an embodiment.
[0033] Figure 6 illustrates a method of determining a preset coding / decoding unit from among an odd number of coding / decoding units performed by an image decoding apparatus according to an embodiment.
[0034] Figure 7 illustrates the order of processing a plurality of coding / decoding units when an image decoding apparatus determines a plurality of coding / decoding units by dividing a current coding / decoding unit according to an embodiment.
[0035] Figure 8 illustrates a process of determining that a current coding / decoding unit will be divided into an odd number of coding / decoding units performed by an image decoding apparatus when the coding / decoding units cannot be processed in a preset order according to an embodiment.
[0036] Figure 9 illustrates a process of determining at least one coding / decoding unit by dividing a current coding / decoding unit performed by an image decoding apparatus according to an embodiment.
[0037] Figure 10Shows that when a second codec unit with a non-square shape satisfies a preset condition according to an embodiment, the shape into which the second codec unit can be divided is restricted, where the second codec unit with a non-square shape is determined when an image decoding device divides a first codec unit.
[0038] Figure 11 Shows the process of dividing a square codec unit performed by an image decoding device when the segmentation shape mode information indicates that the square codec unit will not be divided into four square codec units according to an embodiment.
[0039] Figure 12 Shows that according to an embodiment, the processing order between multiple codec units can be changed depending on the process of dividing the codec units.
[0040] Figure 13 Shows the process of determining the depth of a codec unit as the shape and size of the codec unit change when the codec unit is recursively divided to determine multiple codec units according to an embodiment.
[0041] Figure 14 Shows the depth that can be determined based on the shape and size of the codec unit and the partial index (PID) for distinguishing the codec unit according to an embodiment.
[0042] Figure 15 Shows determining multiple codec units based on multiple preset data units included in a picture according to an embodiment.
[0043] Figure 16 Is a block diagram of an image coding and decoding system.
[0044] Figure 17 Is a block diagram of an image decoding device according to the present disclosure.
[0045] Figure 18 Is a flowchart of an image decoding method according to the present disclosure.
[0046] Figure 19 Shows an image coding device according to the present disclosure.
[0047] Figure 20 Is a flowchart of an image coding method according to the present disclosure.
[0048] Figures 21a to 21c Shows a method of generating motion information according to the present disclosure.
[0049] Figure 22 Shows a method of normalizing the input to a motion information generator according to the present disclosure.
[0050] Figure 23A graph of motion information of an input according to the present disclosure is shown.
[0051] Figure 24 A motion information graph generated according to the present disclosure is shown.
[0052] Figure 25 A motion information generation model and learning method according to the present disclosure are shown.
[0053] Figure 26 A method of generating motion information according to the present disclosure is shown.
[0054] Figure 27 A process of aggregating motion information according to the present disclosure is shown.
[0055] Figure 28 A method of downscaling an input reconstructed picture according to the present disclosure is shown.
[0056] Figure 29 Reference picture resampling (RPR) is shown.
[0057] Figure 30 , Figure 31a and Figure 31b A method of utilizing a motion information generation method in frame rate upconversion (FRUC) according to the present disclosure is shown.
[0058] Figure 32a and Figure 32b A method for refining a generated motion information map according to the present disclosure is shown.
[0059] Figures 33 to 38 The syntax structure and encoding and decoding method according to the present disclosure are shown.
[0060] Figure 39a and Figure 39b The use of the generated motion information in merge mode according to the present disclosure is shown.
[0061] Figure 40 The positions of neighboring blocks that can be used as merge candidates are shown.
[0062] Figures 41a to 41d A method for determining a position corresponding to a current block in a generated motion information map according to the present disclosure is shown.
[0063] Figures 42a to 42c A method of performing sub-block prediction by using generated motion information according to the present disclosure is shown. DETAILED DESCRIPTION
[0064] Since the present disclosure allows for various variations and many embodiments, exemplary embodiments will be shown in the drawings and described in detail in the written description. However, this is not intended to limit the present disclosure to a specific practice mode, and it should be understood that all variations, equivalents, and alternatives that do not depart from the spirit and technical scope of the present disclosure are encompassed therein.
[0065] In the description of the embodiments, when it is considered that some detailed explanations of the related art may unnecessarily obscure the essence of the present disclosure, these explanations are omitted. Additionally, the numbers (e.g., first and second) used in the description of the specification are merely identification codes for distinguishing one component from another.
[0066] Furthermore, in the specification, it will be understood that when elements are "connected" or "coupled" to each other, the elements can be directly connected or coupled to each other, but they can also be connected or coupled to each other through intermediate elements therebetween, unless otherwise specified.
[0067] In the specification, regarding components represented as "parts (units)" or "modules", according to the subdivided functions, two or more components can be combined into one component, or one component can be divided into two or more components. Additionally, each of the components described below can additionally perform some or all of the functions performed by another component in addition to its own main function, and some of the main functions of each component can be fully performed by another component.
[0068] Furthermore, the term "image" or "picture" as used herein can refer to a still image or a moving image of a video (i.e., the video itself).
[0069] Furthermore, the term "sample" as used herein refers to the data that is assigned to the sampling positions of an image and will be processed. For example, the pixel values of an image in the spatial domain or the transform coefficients in the transform domain can be samples. A unit including one or more samples can be defined as a block.
[0070] The "current block" as used herein can represent the block of the largest coding / decoding unit, coding / decoding unit, prediction unit, or transform unit of the current picture to be encoded or decoded.
[0071] In this specification, when a motion vector is in the direction of list 0, this may mean that the motion vector is used to indicate a block in a reference picture included in list 0 (or reference list 0 or reference picture list 0), and when a motion vector is in the direction of list 1, this may mean that the motion vector is used to indicate a block in a reference picture included in list 1 (or reference list 1 or reference picture list 1). When a motion vector is unidirectional, this may mean that the motion vector is used to indicate a block in a reference picture included in list 0 or list 1, and when a motion vector is bidirectional, this may mean that the motion vector includes a motion vector in the direction of list 0 and a motion vector in the direction of list 1. List 0 can be abbreviated as L0, and list 1 can be abbreviated as L1.
[0072] In this specification, "binary splitting" of a block means achieving a split into two sub-blocks, where the width or height of each sub-block is half of the width or height of the block. Specifically, when performing "binary vertical splitting" on the current block, a split is performed in the vertical direction (height direction) at a point that is half the width of the current block, such that two sub-blocks can be generated, where each sub-block has a width that is half the width of the current block and a height that is the same as the height of the current block. When performing "binary horizontal splitting" on the current block, a split is performed in the horizontal direction (width direction) at a point that is half the height of the current block, such that two sub-blocks can be generated, where each sub-block has a height that is half the height of the current block and a width that is the same as the width of the current block.
[0073] In this specification, "ternary splitting" of a block means achieving a split into three sub-blocks, where the three sub-blocks are generated by splitting the width or height of the block in a 1:2:1 ratio. Specifically, when performing "ternary vertical splitting" on the current block, a split is performed in the vertical direction (height direction) at points where the ratio of the width of the current block is 1:2:1, such that two sub-blocks with a width of 1 / 4 of the width of the current block and a height that is the same as the height of the current block can be generated, and one sub-block with a width of 2 / 4 of the width of the current block and a height that is the same as the height of the current block. When performing "ternary horizontal splitting" on the current block, a split is performed in the horizontal direction (width direction) at points where the ratio of the height of the current block is 1:2:1, such that two sub-blocks with a height of 1 / 4 of the height of the current block and a width that is the same as the width of the current block can be generated, and one sub-block with a height of 2 / 4 of the height of the current block and a width that is the same as the width of the current block.
[0074] In this specification, the "quad-tree partitioning" of a block refers to the partitioning that realizes four sub-blocks, where the four sub-blocks are generated by partitioning the width and height of the block in a 1:1 ratio. Specifically, when performing "quad-tree partitioning" on the current block, a partition is performed in the vertical direction (height direction) at a point that is half the width of the current block, and a partition is performed in the horizontal direction (width direction) at a point that is half the height of the current block, such that four sub-blocks can be generated, where the width of each sub-block is 1 / 2 of the width of the current block and the height is 1 / 2 of the height of the current block.
[0075] Hereinafter, reference will be made to Figures 1 to 16 to disclose an image encoding method and apparatus and an image decoding method and apparatus according to an embodiment.
[0076] Figure 1 is a schematic block diagram of an image decoding apparatus according to an embodiment.
[0077] The image decoding apparatus 100 may include a receiver 110 and a decoder 120. The receiver 110 and the decoder 120 may include at least one processor.
[0078] In addition, the receiver 110 and the decoder 120 may include a memory that stores instructions to be run by at least one processor. The receiver 110 may receive a bitstream. The bitstream may include information obtained by image encoding performed by an image encoding apparatus 200 to be described below. In addition, the bitstream may be sent from the image encoding apparatus 200. The image decoding apparatus 100 may be connected to the image encoding apparatus 200 in a wired or wireless manner, and the receiver 110 may receive the bitstream in a wired or wireless manner. The receiver 110 may receive the bitstream from a storage medium such as an optical medium or a hard disk. The decoder 120 may reconstruct an image based on the information obtained from the received bitstream. The decoder 120 may obtain syntax elements for reconstructing the image from the bitstream.
[0079] The decoder 120 may reconstruct an image based on the syntax elements.
[0080] Reference will be made to Figure 2 to describe the operation of the image decoding apparatus 100 in more detail.
[0081] Figure 2 is a flowchart of an image decoding method according to an embodiment.
[0082] According to an embodiment of the present disclosure, the receiver 110 receives a bitstream. The image decoding device 100 may perform operation 210 of obtaining a binary string corresponding to the segmentation shape mode of the coding and decoding unit from the bitstream. Then, the image decoding device 100 may perform operation 220 of determining the segmentation rule of the coding and decoding unit. In addition, the image decoding device 100 may perform operation 230 of splitting the coding and decoding unit into a plurality of coding and decoding units based on at least one of the binary string corresponding to the segmentation shape mode and the segmentation rule. The image decoding device 100 may determine a first range according to the ratio of the height to the width of the coding and decoding unit to determine the segmentation rule, where the first range is the allowable size range of the coding and decoding unit.
[0083] The image decoding device 100 may determine a second range according to the segmentation shape mode of the coding and decoding unit to determine the segmentation rule, where the second range is the allowable size range of the coding and decoding unit.
[0084] Hereinafter, the segmentation of the coding and decoding unit will be described in detail according to an embodiment of the present disclosure. First, a picture may be segmented into one or more slices or one or more tiles. A slice or a tile may be a sequence of one or more maximum coding and decoding units (coding tree units (CTUs)).
[0085] As a concept compared with the maximum coding and decoding unit (CTU), there is a maximum coding and decoding block (coding tree block (CTB)). The maximum coding and decoding block (CTB) represents an N×N block including N×N samples (N is an integer).
[0086] Each color component may be segmented into one or more maximum coding and decoding blocks. When a picture has three sample arrays (sample arrays for the Y component, Cr component, and Cb component), the maximum coding and decoding unit (CTU) includes one maximum coding and decoding block of luminance samples, two corresponding maximum coding and decoding blocks of chrominance samples, and a syntax structure for encoding the luminance samples and chrominance samples. When the picture is a monochrome picture, the maximum coding and decoding unit includes one maximum coding and decoding block of monochrome samples and a syntax structure for encoding the monochrome samples.
[0087] When the picture is a picture encoded in a color plane separated according to color components, the maximum coding and decoding unit includes a syntax structure for encoding the picture and the samples of the picture.
[0088] A maximum coding / decoding block (CTB) can be divided into M×N coding / decoding blocks including M×N samples (M and N are integers). When a picture has sample arrays for Y component, Cr component, and Cb component, a coding unit (CU) includes one coding / decoding block of luminance samples, two corresponding coding / decoding blocks of chrominance samples, and syntax structures for encoding the luminance samples and the chrominance samples. When the picture is a monochrome picture, the coding unit includes one coding / decoding block of monochrome samples and syntax structures for encoding the monochrome samples.
[0089] When the picture is a picture encoded in color planes separated according to color components, the coding unit includes syntax structures for encoding the picture and the samples of the picture. As described above, the maximum coding / decoding block and the maximum coding unit are conceptually distinguished from each other, and the coding block and the coding unit are conceptually distinguished from each other. That is, the (maximum) coding unit refers to a data structure including the (maximum) coding block, and the (maximum) coding block includes corresponding samples and syntax structures corresponding to the (maximum) coding block.
[0090] However, since those of ordinary skill in the art understand that the (maximum) coding unit or the (maximum) coding block refers to a block of a specific size including a specific number of samples, the maximum coding / decoding block and the maximum coding unit or the coding block and the coding unit are referred to without distinction in the following description, unless otherwise described. An image can be divided into maximum coding units (CTUs). The size of each maximum coding unit can be determined based on information obtained from the bitstream. The shape of each maximum coding unit can be a square shape of the same size.
[0091] However, the embodiments are not limited thereto. For example, information about the maximum size of a luminance coding / decoding block can be obtained from the bitstream.
[0092] For example, the maximum size of the luminance coding / decoding block indicated by the information about the maximum size of the luminance coding / decoding block can be one of 4×4, 8×8, 16×16, 32×32, 64×64, 128×128, and 256×256. For example, information about the luminance block size difference and the maximum size of the luminance coding / decoding block that can be divided into two can be obtained from the bitstream. The information about the luminance block size difference can refer to the size difference between the maximum luminance coding unit and the maximum luminance coding block that can be divided into two. Accordingly, when the information about the maximum size of the luminance coding / decoding block that can be divided into two and the information about the luminance block size difference obtained from the bitstream are combined with each other, the size of the maximum luminance coding unit can be determined. The size of the maximum chrominance coding unit can be determined by using the size of the maximum luminance coding unit.
[0093] For example, when the ratio of color format Y:Cb:Cr is 4:2:0, the size of a chrominance block can be half the size of a luminance block, and the size of the largest chrominance coding / decoding unit can be half the size of the largest luminance coding / decoding unit. According to an embodiment, since information on the maximum size of a luminance coding / decoding block that can be binary split is obtained from a bitstream, the maximum size of a luminance coding / decoding block that can be binary split can be variably determined. In contrast, the maximum size of a luminance coding / decoding block that can be ternary split can be fixed.
[0094] For example, the maximum size of a luminance coding / decoding block that can be ternary split in an I picture can be 32×32, and the maximum size of a luminance coding / decoding block that can be ternary split in a P picture or a B picture can be 64×64. In addition, the largest coding / decoding unit can be hierarchically split into coding / decoding units based on split shape mode information obtained from the bitstream.
[0095] At least one of information indicating whether to perform quad split, information indicating whether to perform multi split, split direction information, and split type information can be obtained from the bitstream as split shape mode information.
[0096] For example, the information indicating whether to perform quad split can indicate whether the current coding / decoding unit is quad split (QUAD_SPLIT).
[0097] When the current coding / decoding unit is not quad split, the information indicating whether to perform multi split can indicate whether the current coding / decoding unit is no longer split (NO_SPLIT) or is binary split / ternary split.
[0098] When the current coding / decoding unit is binary split or ternary split, the split direction information indicates that the current coding / decoding unit is split in one of the horizontal and vertical directions.
[0099] When the current coding / decoding unit is split in the horizontal or vertical direction, the split type information indicates whether the current coding / decoding unit is binary split or ternary split. The split mode of the current coding / decoding unit can be determined according to the split direction information and the split type information.
[0100] The splitting mode when performing binary splitting on the current codec unit in the horizontal direction can be determined as the binary horizontal splitting mode (SPLIT_BT_HOR), the splitting mode when performing ternary splitting on the current codec unit in the horizontal direction can be determined as the ternary horizontal splitting mode (SPLIT_TT_HOR), the splitting mode when performing binary splitting on the current codec unit in the vertical direction can be determined as the binary vertical splitting mode (SPLIT_BT_VER), and the splitting mode when performing ternary splitting on the current codec unit in the vertical direction can be determined as the ternary vertical splitting mode (SPLIT_TT_VER). The image decoding device 100 can obtain splitting shape mode information from a binary string in the bitstream. The form of the bitstream received by the image decoding device 100 can include fixed-length binary codes, unary codes, truncated unary codes, predetermined binary codes, etc. A binary string is information in the form of binary digits. A binary string can include at least one bit. The image decoding device 100 can obtain splitting shape mode information corresponding to the binary string based on splitting rules.
[0101] The image decoding device 100 can determine whether to perform quaternary splitting on the codec unit, whether to split the codec unit, the splitting direction, and the splitting type based on a binary string. The codec unit can be smaller than the maximum codec unit or the same as the maximum codec unit. For example, since the maximum codec unit is the codec unit with the largest size, the maximum codec unit is one of the codec units. When the splitting shape mode information about the maximum codec unit indicates that no splitting is performed, the size of the codec unit determined in the maximum codec unit is the same as the maximum codec unit. When the splitting shape mode information about the maximum codec unit indicates that splitting is performed, the maximum codec unit can be split into codec units. In addition, when the splitting shape mode information about the codec unit indicates that splitting is performed, the codec unit can be split into smaller codec units. However, the splitting of the image is not limited to this, and the maximum codec unit and the codec unit may not be distinguished.
[0102] Reference will be made to Figures 3 to 16 The splitting of the codec unit will be described in more detail. In addition, one or more prediction blocks for prediction can be determined from the codec unit. The prediction block can be the same as or smaller than the codec unit. In addition, one or more transform blocks for transformation can be determined from the codec unit.
[0103] The transform block can be equal to or smaller than the codec unit.
[0104] The shapes and sizes of the transform block and the prediction block can be independent of each other. In another embodiment, prediction can be performed by using the codec unit as a prediction unit.
[0105] In addition, the transformation can be performed by using a codec unit as a transform block. The current block and adjacent blocks of the present disclosure can indicate one of a maximum codec unit, a codec unit, a prediction block, and a transform block. In addition, the current block or the current codec unit is the block that is currently being decoded or encoded or the block that is currently being split. An adjacent block can be a block reconstructed before the current block. The adjacent block can be adjacent to the current block spatially or temporally.
[0106] The adjacent block can be located at one of the lower left, left, upper left, upper, upper right, right, and lower right of the current block.
[0107] Figure 3 The process of determining at least one codec unit by splitting the current codec unit by an image decoding apparatus according to an embodiment is shown. The block shape can include 4N×4N, 4N×2N, 2N×4N, 4N×N, N×4N, 32N×N, N×32N, 16N×N, N×16N, 8N×N, or N×8N. Here, N can be a positive integer.
[0108] The block shape information is information indicating at least one of the shape, direction, width-to-height ratio, and size of the codec unit. The shape of the codec unit can include a square and a non-square. When the width and height of the codec unit are the same (i.e., when the block shape of the codec unit is 4N×4N), the image decoding apparatus 100 can determine the block shape information of the codec unit as a square.
[0109] The image decoding apparatus 100 can determine the shape of the codec unit as a non-square. When the width and height of the codec unit are different from each other (i.e., when the block shape of the codec unit is 4N×2N, 2N×4N, 4N×N, N×4N, 32N×N, N×32N, 16N×N, N×16N, 8N×N, or N×8N), the image decoding apparatus 100 can determine the block shape information of the codec unit as a non-square. When the shape of the codec unit is non-square, the image decoding apparatus 100 can determine the width-to-height ratio in the block shape information of the codec unit as at least one of 1:2, 2:1, 1:4, 4:1, 1:8, 8:1, 1:16, 16:1, 1:32, and 32:1. In addition, the image decoding apparatus 100 can determine whether the codec unit is in the horizontal direction or the vertical direction based on the width length and height length of the codec unit.
[0110] In addition, the image decoding apparatus 100 may determine the size of a codec unit based on at least one of the width length, height length, and area of the codec unit. According to an embodiment, the image decoding apparatus 100 may determine the shape of a codec unit by using block shape information, and may determine a method of dividing the codec unit by using split shape mode information.
[0111] That is, the method of dividing the codec unit indicated by the split shape mode information may be determined based on the block shape indicated by the block shape information used by the image decoding apparatus 100. The image decoding apparatus 100 may obtain the split shape mode information from a bitstream. However, the embodiment is not limited thereto, and the image decoding apparatus 100 and the image encoding apparatus 200 may determine pre-agreed split shape mode information based on the block shape information. The image decoding apparatus 100 may determine pre-agreed split shape mode information for a maximum codec unit or a minimum codec unit. For example, the image decoding apparatus 100 may determine the split shape mode information for the maximum codec unit as quad-tree splitting. In addition, the image decoding apparatus 100 may determine the split shape mode information for the minimum codec unit as "no splitting". Specifically, the image decoding apparatus 100 may determine the size of the maximum codec unit as 256×256. The image decoding apparatus 100 may determine the pre-agreed split shape mode information as quad-tree splitting. Quad-tree splitting is a split shape mode in which both the width and height of a codec unit are bisected. The image decoding apparatus 100 may obtain a codec unit of size 128×128 from the maximum codec unit of size 256×256 based on the split shape mode information. In addition, the image decoding apparatus 100 may determine the size of the minimum codec unit as 4×4.
[0112] The image decoding apparatus 100 may obtain split shape mode information indicating "no splitting" for the minimum codec unit. According to an embodiment, the image decoding apparatus 100 may use block shape information indicating that the current codec unit has a square shape. For example, the image decoding apparatus 100 may determine whether to split a square codec unit, whether to vertically split a square codec unit, whether to horizontally split a square codec unit, or whether to divide a square codec unit into four codec units based on the split shape mode information.
[0113] Reference Figure 3 , when the block shape information of the current codec unit 300 indicates a square shape, the decoder 120 may not split the codec unit 310a having the same size as the current codec unit 300 based on the split shape mode information indicating no splitting, or may determine codec units 310b, 310c, 310d, 310e, or 310f split based on the split shape mode information indicating a specific splitting method. ReferenceFigure 3 , according to an embodiment, the image decoding device 100 may determine two codec units 310b obtained by vertically dividing the current codec unit 300 based on the segmentation shape pattern information indicating vertical segmentation. The image decoding device 100 may determine two codec units 310c obtained by horizontally dividing the current codec unit 300 based on the segmentation shape pattern information indicating horizontal segmentation. The image decoding device 100 may determine four codec units 310d obtained by vertically and horizontally dividing the current codec unit 300 based on the segmentation shape pattern information indicating vertical and horizontal segmentation. According to an embodiment, the image decoding device 100 may determine three codec units 310e obtained by vertically dividing the current codec unit 300 based on the segmentation shape pattern information indicating ternary segmentation in the vertical direction. The image decoding device 100 may determine three codec units 310f obtained by horizontally dividing the current codec unit 300 based on the segmentation shape pattern information indicating ternary segmentation in the horizontal direction. However, the method for dividing the square codec unit is not limited to the above method, and the segmentation shape pattern information may indicate various methods.
[0114] Specific methods for dividing the square codec unit will be described in detail through various embodiments below.
[0115] Figure 4 The process of the image decoding device according to an embodiment determining at least one codec unit by dividing a non-square codec unit is shown. According to an embodiment, the image decoding device 100 may use the block shape information indicating that the current codec unit has a non-square shape. The image decoding device 100 may determine whether to divide the non-square current codec unit or whether to divide the non-square current codec unit by using a specific division method based on the segmentation shape pattern information. Refer to Figure 4 , when the block shape information of the current codec unit 400 or 450 indicates a non-square shape, the image decoding device 100 may determine the codec unit 410 or 460 having the same size as the current codec unit 400 or 450 based on the segmentation shape pattern information indicating no division, or may determine the codec units 420a and 420b, 430a to 430c, 470a and 470b, or 480a to 480c divided based on the segmentation shape pattern information indicating a specific division method.
[0116] A specific method for splitting a non-square codec unit will be described in detail through various embodiments. According to an embodiment, the image decoding device 100 may determine a method for splitting a codec unit by using split shape mode information, and in this case, the split shape mode information may indicate the number of one or more codec units generated by splitting the codec unit.
[0117] Referring to Figure 4 , when the split shape mode information indicates that the current codec unit 400 or 450 is split into two codec units, the image decoding device 100 may determine two codec units 420a and 420b or 470a and 470b included in the current codec unit 400 or 450 by splitting the current codec unit 400 or 450 based on the split shape mode information. According to an embodiment, when the image decoding device 100 splits the non-square current codec unit 400 or 450 based on the split shape mode information, the image decoding device 100 may consider the long side position of the non-square current codec unit 400 or 450 to split the current codec unit.
[0118] For example, considering the shape of the current codec unit 400 or 450, the image decoding device 100 may determine a plurality of codec units by splitting the current codec unit 400 or 450 in the long side direction of the current codec unit 400 or 450. According to an embodiment, when the split shape mode information indicates that the codec unit is split (ternary split) into an odd number of blocks, the image decoding device 100 may determine an odd number of codec units included in the current codec unit 400 or 450.
[0119] For example, when the split shape mode information indicates that the current codec unit 400 or 450 is to be split into three codec units, the image decoding device 100 may split the current codec unit 400 or 450 into three codec units 430a, 430b, and 430c or 480a, 480b, and 480c. According to an embodiment, the ratio of the width to the height of the current codec unit 400 or 450 may be 4:1 or 1:4. When the ratio of the width to the height is 4:1, the block shape information may indicate the horizontal direction because the width length is longer than the height length. When the ratio of the width to the height is 1:4, the block shape information may indicate the vertical direction because the width length is shorter than the height length. The image decoding device 100 may determine to split the current codec unit into an odd number of blocks based on the split shape mode information. In addition, the image decoding device 100 may determine the split direction of the current codec unit 400 or 450 based on the block shape information of the current codec unit 400 or 450. For example, when the current codec unit 400 is in the vertical direction, the image decoding device 100 may determine the codec units 430a, 430b, and 430c by splitting the current codec unit 400 in the horizontal direction.
[0120] In addition, when the current codec unit 450 is in the horizontal direction, the image decoding device 100 may determine the codec units 480a, 480b, and 480c by splitting the current codec unit 450 in the vertical direction. According to an embodiment, the image decoding device 100 may determine an odd number of codec units included in the current codec unit 400 or 450, and not all of the determined codec units may have the same size. For example, the size of a specific codec unit 430b or 480b among the determined odd number of codec units 430a, 430b, and 430c or 480a, 480b, and 480c may be different from the sizes of the other codec units 430a and 430c or 480a and 480c.
[0121] That is, the codec units that can be determined by splitting the current codec unit 400 or 450 may have various sizes, and in some cases, all of the odd number of codec units 430a, 430b, and 430c or 480a, 480b, and 480c may have different sizes. According to an embodiment, when the split shape mode information indicates that the codec unit is to be split into an odd number of blocks, the image decoding device 100 may determine an odd number of codec units included in the current codec unit 400 or 450, and moreover, a specific constraint may be imposed on at least one of the odd number of codec units generated by splitting the current codec unit 400 or 450. Refer to Figure 4, the image decoding device 100 may set the decoding process of the central codec unit 430b or 480b among the three codec units 430a, 430b, and 430c or 480a, 480b, and 480c generated when splitting the current codec unit 400 or 450 to be different from the decoding processes of the other codec units 430a and 430c or 480a and 480c.
[0122] For example, different from the other codec units 430a and 430c or 480a and 480c, the image decoding device 100 may constrain the codec unit 430b or 480b at the central position from being split any further or only being split a specific number of times.
[0123] Figure 5 The process of splitting a codec unit by an image decoding device based on at least one of block shape information and split shape mode information according to an embodiment is shown. According to an embodiment, the image decoding device 100 may determine whether to split a square first codec unit 500 into codec units based on at least one of block shape information and split shape mode information. According to an embodiment, when the split shape mode information indicates splitting the first codec unit 500 in the horizontal direction, the image decoding device 100 may determine a second codec unit 510 by splitting the first codec unit 500 in the horizontal direction. The first codec unit, second codec unit, and third codec unit used according to an embodiment are terms for understanding the relationship before and after splitting a codec unit. For example, the second codec unit may be determined by splitting the first codec unit, and the third codec unit may be determined by splitting the second codec unit.
[0124] It will be understood that the relationship among the first codec unit, second codec unit, and third codec unit follows the above description. According to an embodiment, the image decoding device 100 may determine whether to split the determined second codec unit 510 into codec units based on the split shape mode information. Refer to Figure 5, the image decoding apparatus 100 may divide or not divide the non-square second decoded unit 510 determined by dividing the first decoded unit 500 into one or more third decoded units 520a, 520b, 520c, and 520d based on the segmentation shape mode information. The image decoding apparatus 100 may obtain the segmentation shape mode information, and may obtain a plurality of second decoded units (e.g., 510) of various shapes by dividing the first decoded unit 500 based on the obtained segmentation shape mode information, and may divide the second decoded unit 510 by using the division method of the first decoded unit 500 based on the segmentation shape mode information. According to an embodiment, when the first decoded unit 500 is divided into the second decoded unit 510 based on the segmentation shape mode information of the first decoded unit 500, the second decoded unit 510 may be divided into third decoded units (e.g., 520a or 520b, 520c, and 520d) based on the segmentation shape mode information of the second decoded unit 510. That is, the decoded units may be recursively divided based on the segmentation shape mode information of each decoded unit.
[0125] Accordingly, a square decoded unit may be determined by dividing a non-square decoded unit, and a non-square decoded unit may be determined by recursively dividing a square decoded unit. Refer to Figure 5 , among the odd-numbered third decoded units 520b, 520c, and 520d determined by dividing the non-square second decoded unit 510, a specific decoded unit (e.g., a decoded unit located at the center position or a square decoded unit) may be recursively divided. According to an embodiment, the non-square third decoded unit 520b among the odd-numbered third decoded units 520b, 520c, and 520d may be divided into a plurality of fourth decoded units in the horizontal direction. Among the plurality of fourth decoded units 530a, 530b, 530c, and 530d, the non-square fourth decoded unit 530b or 530d may be re-divided into a plurality of decoded units. For example, the non-square fourth decoded unit 530b or 530d may be re-divided into an odd number of decoded units.
[0126] A method for recursively partitioning a codec unit will be described below through various embodiments. According to an embodiment, the image decoding device 100 may partition each of the third codec units 520a, 520b, 520c, and 520d into codec units based on the partition shape mode information. In addition, the image decoding device 100 may determine not to partition the second codec unit 510 based on the partition shape mode information. According to an embodiment, the image decoding device 100 may partition the non-square second codec unit 510 into an odd number of third codec units 520b, 520c, and 520d. The image decoding device 100 may impose a specific constraint on a specific third codec unit among the odd number of third codec units 520b, 520c, and 520d.
[0127] For example, the image decoding device 100 may constrain the third codec unit 520c at the center position among the odd number of third codec units 520b, 520c, and 520d to no longer be partitioned or be partitioned a set number of times. Refer to Figure 5 , the image decoding device 100 may constrain the third codec unit 520c at the center position among the odd number of third codec units 520b, 520c, and 520d included in the non-square second codec unit 510 to no longer be partitioned, be partitioned by using a specific partitioning method (e.g., only partitioned into four codec units or partitioned by using the partitioning method of the second codec unit 510), or only be partitioned a specific number of times (e.g., only partitioned n times (where n > 0)).
[0128] However, the constraint on the third codec unit 520c at the center position is not limited to the above examples and may include various constraints for decoding the third codec unit 520c at the center position differently from the other third codec units 520b and 520d.
[0129] According to an embodiment, the image decoding device 100 may obtain partition shape mode information for partitioning the current codec unit from a specific position in the current codec unit.
[0130] Figure 6 A method for the image decoding device according to an embodiment to determine a specific codec unit among an odd number of codec units is shown. Refer to Figure 6 , the partition shape mode information of the current codec unit 600 or 650 may be obtained from a sample at a specific position (e.g., the sample 640 or 690 at the center position) among the plurality of samples included in the current codec unit 600 or 650. However, the specific position in the current codec unit 600 from which at least one piece of partition shape mode information can be obtained is not limited to Figure 6at the central position in, and may include various positions included in the current codec unit 600 (e.g., upper position, lower position, left position, right position, upper left position, lower left position, upper right position, lower right position, etc.).
[0131] The image decoding apparatus 100 may obtain the segmentation shape pattern information from a specific position, and may determine whether to segment the current codec unit into codec units of various shapes and various sizes. According to an embodiment, when the current codec unit is segmented into a specific number of codec units, the image decoding apparatus 100 may select one of the codec units.
[0132] Various methods may be used to select one of the multiple codec units, as will be described through various embodiments below.
[0133] According to an embodiment, the image decoding apparatus 100 may segment the current codec unit into multiple codec units, and may determine the codec unit at a specific position. According to an embodiment, the image decoding apparatus 100 may use the information indicating the positions of an odd number of codec units to determine the codec unit at the central position among the odd number of codec units. Referring to Figure 6 , the image decoding apparatus 100 may determine the odd number of codec units 620a, 620b, and 620c or the odd number of codec units 660a, 660b, and 660c by segmenting the current codec unit 600 or the current codec unit 650. The image decoding apparatus 100 may determine the middle codec unit 620b or the middle codec unit 660b by using the information about the positions of the odd number of codec units 620a, 620b, and 620c or the odd number of codec units 660a, 660b, and 660c. For example, the image decoding apparatus 100 may determine the codec unit 620b at the central position by determining the positions of the codec units 620a, 620b, and 620c based on the information indicating the positions of specific samples included in the codec units 620a, 620b, and 620c.
[0134] Specifically, the image decoding apparatus 100 may determine the codec unit 620b at the center position by determining the positions of the codec units 620a, 620b, and 620c based on the information indicating the positions of the upper-left samples 630a, 630b, and 630c of the codec units 620a, 620b, and 620c. According to an embodiment, the information indicating the positions of the upper-left samples 630a, 630b, and 630c included in the codec units 620a, 620b, and 620c respectively may include information about the positions or coordinates of the codec units 620a, 620b, and 620c in the picture. According to an embodiment, the information indicating the positions of the upper-left samples 630a, 630b, and 630c included in the codec units 620a, 620b, and 620c respectively may include information indicating the width or height of the codec units 620a, 620b, and 620c included in the current codec unit 600, and the width or height may correspond to the information indicating the difference between the coordinates of the codec units 620a, 620b, and 620c in the picture.
[0135] That is, by directly using the information about the positions or coordinates of the codec units 620a, 620b, and 620c in the picture, or by using the information about the width or height of the codec units corresponding to the difference between the coordinates, the image decoding apparatus 100 may determine the codec unit 620b at the center position. According to an embodiment, the information indicating the position of the upper-left sample 630a of the upper codec unit 620a may include the coordinates (xa, ya), the information indicating the position of the upper-left sample 630b of the middle codec unit 620b may include the coordinates (xb, yb), and the information indicating the position of the upper-left sample 630c of the lower codec unit 620c may include the coordinates (xc, yc). The image decoding apparatus 100 may determine the middle codec unit 620b by using the coordinates of the upper-left samples 630a, 630b, and 630c included in the codec units 620a, 620b, and 620c respectively. For example, when the coordinates of the upper-left samples 630a, 630b, and 630c are sorted in ascending or descending order, the codec unit 620b including the coordinates (xb, yb) of the sample 630b at the center position may be determined as the codec unit at the center position among the codec units 620a, 620b, and 620c determined by dividing the current codec unit 600. However, the coordinates indicating the positions of the upper-left samples 630a, 630b, and 630c may include the coordinates indicating the absolute position in the picture, or may use the coordinates (dxb, dyb) indicating the relative position of the upper-left sample 630b of the middle codec unit 620b with respect to the position of the upper-left sample 630a of the upper codec unit 620a and the coordinates (dxc, dyc) indicating the relative position of the upper-left sample 630c of the lower codec unit 620c.
[0136] The method of determining a codec unit at a specific position by using the coordinates of samples included in a codec unit as information indicating the positions of the samples is not limited to the above method and may include various arithmetic methods capable of using sample coordinates. According to an embodiment, the image decoding apparatus 100 may divide a current codec unit 600 into a plurality of codec units 620a, 620b, and 620c, and may select one of the codec units 620a, 620b, and 620c based on a specific criterion.
[0137] For example, the image decoding apparatus 100 may select a codec unit 620b having a size different from that of other codec units from among the codec units 620a, 620b, and 620c. According to an embodiment, the image decoding apparatus 100 may determine the width or height of each of the codec units 620a, 620b, and 620c by using the coordinates (xa, ya) as information indicating the position of the upper-left sample 630a of the upper codec unit 620a, the coordinates (xb, yb) as information indicating the position of the upper-left sample 630b of the middle codec unit 620b, and the coordinates (xc, yc) as information indicating the position of the upper-left sample 630c of the lower codec unit. The image decoding apparatus 100 may determine the respective sizes of the codec units 620a, 620b, and 620c by using the coordinates (xa, ya), (xb, yb), and (xc, yc) indicating the positions of the codec units 620a, 620b, and 620c. According to an embodiment, the image decoding apparatus 100 may determine the width of the upper codec unit 620a as the width of the current codec unit 600. The image decoding apparatus 100 may determine the height of the upper codec unit 620a as yb - ya. According to an embodiment, the image decoding apparatus 100 may determine the width of the middle codec unit 620b as the width of the current codec unit 600. The image decoding apparatus 100 may determine the height of the middle codec unit 620b as yc - yb. According to an embodiment, the image decoding apparatus 100 may determine the width or height of the lower codec unit 620c by using the width or height of the current codec unit 600 and the widths or heights of the upper codec unit 620a and the middle codec unit 620b. The image decoding apparatus 100 may determine a codec unit having a size different from that of other codec units based on the determined widths and heights of the codec units 620a, 620b, and 620c. Refer Figure 6 , the image decoding apparatus 100 may determine the middle codec unit 620b having a size different from that of the upper codec unit 620a and the lower codec unit 620c as the codec unit at the specific position.
[0138] However, the above process in which the image decoding device 100 determines a codec unit having a size different from that of other codec units merely corresponds to an example of determining a codec unit at a specific position by using the size of a codec unit determined based on sample coordinates. Therefore, various processes of determining a codec unit at a specific position by comparing the sizes of codec units determined based on the coordinates of a specific sample can be used. The image decoding device 100 may determine the width or height of each of the codec units 660a, 660b, and 660c by using the coordinate (xd, yd) which is information indicating the position of the upper left sample 670a of the left codec unit 660a, the coordinate (xe, ye) which is information indicating the position of the upper left sample 670b of the middle codec unit 660b, and the coordinate (xf, yf) which is information indicating the position of the upper left sample 670c of the right codec unit 660c.
[0139] The image decoding device 100 may determine the respective sizes of the codec units 660a, 660b, and 660c by using the coordinates (xd, yd), (xe, ye), and (xf, yf) indicating the positions of the codec units 660a, 660b, and 660c. According to an embodiment, the image decoding device 100 may determine the width of the left codec unit 660a as xe - xd. The image decoding device 100 may determine the height of the left codec unit 660a as the height of the current codec unit 650. According to an embodiment, the image decoding device 100 may determine the width of the middle codec unit 660b as xf - xe. The image decoding device 100 may determine the height of the middle codec unit 660b as the height of the current codec unit 650. According to an embodiment, the image decoding device 100 may determine the width or height of the right codec unit 660c by using the width or height of the current codec unit 650 and the widths or heights of the left codec unit 660a and the middle codec unit 660b. The image decoding device 100 may determine a codec unit having a size different from that of other codec units based on the determined widths and heights of the codec units 660a, 660b, and 660c. Referring Figure 6 , the image decoding device 100 may determine the middle codec unit 660b having a size different from that of the left codec unit 660a and the right codec unit 660c as the codec unit at a specific position.
[0140] However, the above process in which the image decoding device 100 determines a codec unit having a size different from that of other codec units merely corresponds to an example of determining a codec unit at a specific position by using the size of a codec unit determined based on sample coordinates. Therefore, various processes of determining a codec unit at a specific position by comparing the sizes of codec units determined based on the coordinates of a specific sample can be used.
[0141] However, the position of the sample considered for determining the position of the codec unit is not limited to the above-mentioned upper left position, and information about an arbitrary position of the sample included in the codec unit may be used. According to an embodiment, by considering the shape of the current codec unit, the image decoding device 100 may select a codec unit at a specific position from an odd number of codec units determined by dividing the current codec unit. For example, when the current codec unit has a non-square shape in which the width is longer than the height, the image decoding device 100 may determine the codec unit at a specific position in the horizontal direction. That is, the image decoding device 100 may determine one of the codec units at different positions in the horizontal direction, and may impose constraints on the codec unit. When the current codec unit has a non-square shape in which the height is longer than the width, the image decoding device 100 may determine the codec unit at a specific position in the vertical direction.
[0142] That is, the image decoding device 100 may determine one of the codec units at different positions in the vertical direction, and may impose constraints on the codec units. According to an embodiment, the image decoding device 100 may determine the codec unit at a specific position from among the even-numbered codec units using information indicating the respective positions of the even-numbered codec units. The image decoding device 100 may determine the even-numbered codec units by splitting (binary splitting) the current codec unit, and may determine the codec unit at a specific position by using information about the positions of the even-numbered codec units.
[0143] The operations related to it can correspond to those already referenced above Figure 6 The operation of determining the codec unit at a specific position (e.g., the center position) from among the odd number of codec units is described in detail, so a detailed description thereof will be omitted. According to an embodiment, when a non-square current codec unit is divided into a plurality of codec units, specific information about the codec unit at a specific position may be used in the division operation to determine the codec unit at the specific position from among the plurality of codec units.
[0144] For example, the image decoding device 100 may use at least one of the block shape information and the partition shape mode information stored in the sample included in the intermediate codec unit in the partition operation to determine the codec unit at the center position from among the plurality of codec units determined by partitioning the current codec unit. Figure 6, the image decoding apparatus 100 may divide a current codec unit 600 into a plurality of codec units 620a, 620b, and 620c based on split shape mode information, and may determine the codec unit 620b at the center position among the plurality of codec units 620a, 620b, and 620c. Further, the image decoding apparatus 100 may determine the codec unit 620b at the center position in consideration of the position from which the split shape mode information is obtained. That is, the split shape mode information of the current codec unit 600 may be obtained from the sample 640 at the center position of the current codec unit 600, and when the current codec unit 600 is divided into a plurality of codec units 620a, 620b, and 620c based on the split shape mode information, the codec unit 620b including the sample 640 may be determined as the codec unit at the center position.
[0145] However, the information for determining the codec unit at the center position is not limited to the split shape mode information, and various types of information may be used to determine the codec unit at the center position. According to an embodiment, specific information for identifying the codec unit at a specific position may be obtained from a specific sample included in the codec unit to be determined. Refer to Figure 6 , the image decoding apparatus 100 may use the split shape mode information obtained from the sample at a specific position (e.g., the sample at the center position of the current codec unit 600) in the current codec unit 600 to determine the codec unit at a specific position (e.g., the codec unit at the center position among the plurality of split codec units) among the plurality of codec units 620a, 620b, and 620c determined by splitting the current codec unit 600. That is, the image decoding apparatus 100 may determine the sample at a specific position by considering the block shape of the current codec unit 600, may determine the codec unit 620b including the sample from which specific information (e.g., split shape mode information) can be obtained among the plurality of codec units 620a, 620b, and 620c determined by splitting the current codec unit 600, and may impose a specific constraint on the codec unit 620b. Refer to Figure 6 , according to an embodiment, during the decoding process, the image decoding apparatus 100 may determine the sample 640 at the center position of the current codec unit 600 as the sample from which specific information can be obtained, and may impose a specific constraint on the codec unit 620b including the sample 640.
[0146] However, the positions of samples from which specific information can be obtained are not limited to the above positions, and may include any position of the samples included in the codec unit 620b to be determined for constraint. According to an embodiment, the positions of samples from which specific information can be obtained may be determined based on the shape of the current codec unit 600. According to an embodiment, the block shape information may indicate whether the current codec unit has a square shape or a non-square shape, and the positions of samples from which specific information can be obtained may be determined based on the shape. For example, by using at least one of the information on the width of the current codec unit and the information on the height of the current codec unit, the image decoding device 100 may determine a sample located at a boundary for dividing at least one of the width and height of the current codec unit into two halves as a sample from which specific information can be obtained.
[0147] As another example, when the block shape information of the current codec unit indicates a non-square shape, the image decoding device 100 may determine one of the samples adjacent to the boundary for dividing the long side of the current codec unit into two halves as a sample from which specific information can be obtained. According to an embodiment, when the current codec unit is divided into a plurality of codec units, the image decoding device 100 may use the split shape mode information to determine a codec unit at a specific position among the plurality of codec units. According to an embodiment, the image decoding device 100 may obtain the split shape mode information from the samples at specific positions in the codec unit, and may divide the plurality of codec units generated by splitting the current codec unit by using the split shape mode information obtained from the samples at specific positions in each of the plurality of codec units. That is, the codec unit may be recursively divided based on the split shape mode information obtained from the samples at specific positions in each codec unit.
[0148] The above has been referred to Figure 5 The process of recursively dividing the codec unit has been described, so the detailed description thereof will be omitted.
[0149] According to an embodiment, the image decoding device 100 may determine one or more codec units by splitting the current codec unit, and may determine the order of decoding the one or more codec units based on a specific block (e.g., the current codec unit).
[0150] Figure 7 FIG. shows the order of processing a plurality of codec units when the image decoding device determines a plurality of codec units by splitting the current codec unit according to an embodiment.
[0151] According to an embodiment, based on the segmentation shape pattern information, the image decoding apparatus 100 may determine second decoding units 710a and 710b by dividing a first decoding unit 700 in a vertical direction, may determine second decoding units 730a and 730b by dividing the first decoding unit 700 in a horizontal direction, or may determine second decoding units 750a, 750b, 750c, and 750d by dividing the first decoding unit 700 in both the vertical and horizontal directions. Refer to Figure 7 , the image decoding apparatus 100 may determine to process the second decoding units 710a and 710b determined by dividing the first decoding unit 700 in the vertical direction in a horizontal direction order 710c. The image decoding apparatus 100 may determine to process the second decoding units 730a and 730b determined by dividing the first decoding unit 700 in the horizontal direction in a vertical direction order 730c.
[0152] The image decoding apparatus 100 may determine the second decoding units 750a, 750b, 750c, and 750d determined by dividing the first decoding unit 700 in both the vertical and horizontal directions according to a specific order (e.g., raster scan order or Z-scan order 750e) of processing the decoding units in one row and then processing the decoding units in the next row. According to an embodiment, the image decoding apparatus 100 may recursively divide the decoding units. Refer to Figure 7 , the image decoding apparatus 100 may determine a plurality of decoding units 710a and 710b, 730a and 730b, or 750a, 750b, 750c, and 750d by dividing the first decoding unit 700, and may recursively divide each of the determined plurality of decoding units 710a, 710b, 730a, 730b, 750a, 750b, 750c, and 750d. The dividing method of the plurality of decoding units 710a and 710b, 730a and 730b, or 750a, 750b, 750c, and 750d may correspond to the dividing method of the first decoding unit 700. Accordingly, each of the plurality of decoding units 710a and 710b, 730a and 730b, or 750a, 750b, 750c, and 750d may be independently divided into a plurality of decoding units.
[0153] Refer to Figure 7 , the image decoding apparatus 100 may determine the second decoding units 710a and 710b by dividing the first decoding unit 700 in the vertical direction, and may determine to independently divide or not divide each of the second decoding units 710a and 710b.
[0154] According to an embodiment, the image decoding apparatus 100 may determine third decoding units 720a and 720b by dividing the second decoding unit 710a on the left side in the horizontal direction, and may not divide the second decoding unit 710b on the right side. According to an embodiment, the processing order of the coding / decoding units may be determined based on the process of dividing the coding / decoding units. In other words, the processing order of the divided coding / decoding units may be determined based on the processing order of the coding / decoding unit immediately before being divided. The image decoding apparatus 100 may determine the processing order of the third decoding units 720a and 720b determined by dividing the second decoding unit 710a on the left side independently of the second decoding unit 710b on the right side. Since the third decoding units 720a and 720b are determined by dividing the second decoding unit 710a on the left side in the horizontal direction, the third decoding units 720a and 720b may be processed in the vertical direction order 720c. Since the second decoding unit 710a on the left side and the second decoding unit 710b on the right side are processed in the horizontal direction order 710c, the second decoding unit 710b on the right side may be processed after the third decoding units 720a and 720b included in the second decoding unit 710a on the left side are processed in the vertical direction order 720c.
[0155] The process of determining the processing order of the coding / decoding units based on the coding / decoding unit before being divided is not limited to the above example, and various methods may be used to independently process the divided coding / decoding units determined to be in various shapes in a specific order.
[0156] Figure 8 The process in which the image decoding apparatus determines that the current coding / decoding unit will be divided into an odd number of coding / decoding units when the coding / decoding units cannot be processed in a specific order according to an embodiment is shown. According to an embodiment, the image decoding apparatus 100 may determine that the current coding / decoding unit will be divided into an odd number of coding / decoding units based on the obtained division shape pattern information. Refer to Figure 8 , the first square coding / decoding unit 800 may be divided into non-square second coding / decoding units 810a and 810b, and the second coding / decoding units 810a and 810b may be independently divided into third coding / decoding units 820a and 820b and 820c, 820d, and 820e.
[0157] According to an embodiment, the image decoding apparatus 100 may determine a plurality of third decoding units 820a and 820b by dividing the second decoding unit 810a on the left side in the horizontal direction, and may divide the second decoding unit 810b on the right side into an odd number of third decoding units 820c, 820d, and 820e. According to an embodiment, the image decoding apparatus 100 may determine whether any decoding unit is divided into an odd number of decoding units by determining whether the third decoding units 820a and 820b and 820c, 820d, and 820e can be processed in a specific order. Refer to Figure 8 , the image decoding apparatus 100 may determine the third decoding units 820a and 820b and 820c, 820d, and 820e by recursively dividing the first decoding unit 800. The image decoding apparatus 100 may determine whether any one of the first decoding unit 800, the second decoding units 810a and 810b, or the third decoding units 820a and 820b and 820c, 820d, and 820e is divided into an odd number of decoding units based on at least one of the block shape information and the division shape pattern information. For example, the right decoding unit among the second decoding units 810a and 810b may be divided into an odd number of third decoding units 820c, 820d, and 820e.
[0158] The processing order of the multiple codec units included in the first codec unit 800 may be a specific order (e.g., the Z-scan order 830), and the image decoding device 100 may determine whether the third codec units 820c, 820d, and 820e determined by dividing the second codec unit 810b on the right side into an odd number of codec units satisfy the conditions for processing in a specific order. According to an embodiment, the image decoding device 100 may determine whether the third codec units 820a, 820b, 820c, 820d, and 820e included in the first codec unit 800 satisfy the conditions for processing in a specific order, and the condition relates to whether at least one of the width and height of the second codec units 810a and 810b will be divided in half along the boundaries of the third codec units 820a, 820b, 820c, 820d, and 820e. For example, the third codec units 820a and 820b determined when the height of the non-square-shaped second codec unit 810a on the left side is divided in half may satisfy the condition. It may be determined that the third codec units 820c, 820d, and 820e do not satisfy the condition because the boundaries of the third codec units 820c, 820d, and 820e determined when the second codec unit 810b on the right side is divided into three codec units cannot divide the width or height of the second codec unit 810b on the right side in half. When the above-described condition is not satisfied, the image decoding device 100 may decide to break the scan order and may determine that the second codec unit 810b on the right side will be divided into an odd number of codec units based on the determined result.
[0159] According to an embodiment, when a codec unit is divided into an odd number of codec units, the image decoding device 100 may impose specific constraints on the codec units at specific positions among the divided codec units.
[0160] The constraints or specific positions have been described above through various embodiments, so the detailed description thereof will be omitted. Figure 9 The process of an image decoding device according to an embodiment for determining at least one codec unit by dividing a first codec unit is shown. According to an embodiment, the image decoding device 100 may divide the first codec unit 900 based on the division shape mode information obtained by the receiver 110. The square first codec unit 900 may be divided into four square codec units or may be divided into multiple non-square codec units.
[0161] For example, referring to Figure 9, when the first codec unit 900 has a square shape and the split shape mode information indicates that the first codec unit 900 is to be split into non-square codec units, the image decoding device 100 may split the first codec unit 900 into a plurality of non-square codec units. Specifically, when the split shape mode information indicates that an odd number of codec units are to be determined by splitting the first codec unit 900 in the horizontal or vertical direction, the image decoding device 100 may split the square first codec unit 900 into an odd number of codec units. For example, the second codec units 910a, 910b, and 910c determined by splitting the square first codec unit 900 in the vertical direction, or the second codec units 920a, 920b, and 920c determined by splitting the square first codec unit 900 in the horizontal direction. According to an embodiment, the image decoding device 100 may determine whether the second codec units 910a, 910b, 910c, 920a, 920b, and 920c included in the first codec unit 900 satisfy the conditions for processing in a specific order, and the condition relates to whether at least one of the width and height of the first codec unit 900 will be split in half along the boundaries of the second codec units 910a, 910b, 910c, 920a, 920b, and 920c. Refer to Figure 9 , since the boundaries of the second codec units 910a, 910b, and 910c determined by splitting the square first codec unit 900 in the vertical direction do not split the width of the first codec unit 900 in half, it can be determined that the first codec unit 900 does not satisfy the conditions for processing in a specific order. In addition, since the boundaries of the second codec units 920a, 920b, and 920c determined by splitting the square first codec unit 900 in the horizontal direction do not split the height of the first codec unit 900 in half, it can be determined that the first codec unit 900 does not satisfy the conditions for processing in a specific order.
[0162] When the above conditions are not satisfied, the image decoding device 100 may decide to break the scanning order, and may determine that the first codec unit 900 is to be split into an odd number of codec units based on the determined result.
[0163] According to an embodiment, when a codec unit is split into an odd number of codec units, the image decoding device 100 may impose specific constraints on the codec units at specific positions among the split codec units.
[0164] The constraints or specific positions have been described above through various embodiments, so the detailed description thereof will be omitted.
[0165] According to an embodiment, the image decoding apparatus 100 may determine codec units of various shapes by dividing a first codec unit. Referring to Figure 9 , the image decoding apparatus 100 may divide a square first codec unit 900 or a non-square first codec unit 930 or 950 into codec units of various shapes. Figure 10 FIG. shows that when a second codec unit having a non-square shape satisfies a specific condition according to an embodiment, the shapes into which the second codec unit can be divided are restricted, where the second codec unit having a non-square shape is determined when the image decoding apparatus divides the first codec unit. According to an embodiment, the image decoding apparatus 100 may determine to divide a square first codec unit 1000 into non-square second codec units 1010a and 1010b or 1020a and 1020b based on the division shape pattern information obtained by the receiver 110. The second codec units 1010a and 1010b or 1020a and 1020b may be independently divided. Accordingly, the image decoding apparatus 100 may determine whether or not to divide each of the second codec units 1010a and 1010b or 1020a and 1020b into a plurality of codec units based on the division shape pattern information of each of the second codec units 1010a and 1010b or 1020a and 1020b. According to an embodiment, the image decoding apparatus 100 may determine third codec units 1012a and 1012b by dividing a second codec unit 1010a on the left side of a non-square in the horizontal direction, where the second codec unit 1010a on the left side of the non-square is determined by dividing the first codec unit 1000 in the vertical direction.
[0166] However, when dividing the second codec unit 1010a on the left side in the horizontal direction, the image decoding apparatus 100 may restrict the second codec unit 1010b on the right side from being divided in the horizontal direction in which the second codec unit 1010a on the left side is divided. When determining third codec units 1014a and 1014b by dividing the second codec unit 1010b on the right side in the same direction, since the second codec unit 1010a on the left side and the second codec unit 1010b on the right side are independently divided in the horizontal direction, the third codec units 1012a, 1012b, 1014a, and 1014b may be determined.
[0167] However, this case is equivalent to the case where the image decoding apparatus 100 divides the first codec unit 1000 into four square second codec units 1030a, 1030b, 1030c, and 1030d based on the division shape pattern information, and may be inefficient in terms of image decoding.
[0168] According to an embodiment, the image decoding device 100 may determine third decoded units 1022a and 1022b or 1024a and 1024b by dividing non-square second decoded units 1020a or 1020b in the vertical direction, where the non-square second decoded units 1020a or 1020b are determined by dividing a first decoded unit 1000 in the horizontal direction. However, when dividing a second decoded unit (e.g., the upper second decoded unit 1020a) in the vertical direction, for the reasons described above, the image decoding device 100 may constrain another second decoded unit (e.g., the lower second decoded unit 1020b) not to be divided in the vertical direction of the upper second decoded unit 1020a that is divided. Figure 11 Illustrated is a process of an image decoding device dividing a square decoded unit when the division shape mode information indicates that the square decoded unit will not be divided into four square decoded units according to an embodiment. According to an embodiment, the image decoding device 100 may determine second decoded units 1110a and 1110b or 1120a and 1120b, etc. by dividing a first decoded unit 1100 based on the division shape mode information.
[0169] The division shape mode information may include information on various methods of dividing decoded units. However, the information on various division methods may not include information for dividing a decoded unit into four square decoded units. According to such division shape mode information, the image decoding device 100 may not divide the square first decoded unit 1100 into four square second decoded units 1130a, 1130b, 1130c, and 1130d.
[0170] The image decoding device 100 may determine non-square second decoded units 1110a and 1110b or 1120a and 1120b, etc. based on the division shape mode information. According to an embodiment, the image decoding device 100 may independently divide the non-square second decoded units 1110a and 1110b or 1120a and 1120b, etc. Each of the second decoded units 1110a and 1110b or 1120a and 1120b, etc. may be recursively divided in a specific order, and the division method may correspond to the method of dividing the first decoded unit 1100 based on the division shape mode information. For example, the image decoding device 100 may determine square third decoded units 1112a and 1112b by dividing the left second decoded unit 1110a in the horizontal direction, and may determine square third decoded units 1114a and 1114b by dividing the right second decoded unit 1110b in the horizontal direction.
[0171] Furthermore, the image decoding device 100 can determine square third-stage decoding units 1116a, 1116b, 1116c, and 1116d by dividing the second-stage decoding unit 1110a on the left and the second-stage decoding unit 1110b on the right in the horizontal direction. In this case, decoding units having the same shape as the four square second-stage decoding units 1130a, 1130b, 1130c, and 1130d divided from the first-stage decoding unit 1100 can be determined. As another example, the image decoding device 100 can determine square third-stage decoding units 1122a and 1122b by dividing the upper second-stage decoding unit 1120a in the vertical direction, and can determine square third-stage decoding units 1124a and 1124b by dividing the lower second-stage decoding unit 1120b in the vertical direction.
[0172] Furthermore, the image decoding device 100 can determine square third-stage decoding units 1126a, 1126b, 1126c, and 1126d by dividing the upper second-stage decoding unit 1120a and the lower second-stage decoding unit 1120b in the vertical direction.
[0173] In this case, decoding units having the same shape as the four square second-stage decoding units 1130a, 1130b, 1130c, and 1130d divided from the first-stage decoding unit 1100 can be determined. Figure 12 It is shown that the processing order among multiple decoding units according to an embodiment can vary according to the process of dividing the decoding units. According to an embodiment, the image decoding device 100 can divide the first-stage decoding unit 1200 based on the division shape mode information. When the block shape indicates a square shape and the division shape mode information indicates dividing the first-stage decoding unit 1200 in at least one of the horizontal and vertical directions, the image decoding device 100 can determine second-stage decoding units (e.g., 1210a and 1210b or 1220a and 1220b, etc.) by dividing the first-stage decoding unit 1200.
[0174] Reference Figure 12, non-square second codec units 1210a and 1210b or 1220a and 1220b determined by splitting the first codec unit 1200 only in the horizontal or vertical direction can be independently split based on the split shape pattern information of each codec unit. For example, the image decoding apparatus 100 can determine third codec units 1216a, 1216b, 1216c, and 1216d by splitting the second codec units 1210a and 1210b in the horizontal direction, where the second codec units 1210a and 1210b are generated by splitting the first codec unit 1200 in the vertical direction, and can determine third codec units 1226a, 1226b, 1226c, and 1226d by splitting the second codec units 1220a and 1220b in the vertical direction, where the second codec units 1220a and 1220b are generated by splitting the first codec unit 1200 in the horizontal direction. The process of splitting the second codec units 1210a and 1210b or 1220a and 1220b has been described above with reference to Figure 11 The process of splitting the second codec units 1210a and 1210b or 1220a and 1220b has been described, and thus a detailed description thereof will be omitted. According to an embodiment, the image decoding apparatus 100 can process the codec units in a specific order.
[0175] The operation of processing the codec units in a specific order has been described above with reference to Figure 7 and thus a detailed description thereof will be omitted.
[0176] With reference to Figure 12 , the image decoding apparatus 100 can determine four square third codec units 1216a, 1216b, 1216c, and 1216d and 1226a, 1226b, 1226c, and 1226d by splitting the square first codec unit 1200.
[0177] According to an embodiment, the image decoding apparatus 100 may determine the processing order of the third codec units 1216a, 1216b, 1216c, 1216d, and 1226a, 1226b, 1226c, and 1226d based on the divided shapes into which the first codec unit 1200 is divided. According to an embodiment, the image decoding apparatus 100 may determine the third codec units 1216a, 1216b, 1216c, and 1216d by dividing the second codec units 1210a and 1210b in a horizontal direction, where the second codec units 1210a and 1210b are generated by dividing the first codec unit 1200 in a vertical direction, and may process the third codec units 1216a, 1216b, 1216c, and 1216d in the processing order 1217 for initially processing the third codec units 1216a and 1216c included in the second codec unit 1210a on the left side in the vertical direction, and then processing the third codec units 1216b and 1216d included in the second codec unit 1210b on the right side in the vertical direction. According to an embodiment, the image decoding apparatus 100 may determine the third codec units 1226a, 1226b, 1226c, and 1226d by dividing the second codec units 1220a and 1220b in a vertical direction, where the second codec units 1220a and 1220b are generated by dividing the first codec unit 1200 in a horizontal direction, and may process the third codec units 1226a, 1226b, 1226c, and 1226d in the processing order 1227 for initially processing the third codec units 1226a and 1226b included in the second codec unit 1220a above in the horizontal direction, and then processing the third codec units 1226c and 1226d included in the second codec unit 1220b below in the horizontal direction.
[0178] Reference Figure 12 , the square third codec units 1216a, 1216b, 1216c, 1216d, and 1226a, 1226b, 1226c, and 1226d may be determined by dividing the second codec units 1210a and 1210b and 1220a and 1220b, respectively.
[0179] Although the second codec units 1210a and 1210b are determined by vertically dividing the first codec unit 1200, different from the second codec units 1220a and 1220b determined by horizontally dividing the first codec unit 1200, the third codec units 1216a, 1216b, 1216c, 1216d and 1226a, 1226b, 1226c, and 1226d divided therefrom ultimately show codec units of the same shape divided from the first codec unit 1200. Thus, by recursively dividing codec units in different ways based on the division shape pattern information, even when the codec units are finally determined to have the same shape, the image decoding apparatus 100 can process multiple codec units in different orders. Figure 13 Illustrates a process of determining the depth of a codec unit as the shape and size of the codec unit change when the codec unit is recursively divided to determine multiple codec units according to an embodiment. According to an embodiment, the image decoding apparatus 100 may determine the depth of a codec unit based on a specific criterion.
[0180] For example, the specific criterion may be the long side length of the codec unit. When the long side length of the codec unit before division is 2n times (n > 0) the long side length of the currently divided codec unit, the image decoding apparatus 100 may determine that the depth of the current codec unit is increased by n times from the depth of the codec unit before division. In the following description, a codec unit with an increased depth is represented as a codec unit with a lower depth. Refer to Figure 13 , according to an embodiment, the image decoding apparatus 100 may determine the second codec unit 1302 and the third codec unit 1304 with lower depths by dividing the square first codec unit 1300 based on block shape information indicating a square shape (e.g., the block shape information may be represented as "0:SQUARE"). Assuming that the size of the square first codec unit 1300 is 2N×2N, the size of the second codec unit 1302 determined by dividing the width and height of the first codec unit 1300 into 1 / 2 may be N×N.
[0181] Furthermore, the size of the third codec unit 1304 determined by dividing the width and height of the second codec unit 1302 into 1 / 2 may be N / 2×N / 2.
[0182] In this case, the width and height of the third-stage decoding unit 1304 are 1 / 4 times those of the first-stage decoding unit 1300. When the depth of the first-stage decoding unit 1300 is D, the depth of the second-stage decoding unit 1302 with a width and height that are 1 / 2 times those of the first-stage decoding unit 1300 can be D + 1, and the depth of the third-stage decoding unit 1304 with a width and height that are 1 / 4 times those of the first-stage decoding unit 1300 can be D + 2.
[0183] According to an embodiment, the image decoding apparatus 100 may determine the second-stage decoding units 1312 or 1322 and the third-stage decoding units 1314 or 1324 with a lower depth by dividing the non-square first-stage decoding units 1310 or 1320 based on block shape information indicating a non-square shape (e.g., the block shape information may be represented as "1:NS_VER" indicating a non-square shape with a height longer than the width, or may be represented as "2:NS_HOR" indicating a non-square shape with a width longer than the height). The image decoding apparatus 100 may determine the second-stage decoding units 1302, 1312, or 1322 by dividing at least one of the width and height of the first-stage decoding unit 1310 having a size of N×2N.
[0184] That is, the image decoding apparatus 100 may determine the second-stage decoding unit 1302 having a size of N×N or the second-stage decoding unit 1322 having a size of N×N / 2 by dividing the first-stage decoding unit 1310 in the horizontal direction, or may determine the second-stage decoding unit 1312 having a size of N / 2×N by dividing the first-stage decoding unit 1310 in both the horizontal and vertical directions. According to an embodiment, the image decoding apparatus 100 may determine the second-stage decoding units (e.g., 1302, 1312, or 1322) by dividing at least one of the width and height of the first-stage decoding unit 1320 having a size of 2N×N.
[0185] That is, the image decoding apparatus 100 may determine the second-stage decoding unit 1302 having a size of N×N or the second-stage decoding unit 1312 having a size of N / 2×N by dividing the first-stage decoding unit 1320 in the vertical direction, or may determine the second-stage decoding unit 1322 having a size of N×N / 2 by dividing the first-stage decoding unit 1320 in both the horizontal and vertical directions. According to an embodiment, the image decoding apparatus 100 may determine the third-stage decoding units (e.g., 1304, 1314, or 1324) by dividing at least one of the width and height of the second-stage decoding unit 1302 having a size of N×N.
[0186] That is to say, the image decoding device 100 can determine a third coding and decoding unit 1304 with a size of N / 2×N / 2, a third coding and decoding unit 1314 with a size of N / 4×N / 2, or a third coding and decoding unit 1324 with a size of N / 2×N / 4 by dividing the second coding and decoding unit 1302 in the vertical and horizontal directions. According to an embodiment, the image decoding device 100 can determine a third coding and decoding unit (e.g., 1304, 1314, or 1324) by dividing at least one of the width and height of the second coding and decoding unit 1312 with a size of N / 2×N.
[0187] That is to say, the image decoding device 100 can determine a third coding and decoding unit 1304 with a size of N / 2×N / 2 or a third coding and decoding unit 1324 with a size of N / 2×N / 4 by dividing the second coding and decoding unit 1312 in the horizontal direction, or can determine a third coding and decoding unit 1314 with a size of N / 4×N / 2 by dividing the second coding and decoding unit 1312 in the vertical and horizontal directions. According to an embodiment, the image decoding device 100 can determine a third coding and decoding unit (e.g., 1304, 1314, or 1324) by dividing at least one of the width and height of the second coding and decoding unit 1322 with a size of N×N / 2. That is to say, the image decoding device 100 can determine a third coding and decoding unit 1304 with a size of N / 2×N / 2 or a third coding and decoding unit 1314 with a size of N / 4×N / 2 by dividing the second coding and decoding unit 1322 in the vertical direction, or can determine a third coding and decoding unit 1324 with a size of N / 2×N / 4 by dividing the second coding and decoding unit 1322 in the vertical and horizontal directions.
[0188] According to an embodiment, the image decoding device 100 can divide a square coding and decoding unit (e.g., 1300, 1302, or 1304) in the horizontal or vertical direction. For example, the image decoding device 100 can determine a first coding and decoding unit 1310 with a size of N×2N by dividing the first coding and decoding unit 1300 with a size of 2N×2N in the vertical direction, or can determine a first coding and decoding unit 1320 with a size of 2N×N by dividing the first coding and decoding unit 1300 in the horizontal direction.
[0189] According to an embodiment, when determining the depth based on the length of the longest side of the coding and decoding unit, the depth of the coding and decoding unit determined by dividing the first coding and decoding unit 1300 with a size of 2N×2N in the horizontal or vertical direction can be the same as the depth of the first coding and decoding unit 1300.
[0190] According to an embodiment, the width and height of the third-tier decoding unit 1314 or 1324 can be 1 / 4 times that of the first-tier decoding unit 1310 or 1320. When the depth of the first-tier decoding unit 1310 or 1320 is D, the depth of the second-tier decoding unit 1312 or 1322 with a width and height that are 1 / 2 times that of the first-tier decoding unit 1310 or 1320 can be D + 1, and the depth of the third-tier decoding unit 1314 or 1324 with a width and height that are 1 / 4 times that of the first-tier decoding unit 1310 or 1320 can be D + 2. Figure 14 Shows the depth that can be determined based on the shape and size of the encoding / decoding unit according to an embodiment and the partial index (PID) for distinguishing the encoding / decoding unit.
[0191] According to an embodiment, the image decoding device 100 can determine second-tier decoding units of various shapes by dividing the square first-tier decoding unit 1400. Refer to Figure 14 , the image decoding device 100 can determine the second-tier decoding units 1402a and 1402b, 1404a and 1404b, and 1406a, 1406b, 1406c, and 1406d by dividing the first-tier decoding unit 1400 in at least one of the vertical and horizontal directions based on the division shape mode information. That is, the image decoding device 100 can determine the second-tier decoding units 1402a and 1402b, 1404a and 1404b, and 1406a, 1406b, 1406c, and 1406d based on the division shape mode information of the first-tier decoding unit 1400.
[0192] According to an embodiment, the depth of the second-tier decoding units 1402a and 1402b, 1404a and 1404b, and 1406a, 1406b, 1406c, and 1406d determined based on the division shape mode information of the square first-tier decoding unit 1400 can be determined based on their long side lengths. For example, since the side length of the square first-tier decoding unit 1400 is equal to the long side lengths of the non-square second-tier decoding units 1402a and 1402b and 1404a and 1404b, the first-tier decoding unit 1400 and the non-square second-tier decoding units 1402a and 1402b and 1404a and 1404b can have the same depth, for example, D.
[0193] However, when the image decoding apparatus 100 divides the first codec unit 1400 into four square second codec units 1406a, 1406b, 1406c, and 1406d based on the segmentation shape mode information, since the side length of the square second codec units 1406a, 1406b, 1406c, and 1406d is 1 / 2 times the side length of the first codec unit 1400, the depth of the second codec units 1406a, 1406b, 1406c, and 1406d can be D + 1 which is 1 deeper than the depth D of the first codec unit 1400. According to an embodiment, the image decoding apparatus 100 can determine a plurality of second codec units 1412a and 1412b and 1414a, 1414b, and 1414c by dividing the first codec unit 1410 having a height longer than the width in the horizontal direction based on the segmentation shape mode information. According to an embodiment, the image decoding apparatus 100 can determine a plurality of second codec units 1422a and 1422b and 1424a, 1424b, and 1424c by dividing the first codec unit 1420 having a width longer than the height in the vertical direction based on the segmentation shape mode information.
[0194] According to an embodiment, the depth of the second codec units 1412a and 1412b and 1414a, 1414b, and 1414c or 1422a and 1422b and 1424a, 1424b, and 1424c determined based on the segmentation shape mode information of the non-square first codec unit 1410 or 1420 can be determined based on its long side length. For example, since the side length of the square second codec units 1412a and 1412b is 1 / 2 times the long side length of the non-square first codec unit 1410 having a height longer than the width, the depth of the square second codec units 1412a and 1412b is D + 1 which is 1 deeper than the depth D of the non-square first codec unit 1410. Further, the image decoding apparatus 100 can divide the non-square first codec unit 1410 into an odd number of second codec units 1414a, 1414b, and 1414c based on the segmentation shape mode information. The odd number of second codec units 1414a, 1414b, and 1414c can include non-square second codec units 1414a and 1414c and a square second codec unit 1414b.
[0195] In this case, since the length of the long side of the non-square second-tier decoding units 1414a and 1414c and the side length of the square second-tier decoding unit 1414b are 1 / 2 times the length of the long side of the first-tier decoding unit 1410, the depth of the second-tier decoding units 1414a, 1414b, and 1414c can be D + 1, which is 1 deeper than the depth D of the non-square first-tier decoding unit 1410. The image decoding device 100 can determine the depth of the decoding units divided from the first-tier decoding unit 1420 having a non-square shape with a width longer than the height by using the method of determining the depth of the decoding units divided from the first-tier decoding unit 1410 as described above. According to an embodiment, when an odd number of divided decoding units do not have the same size, the image decoding device 100 can determine the PID for identifying the divided decoding units based on the size ratio between the decoding units. Refer to Figure 14 , among the odd number of divided decoding units 1414a, 1414b, and 1414c, the decoding unit 1414b at the central position can have the same width as the other decoding units 1414a and 1414c and twice the height of the other decoding units 1414a and 1414c. That is, in this case, the decoding unit 1414b at the central position can include two of the other decoding units 1414a or 1414c. Accordingly, when the PID of the decoding unit 1414b at the central position based on the scanning order is 1, the PID of the decoding unit 1414c adjacent to the decoding unit 1414b can be increased by 2 and thus can be 3.
[0196] That is, there may be a discontinuity in the PID values. According to an embodiment, the image decoding device 100 can determine whether an odd number of divided decoding units do not have the same size based on whether there is a discontinuity in the PID for identifying the divided decoding units. According to an embodiment, the image decoding device 100 can determine whether to use a specific segmentation method based on the PID values for identifying the multiple decoding units determined by dividing the current decoding unit. Refer to Figure 14 , the image decoding device 100 can determine an even number of decoding units 1412a and 1412b or an odd number of decoding units 1414a, 1414b, and 1414c by dividing the first-tier decoding unit 1410 having a rectangular shape with a height longer than the width.
[0197] The image decoding apparatus 100 may use PIDs indicating respective codec units to identify the respective codec units. According to an embodiment, PIDs may be obtained from samples at specific positions (e.g., upper left samples) of each codec unit. According to an embodiment, by using the PIDs for distinguishing codec units, the image decoding apparatus 100 may determine a codec unit at a specific position among the divided codec units. According to an embodiment, when the segmentation shape mode information of a first codec unit 1410 having a rectangular shape with a height longer than the width indicates that the codec unit is divided into three codec units, the image decoding apparatus 100 may divide the first codec unit 1410 into three codec units 1414a, 1414b, and 1414c. The image decoding apparatus 100 may assign PIDs to each of the three codec units 1414a, 1414b, and 1414c. The image decoding apparatus 100 may compare the PIDs of an odd number of divided codec units to determine a codec unit at the center position among the codec units. The image decoding apparatus 100 may determine the codec unit 1414b having a PID corresponding to the median value among the PIDs of the codec units as the codec unit at the center position among the codec units determined by dividing the first codec unit 1410. According to an embodiment, when the divided codec units do not have the same size, the image decoding apparatus 100 may determine PIDs for distinguishing the divided codec units based on the size ratio between the codec units. Refer to Figure 14 , the codec unit 1414b generated by dividing the first codec unit 1410 may have the same width as the other codec units 1414a and 1414c and a height twice that of the other codec units 1414a and 1414c. In this case, when the PID of the codec unit 1414b at the center position is 1, the PID of the codec unit 1414c adjacent to the codec unit 1414b may be increased by 2 and thus may be 3. When the PIDs are not increased uniformly as described above, the image decoding apparatus 100 may determine that the codec unit is divided into a plurality of codec units including a codec unit having a size different from that of the other codec units.
[0198] According to an embodiment, when the segmentation shape mode information indicates that the codec unit is divided into an odd number of codec units, the image decoding apparatus 100 may divide the current codec unit in such a way that the size of a codec unit at a specific position (e.g., the codec unit at the center position) among the odd number of codec units is different from that of the other codec units.
[0199] In this case, the image decoding apparatus 100 may determine the codec unit at the center position having a different size by using the PID of the codec unit.
[0200] However, the PID, size, or position of the codec unit at a specific position to be determined is not limited to the above examples, and various PIDs, as well as various positions and sizes of the codec unit, can be used. According to an embodiment, the image decoding apparatus 100 may use a specific data unit from which the codec unit starts to be recursively divided. Figure 15 FIG. illustrates determining a plurality of codec units based on a plurality of specific data units included in a picture according to an embodiment.
[0201] According to an embodiment, a specific data unit may be defined as a data unit from which the codec unit starts to be recursively divided by using split shape pattern information. That is, the specific data unit may correspond to the codec unit with the highest depth for determining a plurality of codec units split from the current picture. In the following description, for ease of explanation, the specific data unit is referred to as a reference data unit. According to an embodiment, the reference data unit may have a specific size and a specific shape.
[0202] According to an embodiment, the reference codec unit may include M×N samples. Here, M and N may be equal to each other and may be integers represented as a power of 2. That is, the reference data unit may have a square shape or a non-square shape and may be divided into an integer number of codec units.
[0203] According to an embodiment, the image decoding apparatus 100 may divide the current picture into a plurality of reference data units. According to an embodiment, the image decoding apparatus 100 may divide the plurality of reference data units split from the current picture by using the split shape pattern information of each reference data unit.
[0204] The process of dividing the reference data unit may correspond to a division process using a quadtree structure. According to an embodiment, the image decoding apparatus 100 may pre-determine the minimum size allowed for the reference data units included in the current picture.
[0205] Accordingly, the image decoding apparatus 100 may determine reference data units having various sizes equal to or greater than the minimum size, and may determine one or more codec units by referring to the determined reference data units by using the split shape pattern information. Refer Figure 15 , the image decoding apparatus 100 may use the square reference codec unit 1500 or the non-square reference codec unit 1502.
[0206] According to an embodiment, the shape and size of a reference coding / decoding unit may be determined based on various data units (e.g., sequence, picture, slice, slice segment, tile, tile group, largest coding / decoding unit, etc.) that can include one or more reference coding / decoding units. According to an embodiment, the receiver 110 of the image decoding apparatus 100 may obtain at least one of reference coding / decoding unit shape information and reference coding / decoding unit size information for each data unit among the various data units from a bitstream. The process of dividing the square reference coding / decoding unit 1500 into one or more coding / decoding units has been described above in the process of dividing the current coding / decoding unit 300 of Figure 3 and the process of dividing the non-square reference coding / decoding unit 1502 into one or more coding / decoding units has been described above in the process of dividing the current coding / decoding unit 400 or 450 of Figure 4 so a detailed description thereof will be omitted. According to an embodiment, the image decoding apparatus 100 may use a PID for identifying the size and shape of a reference coding / decoding unit to determine the size and shape of the reference coding / decoding unit based on some data units predetermined according to specific conditions. That is, the receiver 110 may obtain from the bitstream only the PID for identifying the size and shape of the reference coding / decoding unit for each slice, slice segment, tile, tile group, largest coding / decoding unit among the various data units (e.g., sequence, picture, slice, slice segment, tile, tile group, largest coding / decoding unit, etc.) that are data units satisfying specific conditions (e.g., data units having a size equal to or smaller than a slice). The image decoding apparatus 100 may determine the size and shape of the reference data unit for each data unit satisfying specific conditions by using the PID.
[0207] When the reference coding / decoding unit shape information and the reference coding / decoding unit size information are obtained and used from the bitstream according to each data unit having a relatively small size, the efficiency of using the bitstream may not be high, so only the PID may be obtained and used instead of directly obtaining the reference coding / decoding unit shape information and the reference coding / decoding unit size information. In this case, at least one of the size and shape of the reference coding / decoding unit corresponding to the PID for identifying the size and shape of the reference coding / decoding unit may be predetermined. That is, the image decoding apparatus 100 may determine at least one of the size and shape of the reference coding / decoding unit included in the data unit used as the unit for obtaining the PID by selecting at least one of the predetermined size and shape of the reference coding / decoding unit based on the PID. According to an embodiment, the image decoding apparatus 100 may use one or more reference coding / decoding units included in the largest coding / decoding unit. That is, the largest coding / decoding unit divided from a picture may include one or more reference coding / decoding units, and the coding / decoding units may be determined by recursively dividing each reference coding / decoding unit.
[0208] According to an embodiment, at least one of the width and height of the maximum coding / decoding unit may be an integer multiple of at least one of the width and height of the reference coding / decoding unit. According to an embodiment, the size of the reference coding / decoding unit may be obtained by dividing the maximum coding / decoding unit n times based on a quadtree structure. That is, according to various embodiments, the image decoding apparatus 100 may determine the reference coding / decoding unit by dividing the maximum coding / decoding unit n times based on a quadtree structure, and may divide the reference coding / decoding unit based on at least one of block shape information and split shape mode information. According to an embodiment, the image decoding apparatus 100 may obtain, from a bitstream, block shape information indicating the shape of a current coding / decoding unit or split shape mode information indicating a split method of the current coding / decoding unit, and may use the obtained information.
[0209] The split shape mode information may be included in a bitstream related to various data units.
[0210] For example, the image decoding apparatus 100 may use the split shape mode information included in a sequence parameter set, a picture parameter set, a video parameter set, a slice header, a slice segment header, a tile header, or a tile group header. Further, the image decoding apparatus 100 may obtain, from a bitstream, syntax information (e.g., a syntax element) corresponding to the block shape information or the split shape mode information according to each maximum coding / decoding unit, each reference coding / decoding unit, or a processing block, and may use the obtained syntax element. Hereinafter, a method of determining a split rule according to an embodiment of the present disclosure will be described in detail. The image decoding apparatus 100 may determine a split rule of an image. The split rule may be predetermined between the image decoding apparatus 100 and the image encoding apparatus 200.
[0211] The image decoding apparatus 100 may determine a split rule of an image based on information obtained from a bitstream. The image decoding apparatus 100 may determine a split rule based on information obtained from at least one of a sequence parameter set, a picture parameter set, a video parameter set, a slice header, a slice segment header, a tile header, and a tile group header. The image decoding apparatus 100 may determine a split rule differently according to a frame, a slice, a tile, a temporal layer, a maximum coding / decoding unit, or a coding / decoding unit. The image decoding apparatus 100 may determine a split rule based on the block shape of a coding / decoding unit. The block shape may include the size, shape, width-to-height ratio, and orientation of the coding / decoding unit.
[0212] The image decoding apparatus 100 may predetermine based on the block shape of a coding / decoding unit to determine a split rule. However, the embodiment is not limited thereto. The image decoding apparatus 100 may determine a split rule based on information obtained from a received bitstream. The shape of a coding / decoding unit may include a square and a non-square.
[0213] When the lengths of the width and height of the codec unit are the same, the image decoding apparatus 100 may determine the shape of the codec unit as a square. In addition, when the lengths of the width and height of the codec unit are different, the image decoding apparatus 100 may determine the shape of the codec unit as a non-square. The size of the codec unit may include various sizes such as 4×4, 8×4, 4×8, 8×8, 16×4, 16×8, ……, and 256×256. The size of the codec unit may be classified based on the length of the long side, the length of the short side, or the area of the codec unit. The image decoding apparatus 100 may apply the same segmentation rule to the codec units classified into the same group. For example, the image decoding apparatus 100 may classify the codec units having the same length of the long side as having the same size.
[0214] In addition, the image decoding apparatus 100 may apply the same segmentation rule to the codec units having the same length of the long side. The width-to-height ratio of the codec unit may include 1:2, 2:1, 1:4, 4:1, 1:8, 8:1, 1:16, 16:1, 32:1, 1:32, etc. In addition, the direction of the codec unit may include a horizontal direction and a vertical direction. The horizontal direction may indicate a case where the width length of the codec unit is greater than its height length.
[0215] The vertical direction may indicate a case where the width length of the codec unit is shorter than its height length. The image decoding apparatus 100 may adaptively determine the segmentation rule based on the size of the codec unit. The image decoding apparatus 100 may differently determine the allowable segmentation shape patterns based on the size of the codec unit. For example, the image decoding apparatus 100 may determine whether segmentation is allowed based on the size of the codec unit. The image decoding apparatus 100 may determine the segmentation direction according to the size of the codec unit.
[0216] The image decoding apparatus 100 may determine the allowable segmentation type according to the size of the codec unit. The segmentation rule determined based on the size of the codec unit may be a segmentation rule predetermined between the image encoding apparatus 200 and the image decoding apparatus 100.
[0217] In addition, the image decoding apparatus 100 may determine the segmentation rule based on the information obtained from the bitstream. The image decoding apparatus 100 may adaptively determine the segmentation rule based on the position of the codec unit.
[0218] The image decoding device 100 can adaptively determine a segmentation rule based on the position of the codec unit in the image. In addition, the image decoding device 100 can determine a segmentation rule such that the codec units generated via different segmentation paths do not have the same block shape. However, the embodiments are not limited thereto, and the codec units generated via different segmentation paths have the same block shape. The codec units generated via different segmentation paths can have different decoding processing orders.
[0219] Figure 16 is a block diagram of an image codec system.
[0220] The encoding end 1610 of the image codec system 1600 can transmit an encoded bitstream of an image, and the decoding end 1650 of the image codec system 1600 can receive the bitstream and decode the bitstream to output a reconstructed image. Here, the decoding end 1650 can be configured similarly to the image decoding device 100.
[0221] In the encoding end 1610, the prediction encoder 1615 outputs a reference image through inter-frame prediction and intra-frame prediction, and the transformer and quantizer 1616 quantize the residual data between the reference image and the current input image into quantized transform coefficients and output the quantized transform coefficients. The entropy encoder 1625 encodes the quantized transform coefficients and outputs the encoded quantized transform coefficients as a bitstream. The quantized transform coefficients are reconstructed into data in the spatial domain via the inverse quantizer and inverse transformer 1630, and the reconstructed data in the spatial domain is output as a reconstructed image via the deblocking filter 1635 and the loop filter 1640. The reconstructed image can be used as a reference image for the next input image in the prediction encoder 1615.
[0222] The encoded image data in the bitstream received by the decoding end 1650 is reconstructed into residual data in the spatial domain via the entropy decoder 1655 and the inverse quantizer and inverse transformer 1660. When the reference image output by the prediction decoder 1675 and the residual data are combined, the image data in the spatial domain can be configured, and the deblocking filter 1665 and the loop filter 1670 can output a reconstructed image of the current original image by performing filtering on the image data in the spatial domain. The reconstructed image can be used as a reference image for the next original image by the prediction decoder 1675.
[0223] The loop filter 1640 of the encoding end 1610 performs loop filtering by using filter information input according to user input or system settings. The filter information used by the loop filter 1640 is output to the entropy encoder 1610 and sent to the decoding end 1650 together with the encoded image data. The loop filter 1670 of the decoding end 1650 can perform loop filtering based on the filter information input from the decoding end 1650.
[0224] In the present disclosure, a "tree structure" may refer to a hierarchical structure of one or more coding / decoding units formed according to whether the splitting pattern of the coding / decoding unit is a quaternary split, a binary split, a ternary split, or no split. For example, the hierarchical structure of blocks generated from the current coding / decoding unit according to Figure 5 the splitting process is referred to as a tree structure.
[0225] In the present disclosure, the "availability of a block" refers to whether the block has been encoded or decoded so that information of the block can be obtained. Specifically, in the encoding process, when the current block has been encoded, the encoding information of the current block can be used to encode adjacent blocks, so the current block can be shown as available. When the current block has not been encoded, the current block can be shown as unavailable. Similarly, in the decoding process, when the current block has been decoded, the encoding information of the current block can be used to decode adjacent blocks, so the current block can be shown as available. When the current block has not been decoded, the current block can be shown as unavailable.
[0226] In the present disclosure, the "availability of motion information of a block" refers to whether motion prediction of the block (prediction other than prediction according to the intra mode or intra block copy mode) can be performed so that motion information of the block (motion vector, prediction direction (L0-pred, L1-pred, or Bi-pred), and reference picture index) can be obtained. Specifically, in the encoding process, when motion prediction has been performed on the current block and there is motion information of the current block, the motion information of the current block can be used to perform motion prediction of adjacent blocks, so the motion information of the current block can be shown as available. In the encoding process, when motion prediction has not been performed on the current block, the motion information of the current block can be shown as unavailable. Similarly, in the decoding process, when motion prediction has been performed on the current block and there is motion information of the current block, the motion information of the current block can be used to perform motion prediction of adjacent blocks, so the motion information of the current block can be shown as available. In the decoding process, when motion prediction has not been performed on the current block, the motion information of the current block can be shown as unavailable.
[0227] In the present disclosure, a "merge candidate" may correspond to a motion vector corresponding to an adjacent block of the current block. Since the predicted motion vector of the current block is determined from the motion vectors of adjacent blocks, each predicted motion vector can correspond to an adjacent block. Therefore, in the present disclosure, for the convenience of explanation, a "merge candidate" is described as corresponding to the motion vector of an adjacent block or corresponding to an adjacent block, and there is no difference in meaning between the two expressions.
[0228] In the present disclosure, an "affine merge candidate" may correspond to a control point vector corresponding to an adjacent block or block group of a current block. Since the control point vector is determined from motion vectors of adjacent blocks, or the control point vector is determined based on motion vectors of adjacent blocks belonging to a block group, each control point vector may correspond to a corresponding adjacent block or a corresponding block group. Thus, in the present disclosure, for ease of explanation, an "affine merge candidate" is described as corresponding to a control point vector determined from an adjacent block or block group, or corresponding to an adjacent block or block group, and there is no difference in meaning between the two expressions.
[0229] In the present disclosure, a "motion vector prediction (MVP) candidate" may correspond to a motion vector corresponding to an adjacent block of a current block. Since the predicted motion vector of the current block is determined from motion vectors of adjacent blocks, each predicted motion vector may correspond to an adjacent block. Thus, in the present disclosure, for ease of explanation, an "MVP candidate" is described as corresponding to a motion vector of an adjacent block or corresponding to an adjacent block, and there is no difference in meaning between the two expressions.
[0230] A "merge candidate" is an adjacent block (or motion vector of an adjacent block) used in a merge mode among inter prediction methods, while an "MVP candidate" corresponds to an adjacent block (or motion vector of an adjacent block) used in an AMVP (advanced motion vector prediction) mode among inter prediction methods. In the merge mode, not only can the motion vector of the merge candidate be used to determine the motion vector of the current block, but also the prediction direction (L0-pred, L1-pred, or Bi-pred) and reference picture index of the merge candidate can be used to determine the prediction direction and reference picture index of the current block. On the other hand, in the AMVP mode, the motion vector of the MVP candidate can be used to determine the motion vector of the current block, but the prediction direction and reference picture index of the current block can be determined separately from the prediction direction and reference picture index of the MVP candidate.
[0231] Technical problem
[0232] In the advanced motion vector prediction (AMVP) mode or the merge mode, the motion vector of the block corresponding to the current block (or the collocated block or the co-located block) in the reference picture (or the collocated reference picture) is used as the temporal motion vector predictor (temporal MVP) or the temporal motion vector candidate. When the temporal positions (e.g., picture order count (POC) values) of the reference pictures of the current block and the collocated block are different, the motion vector of the collocated block is used by scaling (linearly) according to the reference picture of the current block. However, when the collocated block is quantized to a high quantization parameter value and the collocated block is encoded or decoded in the skip mode or the merge mode, the accuracy of the motion vector of the collocated block may be reduced. In addition, when the block (or the collocated block) corresponding to the current block in the collocated reference picture is encoded or decoded in the intra mode, there is no motion vector of the collocated block. Thus, when there is no motion vector of the collocated block or the motion vector of the collocated block has low accuracy and the motion vector of the collocated block is used as the temporal motion vector prediction or candidate of the current block, the motion vector prediction performance of the current block may deteriorate, or the motion vector of the current block may be determined as an inaccurate value, resulting in an overall deterioration of the encoding or decoding performance or efficiency.
[0233] A technique is being used to improve the encoding or decoding efficiency by generating a virtual reference picture corresponding to the current picture to be encoded or decoded (e.g., decoded or encoded) via the frame rate up-conversion (FRUC) technique and using the virtual reference picture as a reference picture during the encoding or decoding (e.g., decoding or encoding) of the current picture. However, in this FRUC technique, it is necessary to allocate a picture buffer having the same size as the current picture to encode or decode the current picture, and it is necessary to perform additional complex processing in the pixel domain to generate the virtual reference picture. When processing high-resolution (such as 4K / 8K) images, there are unrealistic problems in terms of complexity, such as external memory bandwidth requirements or processing capabilities.
[0234] Methods and devices according to the present disclosure
[0235] The present disclosure proposes a method and an apparatus for generating motion information corresponding to a current picture by using information (e.g., reconstructed samples, motion information, etc.) of a picture that has been encoded or decoded (e.g., decoded or encoded) before the current picture and using motion information during the encoding or decoding (e.g., decoding or encoding) of the current picture, thereby improving the encoding or decoding efficiency.
[0236] When applying the method according to the present disclosure, motion information corresponding to the current picture can be generated before the current picture is coded or decoded (e.g., decoded or encoded), and the generated motion information can be used as a more accurate motion vector prediction or candidate in the existing AMVP mode or merge mode, or can be used in a new motion vector coding or decoding mode, thereby improving the coding or decoding efficiency. In addition, when applying the method according to the present disclosure, since the motion information is generated in units of blocks of a specific size (e.g., 4×4, 8×8, 16×16, etc.) rather than in units of actual pixels, the same effect as FRUC can be obtained, and it can also be considered more efficient in terms of buffer efficiency and computational complexity.
[0237] In the present disclosure, a picture that has been coded or decoded before the current picture is coded or decoded (e.g., decoded or encoded) is referred to as a "reconstructed picture" or "reconstructed image". For example, the reconstructed picture may include a picture included in a decoded picture buffer (DPB), or may include a reference picture included in a reference list (or reference picture list) (e.g., reference list 0 or reference list 1). Therefore, in the present disclosure, the reconstructed picture can be replaced with a reference picture. In the present disclosure, the number of reconstructed pictures used to generate the motion information of the current picture can be one or more. That is, the number of reconstructed pictures used to generate motion information according to the present disclosure can be one or can be multiple.
[0238] In the present disclosure, "motion information" may include at least one of motion vector information, prediction list utilization information, and prediction mode information. For example, the motion vector information may include at least one of an L0 motion vector (or the x component and / or y component of the L0 motion vector) or an L1 motion vector (or the x component and / or y component of the L1 motion vector). In the present disclosure, the L0 motion vector may be referred to as first motion vector information or mvL0, and the L1 motion vector may be referred to as second motion vector information or mvL1.
[0239] For example, the prediction list utilization information may include at least one of list 0 utilization information (e.g., predFlagL0) or list 1 utilization information (e.g., predFlagL1), and may be referred to as reference list utilization information. In the present disclosure, the list 0 utilization information may be referred to as first prediction list utilization information (or first reference list utilization information or first prediction list utilization flag information), and the list 1 utilization information may be referred to as second prediction list utilization information (or second reference list utilization information or second prediction list utilization flag information). In the present disclosure, when both list 0 and list 1 are used, this may be referred to as bidirectional prediction, and when only one of list 0 or list 1 is used, this may be referred to as unidirectional prediction.
[0240] For example, the prediction mode information may include at least one of information indicating whether the corresponding block is decoded in an intra mode or an inter mode (e.g., pred_mode_flag or CuPredMode) or information indicating whether the corresponding block is decoded in a skip mode (e.g., cu_skip_flag or CuSkipFlag).
[0241] The motion information generated according to the method proposed in the present disclosure may be referred to as "generated motion information". In the present disclosure, the generated motion information may also be referred to as reference motion information, generated reference motion information, etc. Moreover, in the present disclosure, the generated motion information may be abbreviated as generated information or generated data.
[0242] The motion information may be generated or stored in the form of a one-dimensional arrangement or a two-dimensional arrangement in units of blocks of a specific size or in units of M×N pixels (e.g., 4×4, 8×8, or 16×16), and the motion information generated or stored in this form may be referred to as a motion information map. The generated motion information of the current picture may refer to the motion information corresponding to the blocks included in the current picture generated in units of blocks of a specific size (e.g., 4×4, 8×8, or 16×16), and may be referred to as the motion information map of the current picture. In the present disclosure, the generated motion information and the motion information map may be used interchangeably. For example, the motion information map may also be referred to as the generated reference motion map.
[0243] Figure 17 It is a block diagram of the image decoding apparatus 100 according to the present disclosure.
[0244] Figure 17 The illustrated apparatus 100 may be included in the image decoding apparatus 100 described above with reference to Figure 1 and may include a motion information generator 130 and a decoder 120. The motion information generator 130 and the decoder 120 of the apparatus 100 may be implemented by at least one processor included in the apparatus 100.
[0245] According to the present disclosure, the motion information generator 130 may generate (reference) motion information of blocks included in the current picture based on at least one of pixel data of at least one reconstructed picture and / or motion information. For example, the motion information generator 130 may generate motion information corresponding to the current picture before encoding or decoding (e.g., decoding or encoding) the current picture. As a more specific example, the motion information generator 130 may generate motion information for encoding or decoding the current picture before encoding or decoding (e.g., decoding or encoding) an initial strip of the current picture. For example, the motion information generator 130 may generate motion information corresponding to the current picture by using motion information stored in at least one reconstructed picture (e.g., reference picture) (or using motion information stored in at least one reconstructed picture (e.g., reference picture) as an input). For example, the motion information generator 130 may generate motion information corresponding to the current picture by using pixel data of at least one reconstructed picture (e.g., reference picture) (or using pixel data of at least one reconstructed picture (e.g., reference picture) as an input). For example, the motion information generator 130 may generate motion information corresponding to the current picture by using motion information and pixel data stored in at least one reconstructed picture (e.g., reference picture) (or using motion information and pixel data stored in at least one reconstructed picture (e.g., reference picture) as an input).
[0246] For example, the pixel data of the reconstructed picture may be used as an input to the motion information generator 130 without separate processing, or downsampled pixel data obtained by downsampling the pixel data of the reconstructed picture may be used as an input to the motion information generator 130 (e.g., see Figure 28 and its related description).
[0247] The input to the motion information generator 130 according to the present disclosure is not limited to pixel data and motion information of the reconstructed picture, and other information of the reconstructed picture may also be used as an input to the motion information generator 130.
[0248] For example, the input to the motion information generator 130 may include at least one of the following: quantization information of at least one reconstructed picture, prediction samples and residual samples, and distance information from the current picture. For example, the quantization information may include quantization parameter information. For example, the prediction samples may include samples predicted by intra prediction or inter prediction. For example, the residual samples may include samples obtained by inverse-transforming and inverse-quantizing transform coefficient information signaled via a bitstream, or the difference between a reference sample and a sample of the current block. For example, the distance information from the current picture may include picture order count (POC) information of the reconstructed picture or information indicating the POC distance (or difference) between the reconstructed picture and the current picture.
[0249] The motion information generator 130 may generate motion information (or a motion information map) in units of blocks of a specific size (e.g., 4×4, 8×8, or 16×16). The generated motion information output by the motion information generator 130 may include at least one of motion information in units of blocks of a specific size (e.g., 4×4, 8×8, or 16×16), intra-coding / decoding related information, residual coding / decoding related information, illuminance change information, and statistical information.
[0250] For example, as described above, the motion information may include at least one of motion vector information, prediction list utilization information, and prediction mode information for a block of a specific size (e.g., see the description of “Motion Information” in “Methods and Apparatuses According to the Present Disclosure”).
[0251] For example, the intra-coding / decoding related information may include at least one of direct current (DC) transform coefficient information and intra-prediction mode information (e.g., intra-DC mode, intra-plane mode, or intra-angle mode) for a block of a specific size. For example, the residual coding / decoding related information may include information indicating whether residual coding / decoding (e.g., inverse transform and / or inverse quantization) is performed on a block of a specific size. Alternatively, the residual coding / decoding related information may include information indicating whether there is non-zero transform coefficient information for the corresponding block. For example, the illuminance change information may include information on the amount of illuminance change for illuminance compensation for a block of a specific size. For example, the statistical information may include probability information for entropy coding for a block of a specific size, as information available during coding / decoding (e.g., an encoding process and / or a decoding process).
[0252] The motion information generator 130 may process the input and generate the output through various methods. For example, the motion information generator 130 may generate generated information or generated data based on pixel data or motion information (and / or quantization information, prediction samples, residual samples, and distance information from the current picture) of at least one reconstructed picture by using arithmetic operations (e.g., averaging, median, interpolation, or filtering).
[0253] Alternatively, for example, the motion information generator 130 may generate generation information or generation data by using an artificial neural network based on at least one of pixel data or motion information (and / or quantization information, prediction samples, residual samples, and distance information from the current picture) of at least one reconstructed picture. For example, the artificial neural network available in the motion information generator 130 according to the present disclosure may include a convolutional neural network (CNN), a fully connected neural network (FCNN), or a combination thereof. The structures of the convolutional neural network and the fully connected neural network are well known, and thus detailed descriptions thereof will be omitted in this specification. The processing of the motion information generator 130 is not limited to processing using arithmetic operations and processing using an artificial neural network, and various methods may be applied.
[0254] The decoder 120 of the apparatus 100 may decode a current picture based on the motion information generated by the motion information generator 130 and at least one piece of syntax information obtained from a bitstream (e.g., the bitstream obtained by the receiver 110). For example, the decoder 120 may be configured to obtain motion information of a current block based on the motion information generated by the motion information generator 130 and at least one piece of syntax information obtained from the bitstream, and reconstruct the current block based on the obtained motion information. Details of using generation information or generation data (or a motion information map) during decoding according to the present disclosure will now be described in detail.
[0255] Figure 18 is a flowchart of an image decoding method 1800 according to the present disclosure. Figure 18 The method 1800 shown may be performed by the apparatus 100.
[0256] Reference Figure 18 , the apparatus 100 may generate (reference) motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture (operation 1810). For example, operation 1810 may be performed in the motion information generator 130 of the apparatus 100, and reference Figure 17 the description related to the motion information generator 130 given is included herein by reference.
[0257] The apparatus 100 may decode a current picture (operation 1820) based on the (reference) motion information generated in operation 1810 and at least one piece of syntax information obtained from the bitstream. More specifically, in operation 1820, the apparatus 100 may obtain motion information of a current block based on the motion information (or motion information map) generated in operation 1810 and at least one piece of syntax information obtained from the bitstream, and may reconstruct the current block based on the obtained motion information of the current block.
[0258] For example, reconstructing the current block by the apparatus 100 in operation 1820 may include determining whether the current block is decoded in an inter mode or an intra mode based on the motion information of the current block (e.g., prediction mode information (e.g., pred_mode_flag or CuPredMode)). For example, when the current block is decoded in the inter mode, reconstructing the current block by the apparatus 100 in operation 1820 may include performing inter prediction on the current block based on the motion information of the current block (e.g., motion vector information and prediction list utilization information) to obtain a predicted sample (predicted block) of the current block. For example, when the current block is decoded in the intra mode, reconstructing the current block by the apparatus 100 in operation 1820 may include performing intra prediction on the current block based on the intra prediction mode information of the current block (e.g., intra DC mode, intra planar mode, and intra angular mode) to obtain a predicted sample (or predicted block) of the current block. For example, in operation 1820, reconstructing the current block by the apparatus 100 in operation 1820 may include reconstructing the current block based on the obtained predicted sample (or predicted block).
[0259] Figure 19 An image encoding apparatus 200 according to the present disclosure is shown. The image encoding apparatus 200 according to the present disclosure is not limited to Figure 19 the examples, and may include Figure 19 other components not shown in Figure 19 For example, the image encoding apparatus 200 may include at least one processor, and
[0260] Referring to Figure 19 , the apparatus 200 may include a motion information generator 130 and an encoder 220. The motion information generator 130 may operate in the same manner as the Figure 17 motion information generator 130. Thus, the detailed description of the motion information generator 130 includes Figure 17 and its related description as a reference.
[0261] The encoder 220 may determine motion information of a current block in a current picture corresponding to an original picture based on rate-distortion (RD) optimization, and may encode the determined motion information of the current block into a bitstream based on the motion information (or motion information map) generated by the motion information generator 130.
[0262] For example, the encoder 220 may determine prediction mode information (e.g., pred_mode_flag or CuPredMode) indicating whether the current block is encoded / decoded in an intra mode or an inter mode based on RD optimization by using pixel data of a reference picture and / or a current block and reconstructed neighboring blocks of the current picture. For example, when the current block is encoded / decoded in the intra mode, the encoder 220 may identify motion vector information (e.g., a first motion vector and / or a second motion vector) obtained by performing motion estimation (ME) by using pixel data of the current block and the reference picture and prediction list utilization information (e.g., list 0 utilization information and / or list 1 utilization information). For example, based on the motion information (or motion information map) generated by the motion information generator 130, the encoder 220 may determine whether to encode the motion information of the current block determined based on RD optimization in an AMVP mode, a merge mode, or a skip mode, and may encode the motion information of the current block into the bitstream based on the determined mode.
[0263] The bitstream generated by the apparatus 200 may be stored in a storage medium or a recording medium, or may be transmitted via a wired or wireless network in units of network adaptation layer (NAL).
[0264] Figure 20 is a flowchart of an image encoding method 2000 according to the present disclosure. Figure 20 The illustrated method 2000 may be performed by the apparatus 200. Similarly, the bitstream generated by the method 2000 may be stored in a storage medium or a recording medium, or may be transmitted via a wired or wireless network in units of NAL.
[0265] Reference Figure 20 , the apparatus 200 may generate (reference) motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture (operation 2010). For example, operation 2010 may be performed in the motion information generator 130 of the apparatus 200, and Figure 19 and its related description (i.e., Figure 17 and its related description) are included herein by reference.
[0266] Device 200 may encode the current picture into a bitstream based on the motion information generated in operation 2010 (operation 2020). For example, device 200 may determine the motion information of the current block included in the current picture by using the pixel data of the original image (i.e., the current picture) based on RD optimization. As described above with reference to Figure 19 , the motion information of the current block may include prediction mode information (e.g., pred_mode_flag or CuPredMode), motion vector information (e.g., the first motion vector and / or the second motion vector), and prediction list utilization information (e.g., list 0 utilization information and / or list 1 utilization information). Device 200 may encode the determined motion information of the current block into the bitstream based on the motion information generated in operation 2010 (operation 2020). In operation 2020, device 200 may determine the motion information of the current block by performing the operations of Figure 19 encoder 220 and may encode the motion information of the current block into the bitstream.
[0267] FIG. 21 shows a method for generating motion information according to the present disclosure.
[0268] In the example of FIG. 21, a method for generating motion information (or a motion information map) 2140 corresponding to the current picture by using an artificial neural network 2110 is shown. However, in the method according to the present disclosure, another method (e.g., arithmetic operation) may be used instead of the artificial neural network to generate motion information. In addition, for ease of explanation, the example of FIG. 21 shows that the pixel data and / or the motion information map of two reconstructed pictures (e.g., the current reconstructed picture and / or at least one reconstructed picture) are shown as being input into an artificial neural network (ANN) (abbreviated as neural network (NN)) 2110. However, the number of reconstructed pictures input into NN 2110 is not limited to two, and various numbers of reconstructed pictures and their combinations may be used as the input to NN 2110.
[0269] The example of FIG. 21 may be executed by image decoding device 100 and image encoding device 200, and as a more specific example, may be executed by motion information generator 130. Moreover, the example of FIG. 21 may be executed in operation 1810 of image decoding method 1800 and operation 2010 of image encoding method 2000.
[0270] Referring to FIG. 21(a), the input to the NN 2110 may include the pixel data of the reconstructed picture (e.g., POC = T-1 or reference picture) and the pixel data 2130 of the current reconstructed picture (e.g., POC = T). The NN 2110 may be configured to generate the generated motion information map 2150 by using the pixel data of the reconstructed picture 2120 and the current reconstructed picture 2130 as inputs. In the example of FIG. 21(a), the generated motion information map 2150 may be used as the motion information map of the next picture to be coded / decoded, or may be used to refine the motion information map of the current picture.
[0271] Referring to FIG. 21(b), when the current picture (e.g., POC = T) is a bi-predicted picture, the input to the NN 2110 may include the motion information map 2120 stored in the neighboring forward reference picture (e.g., POC = T-1) and the motion information map 2140 stored in the backward reference picture (e.g., POC = T+1). Alternatively, when the current picture (e.g., POC = T) is a bi-predicted picture, the input to the NN 2110 may include the motion information (or motion information map) stored in each reference picture by using the reference pictures in a specific order in each of reference list 0 (or reference picture list 0) and reference list 1 (or reference picture list 1) as the forward reference picture and the backward reference picture (e.g., the first reference picture in each reference list or the reference picture with reference index 0 in each reference list).
[0272] Alternatively, in the example of FIG. 21(b), the backward reference picture may be replaced with the reference picture in reference list 0, and the forward reference picture may be replaced with the reference picture in reference list 1. In other words, when the current picture (e.g., POC = T) is a bi-predicted picture, the input to the NN 2110 may include the motion information (or motion information map) stored in each reference picture by using the reference pictures in a specific order in each of reference list 1 (or reference picture list 1) and reference list 0 (or reference picture list 0) as the forward reference picture and the backward reference picture (e.g., the first reference picture in each reference list or the reference picture with reference index 0 in each reference list).
[0273] Referring to FIG. 21(c), when the current picture is a uni-predicted picture, the input to the NN 2110 may include the motion information (or motion information map) of two reconstructed pictures that are close to the current picture in terms of the POC distance between the reconstructed pictures (e.g., reference pictures) available in the current picture. In the example of FIG. 21(c), the backward reference picture may be replaced with the reconstructed picture (e.g., reference picture) closest to the current picture, and the forward reference picture may be replaced with the second closest reconstructed picture (e.g., reference picture) to the current picture.
[0274] Figure 22 A method for normalizing the input to the motion information generator 130 (e.g., NN 2110) according to the present disclosure is shown. Figure 22 The examples of are merely examples, and the method proposed by the present disclosure is not limited to Figure 22 the examples of.
[0275] In the present disclosure, when motion information (or a motion information map) is input to the motion information generator 130 (e.g., NN 2110), the motion information can be used by normalizing it to a distance K in units of POC. For example, K can be set to the POC distance of the current picture (or the difference from the nearest reconstructed picture (e.g., reference picture)).
[0276] More specifically, the method proposed by the present disclosure may include scaling the motion information (or motion information map) of the input reconstructed picture based on the POC distance (or difference) between the current picture and the reconstructed picture closest to the current picture (e.g., reference picture) and the POC distance (or difference) between the input reconstructed picture and the reference picture of the input reconstructed picture.
[0277] The normalization of the motion information (or motion information map) can be performed by apparatuses 100 and 200, and more specifically, can be performed by the motion information generator 130 of apparatuses 100 and 200. Moreover, the normalization of the motion information (or motion information map) can be performed in operation 1810 of the image decoding method 1800 and operation 2010 of the image encoding method 2000.
[0278] For example, referring to Figure 22 , when the POC of the current picture 2210 is 3, the POC of the nearest forward reference picture 2220 is 2, and the POC of the nearest backward reference picture 2230 is 4, the distance between the current picture and the nearest reconstructed picture is 1, so K can be 1. When performing bi - directional prediction on the forward reference picture 2220 (e.g., POC = 2), as an example of the value of the motion information at a specific position of the reference picture 2220, the forward MV indicates a reference picture 2240 with POC = 0 and MV value (4, 8), and the backward MV indicates a reference picture 2230 with POC = 4 and MV value (-10, -4). The motion information at the specific position has an MV with a POC distance of 2, so it can be scaled (e.g., 1 / 2) to a POC distance of 1, such that the forward MV can be normalized to (2, 4), and the backward MV can be normalized to (-5, -2). Then, the normalized motion information can be used as the input to the motion information generator 130 (e.g., NN 2110).
[0279] As another example, when motion information is input to the motion information generator 130, the POC information (or POC map) of each motion information of the reconstructed picture (e.g., reference picture) can also be input (e.g., see Figure 23 and its related description). The unit where the motion information is input can be a block of a specific size or can be M×N pixels. Specifically, the motion information can be stored in the reconstructed picture in units of 16×16, in units of 8×8, or in units of 4×4. For example, when storing the motion information in units of 8×8 and encoding / decoding an image of 1920×1080, only the motion information with a resolution of 240×135 needs to be processed. For example, when storing the motion information in units of 16×16, only the motion information with a resolution of 120×68 needs to be processed.
[0280] In addition, when there is no motion information in a specific unit (e.g., when encoding / decoding in the intra-frame mode at the corresponding position), the motion information can be filled with a specific value. Specifically, the unit where there is no motion information can be filled with a zero vector (e.g., see Figure 23 and its related description). Moreover, a map indicating whether there is motion information can be created and used as an input to the motion information generator 130 (e.g., NN 2110) (e.g., see Figure 23 and its related description).
[0281] Figure 23 Fig. shows a motion information map of the input according to the present disclosure.
[0282] In Figure 23 's example, for clarifying the explanation of the present disclosure, the explanation focuses on the motion vector information included in the motion information. However, the motion information map according to the present disclosure can include other information in addition to the motion vector information. As described above with reference to "the method and device proposed by the present disclosure" and Figure 17 as described, the input to the motion information generator 130 can include motion information, where the motion information includes at least one of the motion vector information, prediction list utilization information, and prediction mode information of the reconstructed picture, and at least one of the quantization information, prediction samples, residual samples, and distance information from the current picture of the reconstructed picture.
[0283] In addition, as described above, the motion information map of the reconstructed picture can include motion information in the form of a one-dimensional arrangement or a two-dimensional arrangement in units of blocks of a specific size or in units of M×N pixels (e.g., 4×4, 8×8, or 16×16).
[0284] In Figure 23In the example, the motion information diagrams of the reconstructed pictures 2120 and 2140 of FIG. 21 are assumed and described. However, even when using other reconstructed pictures, the present disclosure can be applied in the same / similar manner. Figure 23 The motion information diagram shown can be a motion information diagram used as an input to the motion information generator 130 (e.g., NN 2110).
[0285] Reference Figure 23 , each motion information diagram of the reconstructed pictures 2120 and 2140 of FIG. 21 may include motion vector diagrams 2310 and 2330, POC diagrams 2320 and 2340, and a diagram 2350 regarding the presence or absence of motion vectors. For example, the reconstructed picture 2120 may be a forward reference picture, and the reconstructed picture 2140 may be a backward reference picture.
[0286] In Figure 23 the example, the motion vector diagrams 2310 and 2330 may separately include a diagram 2310 for the L0 motion vector (or the first motion vector) and a diagram 2330 for the L1 motion vector (or the second motion vector). Alternatively, the motion vector diagrams 2310 and 2330 may include the L0 motion vector (or the first motion vector) and the L1 motion vector (or the second motion vector) as a single diagram. The motion vector diagrams 2310 and 2330 may include separate diagrams for the x and y components of the motion vector, or may include one diagram in the form of an (x, y) vector.
[0287] In Figure 23 the example, the POC diagrams 2320 and 2340 may include the POC values of the reference pictures indicated by the motion vectors of the motion vector diagrams 2310 and 2330 (or the reference pictures referenced by the motion vectors at specific positions of the reconstructed pictures). When the motion information is not normalized (or scaled) based on the POC distance (or difference) between the current picture and the nearest reconstructed picture as described above with reference to Figure 22 , the motion information diagram may include the POC diagrams 2320 and 2340. When the motion information is normalized (or scaled) based on the POC distance (or difference) between the current picture and the nearest reconstructed picture, the motion information diagram may not include the POC diagrams 2320 and 2340 (or the POC diagrams 2320 and 2340 may be omitted from the motion information diagram).
[0288] In Figure 23In the example, when there is no motion vector at a specific position of the motion vector maps 2310 and 2330 for reconstructing pictures 2120 and 2140 (e.g., when the position is encoded / decoded in the intra mode), the motion vector at this position can be filled with a specific value (e.g., a zero vector) as described above. To this end, the motion information maps of the reconstructing pictures 2120 and 2140 may include a map 2350 regarding the presence or absence of the motion vector. For example, the map 2350 may include values indicating the presence or absence of the motion vector in the direction mode. As a non-limiting example, when only the L0 motion vector (or the first motion vector) exists, the map 2350 may include a first value (e.g., 0), when only the L1 motion vector (or the second motion vector) exists, the map 2350 may include a second value (e.g., 1), when both the L0 motion vector and the L1 motion vector exist, the map 2350 may include a third value (e.g., 2), and when there is no motion vector, the map 2350 may include a fourth value (e.g., 3). Alternatively, as another example, a map regarding the presence or absence of the motion vector may be separately included for each of the motion vector maps 2310 and 2330. In this case, the map may include flag information having a value (e.g., 0 or 1) indicating whether the corresponding motion vector (L0 motion vector or L1 motion vector) exists.
[0289] Figure 24 Shows a generated motion information map according to the present disclosure.
[0290] In Figure 24 In the example, for clarifying the explanation of the present disclosure, the explanation focuses on the motion vector information included in the generated motion information. However, the motion information map according to the present disclosure may include other information in addition to the motion vector information. As described above with reference to "the methods and devices proposed by the present disclosure" and Figure 17 As described, the generated information or generated data according to the present disclosure may include motion information, the motion information includes at least one of motion vector information, prediction list utilization information, and prediction mode information, and at least one of intra encoding / decoding related information, residual encoding / decoding related information, illuminance change information, and statistical information.
[0291] In addition, as described above, the generated motion information map according to the present disclosure may include motion information in the form of a one-dimensional arrangement or a two-dimensional arrangement in units of blocks of a specific size or in units of M×N pixels (e.g., 4×4, 8×8, or 16×16).
[0292] In Figure 24 In the example, the generated motion information map 2150 of FIG. 21 is assumed and described. However, the present disclosure can be applied in the same / similar manner even when other motion information maps are generated. Figure 24The motion information map shown can be a motion information map generated as an input to a motion information generator 130 (e.g., NN 2110).
[0293] Reference Figure 24 , the generated motion information map 2150 may include a motion vector map 2410, a POC map 2420, and a map 2430 regarding the presence or absence of motion vectors. In Figure 24 's example, a motion vector map 2410, a POC map 2420, and a map 2430 regarding the presence or absence of motion vectors are shown separately for L0 motion vectors and L1 motion vectors. However, for L0 motion vectors and L1 motion vectors, one motion vector map 2410, one POC map 2420, and one map 2430 regarding the presence or absence of motion vectors are included.
[0294] In Figure 24 's example, the motion vector map 2410 may include separate maps for the x and y components of the motion vector, or may include one map in (x,y) vector form.
[0295] In Figure 24 's example, the POC map 2420 may include the POC values of the reference pictures indicated by the motion vectors of the corresponding motion vector map 2410. When the input motion information is normalized (or scaled) based on the POC distance (or difference) between the current picture and the nearest reconstructed picture as described above with reference to Figure 22 , the POC map 2420 may be omitted from the generated motion information map. Alternatively, when the motion information generator 130 (e.g., NN 2110) is configured to output normalized motion vector values, the POC map 2420 may be omitted from the generated motion information map.
[0296] In Figure 24In the example, for each reference list (or each reference picture list), a figure 2430 regarding the presence or absence of motion vectors may be output, or the presence or absence of motion vectors may be output as one figure for reference lists L0 and L1, or the direction pattern of the motion vectors may be expressed as a figure and output. For example, when only L0 motion vectors (or first motion vectors) exist, figure 2430 may include a first value (e.g., 0), when only L1 motion vectors (or second motion vectors) exist, figure 2430 may include a second value (e.g., 1), when both L0 motion vectors and L1 motion vectors exist, figure 2430 may include a third value (e.g., 2), and when no motion vectors exist, figure 2430 may include a fourth value (e.g., 3). Alternatively, as another example, figures regarding the presence or absence of motion vectors may be separately included for each of the L0 motion vectors and the L1 motion vectors. In this case, the figure may include flag information having values (e.g., 0 or 1) indicating whether the corresponding motion vectors (L0 motion vectors or L1 motion vectors) exist.
[0297] Figure 25 shows a motion information generation model and a learning method according to the present disclosure.
[0298] In the present disclosure, the artificial neural network 2110 configured to output the generated motion information may be referred to as a motion information generation model. The motion information generation model 2110 may include a convolutional neural network (CNN), a fully connected neural network (FCNN), or a combination of a CNN and an FCNN. In Figure 25 the example, the motion information generation model 2110 is shown as a CNN. However, the present disclosure is equally / similarly applicable even when other types of neural networks are used.
[0299] As referred to above Figure 25 The motion information generation model and the learning method described above may be executed by apparatuses 100 and 200, and more specifically, may be executed by the motion information generator 130 of apparatuses 100 and 200. Moreover, as referred to above Figure 25 The motion information generation model and the learning method described above may be executed in operation 1810 of the image decoding method 1800 and operation 2010 of the image encoding method 2000.
[0300] In addition, in Figure 25 the example, as shown in FIG. 21, the reconstructed picture (e.g., POC = T - 1 or reference picture) 2120 and the current reconstructed picture (e.g., POC = T) 2130 are used as the input reconstructed pictures. However, the present disclosure is equally / similarly applicable even when other reconstructed pictures are used as the input.
[0301] As referred to Figure 25(a), The motion information generation model 2110 according to the present disclosure can be trained using an objective function based on minimizing a residual signal (similar to the motion vector map (or MV map) calculation of existing video codecs (e.g., motion estimation that minimizes the residual image within a search range)).
[0302] However, existing neural networks (NNs) are trained using a dataset with labels (or ground truth labels) (e.g., synthetic videos), while the motion information generation model 2110 according to the present disclosure is trained in a direction that minimizes the residual signal 2530 between the deformed image 2520 and the existing image (e.g., the reconstructed picture 2130), where the deformed image 2520 is obtained by deforming the reconstructed picture 2120 using the obtained (sample-level) motion information 2510. Thus, the motion information generation model 2110 according to the present disclosure can be trained using real videos because the motion information generation model 2110 does not require labels (or ground truth labels). Additionally, the motion information generation model 2110 according to the present disclosure has the advantage that models trained based on supervised learning, such as existing neural networks (NNs), can also be implemented in a fine-tuning form.
[0303] For example, the objective function for training the motion information generation model 2110 can use the mean squared error between the reconstructed picture 2130 and the deformed image 2520, or the sum or average of the squares of the samples of the residual image 2530. In the present disclosure, the motion information at the sample unit or sample level can be referred to as optical flow or flow, and the motion information map at the sample unit or sample level can be referred to as a motion field or (optical) flow field.
[0304] Furthermore, the motion information generation model 2110 according to the present disclosure can also estimate an occlusion mask by deformation using the motion information generated at the sample unit or sample level. The activated part of the occlusion mask may mean a part where it is difficult to accurately determine the presence or absence of motion information. Thus, when using the motion information generation model 2110 according to the present disclosure, blocks where the occlusion masks overlap can contribute to performance improvement by setting a mode that does not use motion information (or motion information map or MV map). For example, a block determined to be an occlusion area (or a block where the occlusion masks overlap) can be determined as a block encoded and decoded in the intra-frame mode.
[0305] Reference Figure 25(b), in the existing method, a search range is set, and a block with the smallest error from the current block is found within the search range to calculate the motion vector. Since it is necessary to repeat the process of obtaining a block of the same size as the current block by performing interpolation, etc. in units of the motion vector (for example, in units of 1 pixel, 1 / 2 pixel, or 1 / 4 pixel) within the search range and calculating the error from the current block, the complexity is high and the amount of calculation is large. However, since a search range is set, an accurate motion vector may not be found.
[0306] On the other hand, when applying the motion information generation model 2110 according to the present disclosure, the pixel data and / or motion information of the entire picture can be used as input, so that more accurate motion information can be generated with lower complexity and used for encoding and decoding the current picture, thereby providing improved encoding and decoding performance or efficiency.
[0307] Figure 26 A method for generating motion information according to the present disclosure is shown.
[0308] Reference will be made to Figure 26 to describe a method for generating motion information (or a motion information map) by using the pixel data of the reconstructed picture as input in the motion information generation model 2110 (for example, see Figure 25 and its related description). A method for generating motion information described with reference to Figure 26 can be executed by apparatuses 100 and 200, and more specifically, can be executed by the motion information generator 130 of apparatuses 100 and 200. Moreover, the method for generating motion information described with reference to Figure 26 can be executed in operation 1810 of the decoding method 1800 and operation 2010 of the image encoding method 2000.
[0309] In Figure 26 's example, as shown in FIG. 21, the reconstructed picture (for example, POC = T - 1 or a reference picture) 2120 and the reconstructed picture (for example, POC = T) 2130 are used as the input reconstructed pictures. However, even when other reconstructed pictures are used as input, the present disclosure is applied in the same / similar manner.
[0310] In Figure 26 's example, the generation of the motion information map 2150 generated in units of 16×16 blocks is shown. However, this is only an example, and the motion information map 2150 can also be generated in units of other block sizes (for example, 4×4 or 8×8).
[0311] Reference is made to Figure 26, devices 100 and 200 can input the reconstructed pictures 2120 and 2130 into the motion information generation model 2110 to generate the generated motion information map 2510 (at the sample unit or sample level). For example, when the size (or resolution) of the reconstructed pictures 2120 and 2130 is H×W, the size (or resolution) of the generated motion information map 2510 (at the sample unit or sample level) can be the same as the input reconstructed pictures (e.g., H×W).
[0312] Devices 100 and 200 can convert the generated motion information map 2510 (at the sample unit or sample level) according to the motion information map format of the video codec (e.g., the image decoding method 1800 and the image encoding method 2000). For example, devices 100 and 200 can generate the generated motion information map 2150 for the video codec by performing aggregation, scaling, and / or format transformation (e.g., quantization) on the generated motion information map 2510 (at the sample unit or sample level) according to the output format of the motion information generation model 2110.
[0313] For example, in the present disclosure, aggregation may include adjusting the block size of the generated motion information map 2510 (at the sample unit or sample level) according to the block size of the motion information map 2150 used in the video codec (e.g., the image decoding method 1800 and the image encoding method 2000).
[0314] Figure 27 The process of aggregating motion information according to the present disclosure is shown.
[0315] Figure 27 The examples in are for illustrative purposes only, and the present disclosure can be equally / similarly applied to various block sizes.
[0316] As described above, in the present disclosure, the motion information at the sample unit or sample level can be referred to as optical flow or flow, and the motion information map at the sample unit or sample level can be referred to as a motion field or (optical) flow field. Therefore, in the present disclosure, the motion information at the sample unit or sample level generated by the motion information generation model 2110 or the output of the motion information generation model 2110 can be referred to as neural network (NN) motion information (MI) or neural network (NN) motion information (MI) field or field flow.
[0317] Refer to Figure 27, when the generated motion information map 2150 includes motion information in units of 16×16 block sizes, and the motion information generation model 2110 outputs or generates a motion information map (or NN MI field or flow) 2510 in units of 8×8 blocks, apparatuses 100 and 200 may aggregate 2×2 motion information (or flow) into one piece of motion information in the motion information map 2510. For example, the aggregation may include processes such as fixed-position sampling, median filtering, etc. For example, fixed-position sampling may include determining the motion information at a specific position (e.g., the upper left end, the upper right end, the lower left end, or the lower right end) among multiple pieces of motion information to be aggregated as the aggregated motion information. For example, median filtering may include determining the median among multiple pieces of motion information to be aggregated as the value of the aggregated motion information.
[0318] Figure 27 An example is shown in which apparatuses 100 and 200 determine the motion information at a specific position (e.g., the upper left part) as the aggregated motion information according to fixed-position sampling.
[0319] Return reference Figure 26 , in contrast, when the block size or unit of the motion information map (or NN MI field or flow) 2510 is larger than the block size of the generated motion information map 2150, apparatuses 100 and 200 may perform upsampling on the motion information map (or NN MI field or flow) 2510. For example, when the generated motion information map 2150 includes motion information in units of 8×8 block sizes, and the motion information generation model 2110 outputs or generates a motion information map (or NN MI field or flow) 2510 in units of 8×8 block sizes, apparatuses 100 and 200 may upsample one piece of motion information (or flow) to 2×2 motion information. For example, upsampling may be performed by directly copying or repeating the values of the motion information (or flow), but various other types of upsampling (e.g., bicubic interpolation and DCT-IF interpolation) may also be used.
[0320] Reference Figure 26 , apparatuses 100 and 200 may scale the motion information (or NN MI) of the generated motion information map 2510 (at the sample unit or sample level). For example, the scaling according to the present disclosure may convert the resolution (or scale) of the motion information (or NN MI) of the generated motion information map 2510 (at the sample unit or sample level) according to the motion vector resolution (or motion scale) of a video codec (e.g., the image decoding method 1800 and the image encoding method 2000).
[0321] For example, although existing video codecs (e.g., HEVC (High Efficiency Video Coding)) are designed to store motion vectors in units of 1 / 4 pixels, the motion information (or NN MI) of the generated motion information map 2510 can be output or generated in units of one pixel. Therefore, in this case, apparatuses 100 and 200 can perform scaling (e.g., quadrupling of resolution or scale) on the motion information (or NN MI) of the generated motion information map 2510. In this example, apparatuses 100 and 200 can perform 4-fold scaling by shifting the motion information (or NN MI) of the generated motion information map 2510 2 bits to the left.
[0322] Reference Figure 26 , apparatuses 100 and 200 can perform format conversion on the motion information (or NN MI) of the generated motion information map 2510 (at the sample unit or sample level). For example, since the weights of the motion information generation model 2110 can have values in floating-point form, the output of the motion information generation model 2110 can have values in floating-point form. On the other hand, in a video codec (e.g., the image decoding method 1800 and the image encoding method 2000), motion information can be stored in integer form (e.g., 16-bit short integer). Accordingly, apparatuses 100 and 200 can convert the format of the output motion information (or NN MI) of the motion information generation model 2110 according to the motion information storage form (or format) of the video codec.
[0323] For example, although existing codecs (e.g., HEVC) are designed to store motion vectors in the form of 16-bit short integers, the motion information (or NN MI) of the generated motion information map 2510 can be output or generated in floating-point form. Therefore, in this case, apparatuses 100 and 200 can perform quantization on the motion information (or NNMI) of the generated motion information map 2510. For example, quantization can include mapping a value represented in floating-point form to the nearest integer value.
[0324] Figure 28 A method of reconstructing a reduced input picture according to the present disclosure is shown.
[0325] In the case of a high-definition image, if the image size or resolution of the high-definition image can be reduced before the high-definition image is input into the motion information generation model 2110, there may be benefits in terms of complexity and processing time. For example, in the case of a typical convolution operation performed in a convolutional neural network (CNN), the computational complexity may increase proportionally to the input image size. That is, when the input image size is H×W, the computational complexity can increase proportionally to the product of the width W and height H of the input image (i.e., proportionally to O(HW)). Therefore, when the width and height of the input image are both reduced to 1 / N, a reduction in computational complexity of 1 / (N×N) can be expected.
[0326] Accordingly, in the present disclosure, the apparatuses 100 and 200 may perform downsampling (DN) on the input reconstructed pictures 2120 and 2130 (each in terms of its respective width and respective height) by a downsampling factor (e.g., 1 / N), and input the downsampled reconstructed pictures 2810 and 2820 into the motion information generation model 2110 to generate motion information 2510. For example, the downsampling may use bilinear interpolation and Lanczos interpolation, but various other downsampling methods may also be used.
[0327] When the input reconstructed pictures 2120 and 2130 are downsampled as in the present disclosure and used as the input to the motion information generation model 2110, the output of the motion information generation model 2110 may be used as is, without resizing (or resizing the flow field) the motion information map 2510 generated by the motion information generation model 2110 (e.g., see the related descriptions in Figure 25 and Figure 26 ), such as aggregation, so that the computational complexity and processing period can be further reduced.
[0328] The apparatuses 100 and 200 may perform scaling on the motion information 2510 (or NN MI) generated using the downsampled input image (e.g., see Figure 25 and its related descriptions). For example, in addition to scaling the resolution (or scale) of the motion information (or NNMI) according to the video codec, the apparatuses 100 and 200 may also scale the motion information based on the downsampling factor. For example, in the case of motion information unit scaling (MI unit scaling), the apparatuses 100 and 200 may perform scaling on the motion information 2510 (or NN MI) generated by reflecting the downsampling factor (e.g., multiplying by the reciprocal of the downsampling factor) in the existing scaling factor. More specifically, in Figure 26 and Figure 28In the example, since 1 / 8 downsampling is applied to the reconstructed picture of the input and the generated motion information 2510 is scaled by 4 according to the video codec, apparatuses 100 and 200 can perform 32 (= 4 * 8) times scaling on the generated motion information 2510 (or NN MI), and apparatuses 100 and 200 can perform 32 times scaling through a 5-bit left shift operation.
[0329] Apparatuses 100 and 200 can also perform format conversion (e.g., see Figure 25 and its related description) on the motion information 2510 (or NN MI) generated using the downsampled input image, such as quantization.
[0330] In Figure 28 the example, it is shown that the reconstructed picture of the input has a size and / or resolution of 4K (3840×2160), and both the width and height of the reconstructed picture of the input are downsampled by 1 / 8. However, the present disclosure is not limited to this example, and the present disclosure can be applied in the same / similar manner even when using a reconstructed picture of the input with another size (e.g., 2K or 8K) or applying another downsampling factor (e.g., 1 / 2, 1 / 4, or 1 / 16).
[0331] In the present disclosure, when the time interval between the reconstructed pictures used as the input of the motion information generation model or the time interval between the current picture and the reconstructed picture is large (or they are far apart from each other in terms of time), the accuracy of the generated motion information may decrease. Therefore, in the present disclosure, when the time interval between the reconstructed pictures or the time interval between the current picture and the reconstructed picture is greater than or equal to a specific value or a specific threshold, apparatuses 100 and 200 can determine that the time interval between the two pictures (or frames) is "far".
[0332] For example, the specific threshold can be 1 / 25 second. Since 25 frames per second (FPS) is the frame rate of a typical video, 1 / 25 second can be the frame interval typically used for learning by the motion information generation model 2110. Therefore, when the specific threshold can be 1 / 25 second or higher, this may be the case where the specific threshold is equal to or greater than the interval between the pictures (or frames) of a typical video, and thus the interval between the two pictures (or frames) can be determined as "far". In the present disclosure, the threshold for determining "far" is not limited to 1 / 25 second, and various other values can be used.
[0333] Alternatively, the picture order count (POC) distance (or difference) instead of the time unit can be used as a specific threshold. For example, when the POC distance (or difference) is 4 or greater, apparatuses 100 and 200 can determine that the time interval between two pictures (or frames) is "far". Similarly, in the present disclosure, the POC distance (or difference) used to determine "far" is not limited to 4, and various other values can be used.
[0334] When the time interval between two pictures (or frames) is determined to be "far", the following rules can be added when applying the motion information (or motion information map) generated by the motion information generation model 2110 to the encoding or decoding (e.g., decoding or encoding) of the current picture, and other rules can be applied to each mode of the current block of the current picture (or frame). For example, when the current block is encoded or decoded in the skip mode and the time interval between the current picture and the reconstructed picture or the time interval between the reconstructed pictures is determined to be "far", apparatuses 100 and 200 may not use the generated motion information when encoding or decoding the current block. For example, when the current block is encoded or decoded in the merge mode and the time interval between the current picture and the reconstructed picture or the time interval between the reconstructed pictures is determined to be "far", apparatuses 100 and 200 may exclude the generated motion information from the merge candidates. For example, when the current block is encoded or decoded in the AMVP mode and the time interval between the current picture and the reconstructed picture or the time interval between the reconstructed pictures is determined to be "far", apparatuses 100 and 200 may add or insert the generated motion information as a candidate into the motion vector prediction (MVP) candidate list.
[0335] Figure 29 Reference picture resampling (RPR) is shown. RPR refers to a technique of encoding or decoding a current picture by using a picture with a reference size or resolution different from that of the current picture, and is adopted in the ITU-T H.266 / VVC (Versatile Video Coding) standard.
[0336] Even when RPR is applied, the generated motion information (or motion information map) according to the present disclosure can be applied. When RPR is applied, apparatuses 100 and 200 can compare the sizes or resolutions of two pictures (or frames) with each other, and when the sizes or resolutions of the two pictures (or frames) are different from each other, a process of aligning the sizes or resolutions of the two pictures (or frames) can be performed.
[0337] In the present disclosure, apparatuses 100 and 200 may downsample two pictures (or frames) (e.g., reconstructed pictures 2120 and 2130) to a smaller size or resolution, and then may generate motion information (or a motion information map) 2510 by using the downsampled pictures. For example, the same method as the VVC standard may be used as the downsampling method. However, the present disclosure is not limited thereto, and various downsampling methods (e.g., bilinear interpolation and Lanczos interpolation) may be used. The method described above with reference to Figures 25 to 28 may be used as the method for generating the motion information (or the motion information map) 2510.
[0338] Apparatuses 100 and 200 may perform operations (e.g., aggregation, scaling, and format conversion) for replacing the generated motion information (or the motion information map) 2510 with the generated motion information map 2150 (e.g., see Figures 25 to 28 and its related description).
[0339] When it is necessary to convert the generated motion information map 2150 into a higher resolution for encoding and decoding pictures with a higher resolution, apparatuses 100 and 200 may perform upsampling (UP) on the generated motion information map 2150. For example, the same method as the VVC standard may be used as the upsampling method. However, the present disclosure is not limited thereto, and various upsampling methods (e.g., bicubic interpolation or DCT-IF interpolation) may be used.
[0340] The method for generating the motion information (or the motion information map) 2150 according to the present disclosure may be applied not only to the encoding and decoding (e.g., decoding or encoding) of the current picture, but also to the frame rate up-conversion (FRUC) technique. For example, the method for generating the motion information (or the motion information map) 2150 according to the present disclosure may be applied to generating the motion information of the virtual reference picture generated according to the FRUC technique.
[0341] Figure 30 And FIG. 31 shows a method of utilizing the motion information generation method according to the present disclosure in FRUC.
[0342] Reference will be made to Figure 30 The method described with reference to
[0343] and FIG. 31 may be performed by apparatuses 100 and 200 in the image decoding method 1800 and the image encoding method 2000, and more specifically, may be performed by the motion information generator 130 of apparatuses 100 and 200. Figure 30, devices 100 and 200 can obtain motion information from two reconstructed pictures (e.g., L0 reference picture 3010 or L1 reference picture 3020), and can generate a motion information map (or MV map) of the target picture 3030 added by FRUC. For example, devices 100 and 200 can generate a motion information map 3050 by using the L0 reference picture 3010 and the L1 reference picture 3020 as inputs to a motion information generator 130 (e.g., motion information generation model 2110), and can perform interpolation or extrapolation on the generated motion vector map 2150 based on the temporal position of the target picture 3030 to generate a final motion information map 3060 of the target picture 3030.
[0344] In addition, devices 100 and 200 can refine the final motion information map 3060 to generate a refined final motion information map 3070. For example, devices 100 and 200 can generate a refined final motion information map 3070 by using not only the final motion information map 3060 as an input, but also the motion information map 3080 of the L0 reference picture 3010 and the motion information map 3090 of the L1 reference picture 3020 as inputs. For example, the refinement can include operations such as sampling and weighted averaging. Alternatively, for example, an artificial neural network model can be used to perform the refinement, and in this case, devices 100 and 200 can generate a refined final motion information map 3070 by using at least one of the final motion information map 3060, the motion information map 3080 of the L0 reference picture 3010, and the motion information map 3090 of the L1 reference picture 3020 as an input. For example, the artificial neural network can include a convolutional neural network (CNN), a fully connected neural network (FCNN), or a combination thereof.
[0345] FIG. 31(a) shows generating a final motion information map 3060 by performing interpolation on the generated motion information map 3050 when the target picture 3030 is between the L0 reference picture 3010 and the L1 reference picture 3020, and FIG. 31(b) shows generating a final motion information map 3060 by performing extrapolation on the generated motion information map 3050 when the target picture 303 is outside the L0 reference picture 3010 and the L1 reference picture 3020. Although FIG. 31(b) shows performing extrapolation on the generated motion information map 3050 when both the L0 reference picture 3010 and the L1 reference picture 3020 are temporally in front of the target picture 3030, a final motion information map 3060 can also be generated by performing extrapolation on the generated motion information map 3050 even when both the L0 reference picture 3010 and the L1 reference picture 3020 are temporally behind the target picture 3030.
[0346] Through the above reference Figure 30The motion information map 3060 generated by the method described in FIGS. 30 and 31 can be used to encode or decode (e.g., decode or encode) the target picture 3030. For example, when the block to be encoded or decoded in the target picture 3030 is encoded or decoded in the skip mode, apparatuses 100 and 200 can determine the motion information of the block at the position corresponding to the block in the generated motion information map 3060 as the motion information of the block. For example, when the block to be encoded or decoded in the target picture 3030 is encoded or decoded in the merge mode, apparatuses 100 and 200 can determine the motion information of the block by adding or inserting the motion information at the position corresponding to the block in the motion information map 3060 as a candidate into the merge candidate list. For example, when the block to be encoded or decoded in the target picture 3030 is encoded or decoded in the AMVP mode, apparatuses 100 and 200 can determine the motion information of the block by adding or inserting the motion information at the position corresponding to the block in the motion information map 3060 as an MVP candidate into the MVP candidate list.
[0347] FIG. 32 shows a method for refining the generated motion information map according to the present disclosure.
[0348] Apparatuses 100 and 200 can perform refinement on the generated motion information (or motion information map) 2150 and 3060 to improve the accuracy of the generated motion information (or motion information map) 2150 and 3060. Apparatuses 100 and 200 can include a motion information refiner 3230 configured to perform refinement on the generated motion information (or motion information map) 2150 and 3060 according to the present disclosure, and the motion information refiner 3230 can be implemented by at least one processor.
[0349] In the present disclosure, the motion information generator 3230 can operate in the same manner as the motion information generator 130. For example, the motion information refiner 3230 can generate refined motion information 3220 based on the generated motion information (or motion information map) 2150 and 3060 and additionally based on at least one of the pixel data and / or motion information of at least one reconstructed picture. The motion information refiner 3230 can process the input and generate the output according to various methods. For example, the motion information refiner 3230 can generate refined motion information 3220 from the input information by using arithmetic operations (e.g., average, median, interpolation, or filtering), or can generate refined motion information 3220 from the input information by using an artificial neural network 3210. For example, the artificial neural network 3210 can use the same artificial neural network model as the artificial neural network (or motion information generation model 2110) of the motion information generator 130, or can use a separate artificial neural network model.
[0350] As referred to above Figures 17 to 3The description of the motion information generator 130 or the motion information generation model 2110 given above can be equivalently applied to the motion information refiner 3230 or the artificial neural network model used in the motion information refiner 3230. When an artificial neural network is used separately from the motion information generator 130 or the motion information generation model 2110 in the motion information refiner 3230, the motion information generation model 2110 of the motion information generator 130 can be referred to as a first motion information generation model or a first artificial neural network, and the motion information generation model 3210 of the motion information refiner 3230 can be referred to as a second motion information generation model or a second artificial neural network, so as to distinguish the two artificial neural networks from each other in the present disclosure. The above reference Figures 17 to 3 The description of the motion information generator 130 or the motion information generation model 2110 given above is incorporated herein by reference.
[0351] Referring to FIG. 32(a), the input of the motion information refiner 3230 may include the motion information maps 2150 and 3060 generated by the motion information generator 130 or the motion information generation model 2110. In addition, the input of the motion information refiner 3230 may further include additional information. For example, the additional information may include pixel data of at least one reconstructed picture (e.g., a reference picture) or at least one of the motion information maps.
[0352] As a non-limiting example, the input of the motion information refiner 3230 may be the following combinations.
[0353] - Motion information map of the current picture (or current MV map)
[0354] - Motion information map of the reference picture (or reference MV map)
[0355] - Motion information map of the current picture (or current MV map) + motion information map(s) of the reference picture(s) (or reference MV map(s))
[0356] - Motion information map of the current picture (or current MV map) + motion information map(s) of the reference picture(s) (or reference MV map(s)) + pixel data of the current reconstructed picture (or current reconstructed picture) + pixel data of the reference picture(s) (or reference picture(s))
[0357] - Motion information map of the current picture (or current MV map) + pixel data of the current reconstructed picture (or current reconstructed picture) + pixel data of the reference picture(s) (or reference picture(s))
[0358] - Motion information map of the reference picture (or reference MV map) + pixel data of the reference picture (or reference picture)
[0359] As an example without limiting the present disclosure, FIG. 32(b) shows that the refined motion information map 3220 is generated by using the motion information maps 2150 and 3060 of the current picture, the pixel data of the current reconstructed picture 2130, and the pixel data of the reference picture 2120 as the inputs of the motion information generation model 3210 (or the second motion information generation model).
[0360] In addition, for example, in addition to the motion information, the input of the motion information refiner 3230 may further include at least one of quantization information, prediction samples, residual samples, and distance information from the current picture (for example, see Figure 17 and its related descriptions). In addition, for example, in addition to the refined motion information, the output of the motion information refiner 3230 may further include at least one of intra-coding / decoding related information, residual coding / decoding related information, illumination change information, and statistical information (for example, see Figure 17 and its related descriptions).
[0361] The motion information (or motion information map) generated in the present disclosure can be applied to video coding / decoding (e.g., decoding or encoding) in various ways.
[0362] - Before coding / decoding a picture (or frame), the motion information (or motion information map) of the picture (or frame) for which video coding / decoding will be performed can be generated in advance, and the generated motion information (or motion information map) can be used for coding / decoding the picture (or frame).
[0363] For example, when the current block of the picture to be coded / decoded is coded / decoded in the merge mode or the AMVP mode, the motion information at the position corresponding to the current block in the generated motion information map can be used as a merge candidate or an MVP candidate. For example, when the current block is coded / decoded in the skip mode, the motion information at the position corresponding to the current block in the generated motion information map can be determined as the motion information of the current block.
[0364] - The motion information map of the current picture (or the current MV map) can be replaced with the generated motion information map, and the generated motion information map can be used. For example, the motion information of the generated motion information map can be used as temporal motion information during subsequent picture (or frame) coding / decoding (e.g., decoding or encoding).
[0365] - When applying FRUC, the generated motion information map can be used to generate a virtual reference picture of the picture (or frame) to be coded / decoded.
[0366] - The generated motion information map can be used when generating a pseudo-picture (or frame) of a lost picture (or frame).
[0367] A method of using the generated motion information (or motion information map) during video coding or decoding (e.g., decoding or encoding) will now be described in more detail.
[0368] In the present disclosure, new syntax information is defined to indicate whether the generated motion information (or motion information map) is used for coding or decoding the current picture. For ease of explanation, this new syntax information may be referred to as the first information. Additionally, in the present disclosure, although the first information is represented as use_generated_mv_flag, the name of the first information may change.
[0369] For example, according to the value of the first information, the first information may indicate whether the generated motion information (or motion information map) is used during coding or decoding (e.g., decoding or encoding) of the current block. For example, when the value of the first information is 0, the first information may indicate that the generated motion information (or motion information map) is not used during coding or decoding (e.g., decoding or encoding) of the current block, while when the value of the first information is 1, the first information may indicate that the generated motion information (or motion information map) is used during coding or decoding (e.g., decoding or encoding) of the current block. As another example, the value of the first information may be set in the opposite way (i.e., when the value of the first information is 0, the first information may indicate that the generated motion information is used during coding or decoding of the current block, while when the value of the first information is 1, the first information may indicate that the generated motion information is not used during coding or decoding of the current block).
[0370] For example, the current block may be a coding unit (CU) (or coding block), and the first information may be signaled via a bitstream on a per-CU basis. As another example, the current block may be a coding tree unit (CTU) (or coding tree block), may be a transform unit (TU) (or transform block), or may be a prediction unit (PU) (or prediction block), and the first information may be signaled via a bitstream on a per-CTU, per-TU, or per-PU basis. As another example, the first information may be signaled on a per-sequence parameter set (SPS), per-picture parameter set (PPS), per-picture header (PH), or per-slice header (SH) basis.
[0371] For ease of explanation, it will now be assumed that the first information is signaled on a per-CU basis.
[0372] Reference Figure 33When decoding the current picture, the apparatus 100 may obtain first information for the current block from the bitstream (operation 3310).
[0373] When the first information indicates that the generated motion information (or motion information map) is not used for encoding / decoding the current block (e.g., when the value of the first information is 0), the apparatus 100 may reconstruct or decode the current block by using the syntax information of the existing video coding standard. In the present disclosure, the existing video coding standard may refer to the ITU-T H.266 or Versatile Video Coding (VVC) standard, or the ITU-T H.265 or High Efficiency Video Coding (HEVC) standard. The syntax information of the existing video coding standard is described in detail in the ITU-T H.266 / VVC standard document or the ITU-T H.265 / HEVC standard document, and this specification includes the entire content of these standard documents by reference.
[0374] When the first information indicates that the generated motion information (or motion information map) is used for encoding / decoding the current block (e.g., when the value of the first information is 1), the apparatus 100 may determine whether there is generated motion information at the position corresponding to the current block. When it is determined that there is no generated motion information at the position corresponding to the current block, the apparatus 100 may determine the intra prediction mode information (e.g., intra DC mode, intra planar mode, and intra angular mode) of the current block by using multiple intra mode-related syntax information of the existing video coding standard, and perform intra prediction on the current block to obtain prediction samples (operation 3330).
[0375] When it is determined that there is generated motion information at the position corresponding to the current block, the apparatus 100 may determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information of the current block, and may perform inter prediction on the current block to obtain prediction samples (operation 3340). In operation 3340, the apparatus 100 may not obtain the relevant syntax information from the bitstream (or may omit obtaining the relevant syntax information from the bitstream).
[0376] After obtaining prediction samples (or a prediction block) through intra prediction or inter prediction, the apparatus 100 may perform inverse transform and inverse quantization on the current block by using the residual coding-related syntax information of the existing video coding standard to obtain residual samples, and may reconstruct the current block based on the prediction samples and the residual samples.
[0377] In Figure 33In the example of , when encoding the current picture, the apparatus 200 may perform operations corresponding to those of the above-described apparatus 100. For example, the apparatus 200 may determine whether to encode the current block using the generated motion information (or motion information map) regarding the current picture, and whether to encode the current block in the intra mode or in the inter mode.
[0378] Based on the determined result, the apparatus 200 may encode the first information into the bitstream (operation 3310). For example, when it is determined that the generated motion information (or motion information map) is not used to encode the current block, the apparatus 200 may encode the first information into the bitstream with a value of 0, and may encode the syntax information of the existing video coding and decoding standard into the bitstream according to the result of encoding the current block (operation 3320).
[0379] For example, when it is determined that the generated motion information (or motion information map) is used to encode the current block, the apparatus 200 may encode the first information into the bitstream with a value of 1, and may determine whether there is generated motion information at the position corresponding to the current block in the generated motion information map.
[0380] When it is determined that there is no generated motion information at the position corresponding to the current block, the apparatus 100 may determine the intra prediction mode information of the current block, perform intra prediction on the current block to obtain prediction samples, and may encode multiple syntax information (related to the intra mode of the existing video coding and decoding standard) indicating the determined intra prediction mode into the bitstream (operation 3330).
[0381] When it is determined that there is generated motion information at the position corresponding to the current block, the apparatus 100 may determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information of the current block, and may perform inter prediction on the current block to obtain prediction samples (operation 3340). In operation 3340, the apparatus 100 may not encode the relevant syntax information into the bitstream (or may omit encoding the relevant syntax information into the bitstream).
[0382] The apparatus 200 may obtain prediction samples (or prediction blocks) through intra prediction or inter prediction, then perform residual coding and decoding according to the existing video coding and decoding standard, and encode the relevant syntax information into the bitstream (operation 3350).
[0383] As described above, the value of the first information may also be encoded in the opposite way.
[0384] The above reference Figure 33The operations of the described apparatus 100 may be performed in operation 1820 of the image decoding method 1800, and the operations of the apparatus 200 may be performed in operation 2020 of the image encoding method 2000.
[0385] In addition, as described above, the generated motion information (or generated motion information map) may be generated by the motion information generator 130 or the motion information refiner 3230. For example, the generated motion information (or generated motion information map) may be generated by an artificial neural network (e.g., the motion information generation models 2110 and 3210), or may be generated according to various methods that do not use an artificial neural network (e.g., arithmetic operations).
[0386] FIG. 34 shows a syntax structure and an encoding / decoding method according to the present disclosure.
[0387] The syntax structure of FIG. 34 is obtained by changing the syntax information from the Figure 33 syntax structure. Accordingly, Figure 33 the descriptions of operations 3310, 3320, 3330, and 3350 in the syntax structure and the encoding / decoding method are the same as those of FIG. 34 and are incorporated herein by reference.
[0388] In the present disclosure, when using the generated motion information (or motion information map), new syntax information indicating whether there is additionally a motion vector difference (MVD) in the bitstream may be defined. In the present disclosure, this syntax information may be referred to as second information. In the present disclosure, although the second information is represented as mvd_flag, the name of the second information may be changed.
[0389] For example, according to the value of the second information, the second information may indicate whether there is additionally syntax information related to the MVD in the bitstream when using the generated motion information (or motion information map). For example, when the value of the second information is 0, the second information may indicate that there is no syntax information related to the MVD in the bitstream, while when the value of the second information is 1, the second information may indicate that there is syntax information related to the MVD in the bitstream. As another example, the value of the second information may be set in the opposite way.
[0390] In addition, in the present disclosure, new syntax information for generating a candidate list and indicating one piece of motion information in the candidate list can be defined. The candidate list includes motion information at a position corresponding to the current block in the generated motion information map and motion information at positions corresponding to at least one (spatial) neighboring block of the current block. And this new syntax information can be referred to as third information in the present disclosure. In the present disclosure, although the third information is represented as mv_index, the name of the third information can be changed. For example, the value of the third information can be less than or equal to the maximum number of the candidate list, and the third information can indicate a motion information candidate at a position corresponding to the value of the third information among the candidate list.
[0391] Referring to FIG. 34(a), when it is determined that there is generated motion information at a position corresponding to the current block, the apparatus 100 may obtain second information (e.g., mvd_flag) from the bitstream (operation 3410).
[0392] When the second information indicates that there is MVD-related syntax information of the current block in the bitstream, the apparatus 100 may obtain the MVD-related syntax information from the bitstream and obtain the MVD of the current block based on the obtained information. In addition, the apparatus 100 may obtain a motion vector prediction (MVP) from the generated motion information map, obtain the motion vector of the current block based on the MVD obtained from the bitstream and the MVP obtained from the generated motion information map, and perform inter prediction on the current block by using the obtained motion vector to obtain predicted samples.
[0393] Moreover, in operation 3420, the third information (e.g., mv_index) may be signaled via the bitstream to indicate the MVP in the generated motion information map. In this case, the apparatus 100 may obtain the third information from the bitstream and may determine the motion information candidate indicated by the third information as the MVP among the motion information candidates in the generated motion information map.
[0394] When the second information indicates that there is no MVD-related syntax information of the current block in the bitstream, the apparatus 100 may determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information of the current block and may perform inter prediction to obtain the predicted samples of the current block (operation 3430). In operation 3430, the apparatus 100 may not obtain the related syntax information from the bitstream (or may omit obtaining the related syntax information from the bitstream). Thereafter, the apparatus 100 may reconstruct the current block by performing the operations described above with reference to 3350.
[0395] When encoding the current picture, the apparatus 200 may perform operations corresponding to those of the above-described apparatus 100. For example, when it is determined that there is generated motion information at a position corresponding to the current block, the apparatus 200 may determine whether to encode MVD-related syntax information into the bitstream, and may encode second information having a corresponding value into the bitstream based on the determined result (operation 3410). Based on the second information, the apparatus 200 may encode the MVD-related syntax information of the current block into the bitstream, and encode third information indicating the MVP into the bitstream (operation 3420), or may encode the current block by using the motion information at the position corresponding to the current block in the generated motion information map (in the inter prediction mode) (operation 3430). In operation 3430, the apparatus 200 may not encode the related syntax information into the bitstream (or may omit encoding). Thereafter, the apparatus 200 may generate a bitstream by performing the operations described above with reference to 3350.
[0396] Referring to FIG. 34(b), when it is determined that there is generated motion information at a position corresponding to the current block, the apparatus 100 may obtain third information from the bitstream (operation 3440). For example, the apparatus 100 may determine, as the motion information of the current block, the motion information indicated by the third information among the motion information at the position corresponding to the current block in the generated motion information map and the motion information at the positions corresponding to at least one adjacent block of the current block (operation 3450).
[0397] The apparatus 100 may perform inter prediction by using the motion information determined in operation 3450 to obtain a predicted sample. Thereafter, the apparatus 100 may reconstruct the current block by performing the operations described above with reference to 3350.
[0398] In FIG. 34(b), when encoding the current picture, the apparatus 200 may perform operations corresponding to those of the above-described apparatus 100. For example, when it is determined that there is generated motion information at a position corresponding to the current block, the apparatus 200 may determine, among the motion information at the position corresponding to the current block in the generated motion information map and the motion information at the positions corresponding to at least one adjacent block of the current block, the motion information corresponding to the motion information of the current block (operation 3450), and may encode third information indicating the determined motion information into the bitstream (operation 3440). Thereafter, the apparatus 200 may generate a bitstream by performing the operations described above with reference to 3350.
[0399] The operations of the apparatus 100 described above with reference to FIG. 34 may be performed in operation 1820 of the image decoding method 1800, and the operations of the apparatus 200 may be performed in operation 2020 of the image encoding method 2000.
[0400] New flag information can be determined such that one of the syntax structures and encoding / decoding methods of FIGS. 34(a) and 34(b) can be selected, thereby allowing selection of one of two modes. In the present disclosure, this flag information may be referred to as fourth information. In the present disclosure, although the fourth information is represented as no_mvd_flag, the name of the fourth information may change.
[0401] Accordingly, a more accurate motion vector (MV) corresponding to the current block can be found by sending only a small amount of side information, thereby improving the encoding / decoding efficiency.
[0402] Figure 35 A syntax structure and an encoding / decoding method according to the present disclosure are shown.
[0403] By changing the syntax information from the Figure 33 syntax structure in operation 3340, the syntax structure of FIG. 34 is obtained. Accordingly, Figure 33 the descriptions of operations 3310, 3320, 3330, and 3350 in the Figure 35 syntax structure and encoding / decoding method are the same as those of the
[0404] In Figure 35 the example, the fourth information (e.g., no_mvd_flag) can be used to select one of the methods of FIG. 34(a) and the method of FIG. 34(b). For example, according to the value of the fourth information, the fourth information may indicate whether to use the method of signaling MVD-related syntax information (or the method shown in FIG. 34(a)), or whether to use the method of determining the motion information indicated by the third information (e.g., mv_index) as the motion information of the current block among the generated motion information maps (or the method shown in FIG. 34(b)). For example, when the value of the fourth information is 1, the fourth information may indicate using the method of signaling MVD-related syntax information (or the method shown in FIG. 34(a)), while when the value of the fourth information is 0, the fourth information may indicate using the method of determining the motion information indicated by the third information (e.g., mv_index) as the motion information of the current block among the generated motion information maps (or the method shown in FIG. 34(b)).
[0405] As another example, when the value of the fourth information is 0, the fourth information may indicate using the method of signaling MVD-related syntax information (or the method shown in FIG. 34(a)), while when the value of the fourth information is 1, the fourth information may indicate using the method of determining the motion information indicated by the third information (e.g., mv_index) as the motion information of the current block among the generated motion information maps (or the method shown in FIG. 34(b)).
[0406] Based on the fourth information, apparatuses 100 and 200 may decode or encode the current block by performing the operations described above with reference to 3420 of FIG. 34(a) or performing the operations described above with reference to 3440 and 3450 of FIG. 34(b). The descriptions of operations 3420, 3440, and 3450 are included herein by reference.
[0407] Figure 36 Illustrates a syntax structure and encoding / decoding method according to the present disclosure.
[0408] By changing the syntax information from the Figure 33 syntax structure to obtain the Figure 36 syntax structure. Accordingly, Figure 33 the descriptions of operations 3310, 3320, and 3350 in the Figure 36 syntax structure and encoding / decoding method are the same as the
[0409] reference to Figure 36 and are incorporated herein by reference. When performing image decoding, apparatus 100 may determine, based on the first information (e.g., use_generated_mv_flag), whether the motion information (or motion information map) generated for the current picture is used for encoding / decoding the current block (operation 3310). When the first information indicates that the generated motion information (or motion information map) is used for encoding / decoding the current block (e.g., when the value of the first information is 1), apparatus 100 may determine that the current block is not encoded / decoded in the intra mode, determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information of the current block, and perform inter prediction on the current block to obtain prediction samples (operation 3610). In operation 3610, apparatus 100 may not obtain the relevant syntax information from the bitstream (or may omit obtaining the relevant syntax information from the bitstream). Thereafter, apparatus 100 may reconstruct the current block by performing the operations described above with reference to 3350.
[0410] When there is no motion information at the position corresponding to the current block in the generated motion information map, apparatus 100 may search for other motion information and use the found motion information to decode the current block.
[0411] Different from the Figure 33 example, in the Figure 36 example, when the current block is encoded / decoded in the intra mode, apparatus 100 may perform decoding according to the syntax information of the existing video encoding / decoding standard. That is, in the Figure 36 example, when the current block is encoded / decoded in the intra mode, the first information of the current block may be set to indicate that the generated motion information (or motion information map) is not used for encoding / decoding the current block.
[0412] In Figure 36 the example of, when encoding the current picture, the apparatus 100 may perform operations corresponding to the operations of the apparatus 100 described above. For example, the apparatus 200 may determine whether to encode the current block using the generated motion information (or motion information map) of the current picture, and may encode the first information into the bitstream based on the determined result.
[0413] When it is determined that the generated motion information (or motion information map) is used to encode the current block, the apparatus 200 may determine the motion information at the position corresponding to the current block within the generated motion information map as the motion information regarding the current block, and may perform inter prediction on the current block to obtain prediction samples (operation 3610). In operation 3610, the apparatus 200 may not encode the relevant syntax information into the bitstream (or may omit encoding). Thereafter, the apparatus 200 may generate the bitstream by performing the operations described above with reference to 3350.
[0414] The operations of the apparatus 100 described above with reference to Figure 36 may be performed in operation 1820 of the image decoding method 1800, and the operations of the apparatus 200 may be performed in operation 2020 of the image encoding method 2000.
[0415] FIG. 37 illustrates a syntax structure and an encoding / decoding method according to the present disclosure.
[0416] In the example of FIG. 37, the encoding / decoding method using the motion information (or motion information map) generated according to the present disclosure may be set to perform operations instead of the skip mode. Therefore, in the example of FIG. 37, only in the case of an inter picture (or P slice or B slice), the first information may be signaled via the bitstream. In addition, when the first information indicates that the generated motion information (or motion information map) is not used for encoding / decoding the current block, the mode-related syntax information of the video encoding / decoding standard excluding the existing skip mode may be used, and the existing residual encoding / decoding-related syntax information may also be used. On the other hand, in the example of FIG. 37, when the first information indicates that the generated motion information (or motion information map) is used for encoding / decoding the current block, the syntax information of the current block may not be signaled via the bitstream, and the existing residual encoding / decoding-related syntax information may also not exist in the bitstream.
[0417] In the example of FIG. 37(a), when decoding the current picture, the apparatus 100 may determine whether the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice) (operation 3710). When the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice), the apparatus 100 may obtain the first information of the current block from the bitstream (operation 3720).
[0418] When the first information indicates that the generated motion information (or motion information map) is not used for encoding / decoding the current block (e.g., when the value of the first information is 0), the apparatus 100 may exclude the skip mode of the existing video coding standard and reconstruct or decode the current block by using each relevant syntax information (operation 3730).
[0419] When the first information indicates that the generated motion information (or motion information map) is used for encoding / decoding the current block (e.g., when the value of the first information is 1), the apparatus 100 may determine that the current block is not encoded / decoded in intra mode, determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information of the current block, and obtain the samples of the current block (operation 3740). In operation 3740, the apparatus 100 may not obtain the relevant syntax information from the bitstream (or may omit obtaining the relevant syntax information from the bitstream). In operation 3740, no residual coding is performed.
[0420] When there is no motion information at the position corresponding to the current block in the generated motion information map, the apparatus 100 may search for other motion information and decode the current block by using the found motion information. In addition, in operation 3740, when the current block is encoded / decoded in intra mode, the first information of the current block may be set to indicate that the generated motion information (or motion information map) is not used for encoding / decoding the current block.
[0421] In the example of FIG. 37, when encoding the current picture, the apparatus 200 may perform operations corresponding to the operations of the apparatus 100 described above. For example, the apparatus 200 may determine whether the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice), and when the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice), may encode the first information into the bitstream. Of course, the apparatus 200 may determine whether to use the generated motion information (or motion information map) of the current picture for encoding the current block, and whether to encode the current block in intra mode or in inter mode.
[0422] When applying a mode (including an intra mode) after the exclusion skip mode to the current block, the apparatus 200 may encode the first information such that the first information indicates that the generated motion information (or motion information map) is not used for encoding the current block. On the other hand, when applying the skip mode to the current block, the apparatus 200 may encode the first information such that the first information indicates that the generated motion information (or motion information map) is used for encoding the current block.
[0423] For example, when determining that the generated motion information (or motion information map) is not used for encoding the current block, the apparatus 200 may encode the first information with a value of 0 into the bitstream (operation 3720), and may encode the syntax information of the existing video codec standard into the bitstream according to the result of encoding the current block (operation 3730).
[0424] For example, when determining that the generated motion information (or motion information map) is used for encoding the current block, the apparatus 200 may obtain the samples of the current block by using the motion information at the position corresponding to the current block within the generated motion information map (operation 3740). In operation 3740, the apparatus 200 may not encode the relevant syntax information into the bitstream (or may omit obtaining the relevant syntax information from the bitstream). In operation 3740, residual encoding and decoding are not performed.
[0425] FIG. 37(b) is obtained by changing operation 3740 of FIG. 37(a). Accordingly, the descriptions of operations 3710, 3720, and 3730 in the description of FIG. 37(a) are the same as the description of FIG. 37(b) and are incorporated herein by reference.
[0426] In the example of FIG. 37(b), when the first information indicates that the generated motion information (or motion information map) is used for encoding and decoding the current block, instead of using the motion information at the position corresponding to the current block in the generated motion information map without change, a method of indicating the motion information in the generated motion information map by signaling the third information via the bitstream may be used.
[0427] In the example of FIG. 37(b), when decoding the current picture, the apparatus 100 may determine whether the generated motion information map is used for encoding / decoding the current block based on the first information (operation 3720). When the first information indicates that the generated motion information map is used for encoding / decoding the current block (e.g., when the value of the first information is 1), the apparatus 100 may obtain the third information (e.g., mv_index) from the bitstream (operation 3750). For example, the apparatus 100 may obtain the samples of the current block by using the motion information indicated by the third information among the motion information at the position corresponding to the current block and the motion information at the positions corresponding to at least one neighboring block of the current block in the generated motion information map (operation 3760).
[0428] In FIG. 37(b), when encoding the current picture, the apparatus 200 may perform operations corresponding to the operations of the apparatus 100 described above. For example, when determining that the generated motion information map is used for encoding / decoding the current block, the apparatus 200 may determine the motion information corresponding to the motion information of the current block among the motion information at the position corresponding to the current block and the motion information at the positions corresponding to at least one neighboring block of the current block in the generated motion information map (operation 3760), and may encode the third information indicating the determined motion information into the bitstream (operation 3750).
[0429] The operations of the apparatus 100 described above with reference to FIG. 37 may be performed in operation 1820 of the image decoding method 1800, and the operations of the apparatus 200 may be performed in operation 2020 of the image encoding method 2000.
[0430] Figure 38 Illustrates a syntax structure and an encoding / decoding method according to the present disclosure.
[0431] In Figure 38 when the current block is encoded / decoded in the skip mode or in the merge mode, the first information (e.g., use_generated_mv_flag) may be signaled through the bitstream, and the merge mode may be applied to the current block, or the generated motion information map may be used according to the first information. In Figure 38 the example of, when the current block is encoded / decoded in the AMVP mode, the generated motion information map is not used, and the motion information is signaled in the AMVP mode according to the existing video encoding / decoding standards.
[0432] In Figure 38In the example, when decoding the current picture, the apparatus 100 may determine whether the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice) (operation 3810). When the current picture (or slice) is an inter picture (or a picture encoded / decoded in inter mode) (or a P slice or a B slice), the apparatus 100 may obtain information indicating whether the current block is in skip mode (e.g., skip_flag). When the value of the information indicating whether the current block is in skip mode is 0, the information may indicate that the current block is not in skip mode, and when the value of the information indicating whether the current block is in skip mode is 1, the information may indicate that the current block is in skip mode (and vice versa).
[0433] When the current block is not in skip mode, the apparatus 100 may determine whether the current block is an intra block or an inter block (or whether the current block is encoded / decoded in intra mode or inter mode), and may obtain information indicating whether the current block is in merge mode (e.g., merge_flag) from the bitstream (operation 3820). When the value of the information indicating whether the current block is in merge mode is 0, the information may indicate that the current block is not in merge mode, and when the value of the information indicating whether the current block is in merge mode is 1, the information may indicate that the current block is in merge mode (and vice versa).
[0434] When the current block is in merge mode, the apparatus 100 may obtain first information from the bitstream (operation 3830). When the first information indicates that the generated motion information map is not used for decoding the current block, the apparatus 100 may obtain syntax information related to the merge mode (e.g., merge_index) from the bitstream, obtain the motion information of the current block in merge mode, and perform inter prediction to obtain predicted samples (operation 3840).
[0435] When the first information indicates that the generated motion information map is used for decoding the current block, the apparatus 100 may determine the motion information at the position corresponding to the current block in the generated motion information map as the motion information about the current block, and may perform inter prediction to obtain predicted samples (operation 3850). Alternatively, the apparatus 100 may obtain third information (e.g., mv_index) from the bitstream, and perform inter prediction by using the motion information indicated by the third information among the motion information at the position corresponding to the current block in the generated motion information map and the motion information at the positions corresponding to at least one adjacent block of the current block, so as to obtain predicted samples (operation 3850).
[0436] When the current block is not in the merge mode, the apparatus 100 may obtain relevant syntax information from the bitstream according to the AMVP mode of the existing video coding standard, obtain motion information in the AMVP mode, and perform inter prediction to obtain predicted samples (operation 3860).
[0437] Thereafter, the apparatus 100 may obtain residual samples by performing inverse transform and inverse quantization on the current block by using the syntax information related to residual coding of the existing video coding standard, and may reconstruct the current block based on the predicted samples obtained by inter prediction and the residual samples.
[0438] When the current block is in the skip mode, the apparatus 100 may perform the operations related to operations 3830, 3840, and 3850 described above. The descriptions of operations 3830, 3840, and 3450 are included herein for reference. Thus, in Figure 38 the example of, when the current block is in the merge mode or the skip mode, the first information and the related coding method may be applied.
[0439] In Figure 38 the example of, when encoding the current picture, the apparatus 200 may perform operations corresponding to the operations of the apparatus 100 described above.
[0440] FIG. 39 shows the use of the generated motion information in the merge mode according to the present disclosure.
[0441] In the present disclosure, the merge mode refers to a method of generating a candidate list for motion information of the current block and determining, as the motion information of the current block, a candidate indicated by index information signaled through the bitstream from the generated candidate list, where the current block includes at least one of spatial neighboring blocks and temporal neighboring blocks of the current block in the current picture (for example, see Figure 16 and its related description).
[0442] Referring to FIG. 39(a), the generation of the candidate list in the merge mode of the existing video coding standard (for example, the ITU-T H.265 / HEVC standard) is shown. In the example of FIG. 39(a), A0, A1, A2, B0, and B1 represent merge candidates of spatial neighboring blocks of the current block within the current picture including the current block, and Col represents a merge candidate of a temporal neighboring block (or corresponding block or collocated block or juxtaposed block) of the current block in the picture different from the current block. Figure 40 The positions of neighboring blocks that can be used as merge candidates are shown.
[0443] In the example of FIG. 39(a), availableFlagX (X = A0, A1, A2, B0, B1, Col) is information indicating the availability of the motion information of the neighboring blocks of the current block (for example, seeFigure 16 and its related descriptions). When the motion information of an adjacent block is available (e.g., when the value of availableFlagX is 1), the adjacent block can be added or inserted into the merge candidate list. On the other hand, when the motion information of an adjacent block is not available (e.g., when the value of availableFlagX is 0), the adjacent block may not be added or inserted into the merge candidate list. As the number of merge candidates added to the merge candidate list increases, more accurate motion information can be determined for the current block, thereby improving the encoding and decoding efficiency.
[0444] In the present disclosure, the generated motion information can be added as one of the merge candidates, thereby improving the encoding and decoding performance of the merge mode. FIG. 39(b) shows the addition of the generated motion information as one of the merge candidates according to the present disclosure.
[0445] Referring to FIG. 39(b), the merge candidate of the generated motion information (or motion information map) according to the present disclosure is indicated by GeneratedMV, but the name may be changed. When there is motion information at the position corresponding to the current block in the generated motion information map, the merge candidate of the generated motion information (or motion information map) (e.g., GeneratedMV) can be determined to be available. On the other hand, when there is no motion information at the position corresponding to the current block in the generated motion information map, the merge candidate of the generated motion information (or motion information map) (e.g., GeneratedMV) can be determined to be unavailable.
[0446] When the current block is in the merge mode, the apparatus 100 can determine whether the merge candidate of the generated motion information (or motion information map) (e.g., GeneratedMV) is available, and when the merge candidate of the generated motion information (or motion information map) (e.g., GeneratedMV) is available, the merge candidate of the generated motion information (or motion information map) (e.g., GeneratedMV) can be added or inserted into the merge candidate list (operation 3910). For example, the merge candidate (e.g., GeneratedMV) can be added or inserted as a candidate at the beginning of the merge candidate list, but the order can be changed, and the merge candidate (e.g., GeneratedMV) can be added or inserted at any position.
[0447] Then, the apparatus 100 can obtain the merge index information (e.g., merge_idx or merge_index) of the current block from the bitstream, and can determine the merge candidate corresponding to the merge index information in the merge candidate list as the motion information of the current block. The apparatus 100 can perform inter prediction by using the determined motion information and obtain the predicted samples.
[0448] When encoding the current block, apparatus 200 may perform operations corresponding to those of apparatus 100 described above. For example, apparatus 200 may determine motion information of the current block, generate a merge candidate list including generated motion information (or a motion information map) of merge candidates (e.g., GeneratedMV), obtain merge index information indicating a merge candidate corresponding to the determined motion information of the current block, and encode the merge index information into a bitstream.
[0449] The operations of apparatus 100 described above with reference to FIG. 39(b) may be performed in operation 1820 of image decoding method 1800, and the operations of apparatus 200 may be performed in operation 2020 of image encoding method 2000.
[0450] FIG. 41 illustrates a method of determining a position corresponding to a current block in a generated motion information map according to the present disclosure. As described above, in the generated motion information map according to the present disclosure, motion information may be generated and stored in units of blocks of a specific size or in units of M×N pixels (e.g., 4×4, 8×8, or 16×16). When the size of the current block is the same as the size of the unit in which motion information is generated / stored in the generated motion information map, the generated motion information at that position may be used to encode / decode (e.g., decode or encode) the current block.
[0451] In the example of FIG. 41, it is assumed that the current block is a CU (or an encoding / decoding block). However, the present disclosure is not limited thereto, and the current block may be a CTU (or an encoding / decoding tree block), a TU (or a transform block), or a PU (or a prediction block). In addition, the position of the upper left sample of the current block is represented as (Cx, Cy), and the width and height of the current block are represented as W and H, respectively. Moreover, for ease of explanation, the block size or unit in which motion information is generated / stored in the generated motion information map is referred to as a generated motion information storage unit.
[0452] FIG. 41(a) illustrates a method of determining a position corresponding to a current block in a generated motion information map when the generated motion information storage unit is smaller than the size of the current block (or when the size of the current block is larger than the generated motion information storage unit). For the case shown in FIG. 41(a), apparatuses 100 and 200 may select or determine motion information corresponding to the sample position (Cx + W / 2, Cy + H / 2) in the generated motion information map as the motion information corresponding to the current block.
[0453] Referring to FIG. 41(a), as a non-limiting example, when the current block is (Cx, Cy) = (0, 0), W = 32, H = 32, and the generated motion information storage unit is 8×8, Cx + W / 2 = 16, and Cy + H / 2 = 16. Thus, apparatuses 100 and 200 can select or determine the motion information corresponding to the sample position (16, 16) (e.g., the "C" position in FIG. 41(a)) as the motion information corresponding to the current block.
[0454] FIG. 41(b) shows a method of determining a position corresponding to a current block in a generated motion information map when the size of the current block is smaller than the generated motion information storage unit (or when the size of the current block is larger than the generated motion information storage unit). For the case shown in FIG. 41(b), apparatuses 100 and 200 can select or determine the motion information corresponding to the sample position (Cx + W / 2, Cy + H / 2) in the generated motion information map as the motion information corresponding to the current block.
[0455] Referring to FIG. 41(b), as a non-limiting example, when the current block is (Cx, Cy) = (16, 16), W = 4, H = 8, and the generated motion information storage unit is 8×8, Cx + W / 2 = 18, and Cy + H / 2 = 20. Thus, apparatuses 100 and 200 can select or determine the motion information corresponding to the sample position (20, 16) (e.g., the "C" position in FIG. 41(b)) as the motion information corresponding to the current block.
[0456] FIG. 41(c) shows a method of determining a position corresponding to a current block in a generated motion information map when the current block straddles the generated motion information storage unit. Even in this case, apparatuses 100 and 200 can select or determine the motion information corresponding to the sample position (Cx + W / 2, Cy + H / 2) in the generated motion information map as the motion information corresponding to the current block.
[0457] Referring to FIG. 41(c), as a non-limiting example, when the current block is (Cx, Cy) = (12, 16), W = 8, H = 32, and the generated motion information storage unit is 8×8, Cx + W / 2 = 16, and Cy + H / 2 = 20. Thus, apparatuses 100 and 200 can select or determine the motion information corresponding to the sample position (16, 20) (e.g., the "C" position in FIG. 41(c)) as the motion information corresponding to the current block
[0458] When it is determined that there is no generated motion information at the position (e.g., position C) determined by the method described above with reference to FIG. 41, apparatuses 100 and 200 may scan multiple generated motion information within a block corresponding to the size of the current block in sequential order in the generated motion information map, and set or determine the first found motion information as the motion information of the current block.
[0459] For example, in the example of FIG. 41(a), apparatuses 100 and 200 may search for motion information in the generated motion information map in the order of C→a→b→c→d, and set or determine the first found motion information as the motion information of the current block. As another example, apparatuses 100 and 200 may start from the upper left position block (a) after position C and perform a search in raster scan order, and set or determine the first found motion information as the motion information of the current block.
[0460] When no motion information is found in the block corresponding to the size of the current block in the generated motion information map, motion information may be searched starting from the adjacent block position of the block corresponding to the size of the current block, and the first found motion information may be set or determined as the motion information of the current block. All or some neighboring blocks around the block corresponding to the size of the current block in the generated motion information map may be searched in a series of scan orders. For example, when the position corresponding to the upper left side of the current block in the generated motion information map is (Cx, Cy), the width is W, and the height is H, multiple position motion information may be searched from the generated motion information map in the following order, and the first found motion information may be set or determined as the motion information of the current block (e.g., see FIG. 41(d)).
[0461] (Cx - 1, Cy - 1) (e.g., 4102) → (Cx + w / 2, Cy - 1) (e.g., 4104) → (Cx + w, Cy - 1) (e.g., 4106) → (Cx - 1, Cy + h / 2) (e.g., 4108) → (Cx + w, Cy + h / 2) (e.g., 4110) → (Cx - 1, Cy + h) (e.g., 4112) → (Cx + w / 2, Cy + h) (e.g., 4114) → (Cx + w, Cy + h) (e.g., 4116)
[0462] When no motion information is found in the generated motion information map, apparatuses 100 and 200 may set the generated motion information of the current block to a value of 0 (e.g., a zero vector (0, 0)), or may determine that the current block is in the intra-frame mode. Moreover, during the search process, a pruning process may be additionally performed to prevent duplicate candidates.
[0463] FIG. 42 shows a method of performing sub-block prediction by using generated motion information according to the present disclosure.
[0464] FIG. 42(a) shows performing motion compensation (MC) in units of sub-blocks of a current block by using the generated motion information, FIG. 42(b) shows performing affine motion compensation by using the generated motion information as control point vectors of a four-parameter model, and FIG. 42(c) shows performing affine motion compensation by using the generated motion information as control point vectors of a six-parameter model.
[0465] Referring to FIG. 42(a), apparatuses 100 and 200 may divide a current block into sub-blocks, where the size of each sub-block is equal to or greater than the unit block size of the generated motion information (e.g., 4×4, 8×8, or 16×16), and may perform motion compensation by allocating the generated motion information in units of sub-blocks.
[0466] As a non-limiting example, in the example of FIG. 42(a), assuming that the size of the current block is 32×32 and the unit block size of the generated motion information is 8×8, the current block is divided into sub-blocks, where the size of each sub-block is the same as the unit block size of the generated motion information (e.g., 8×8) or larger than the unit block size of the generated motion information (e.g., 16×16), then motion information is obtained from the generated motion information map in units of sub-blocks, and motion compensation for the sub-blocks is performed by using the obtained (generated) motion information.
[0467] Referring to FIG. 42(b), apparatuses 100 and 200 may allocate the generated motion information at positions corresponding to a first control point cp0 and a second control point cp1 from the generated motion information map as first control point motion information and second control point motion information, and may obtain the motion information of each sub-block of the current block by using an affine four-parameter model. In the example of FIG. 42(b), apparatuses 100 and 200 may obtain a sub-block motion vector of each sub-block based on the affine four-parameter model by using Equation 1. In Equation 1, (cp0x, cp0y) represents a first control point motion vector obtained from the generated motion information map, (cp1x, cp1y) represents a second control point motion vector obtained from the generated motion information map, (mvx, mvy) represents a motion vector at a sub-block position (x, y) of the current block, and w represents the width of the current block.
[0468] [Equation 1]
[0469]
[0470] Apparatuses 100 and 200 may perform motion compensation for each sub-block by using the generated motion information and the motion vector of each sub-block obtained based on the affine four-parameter model.
[0471] Referring to FIG. 42(c), apparatuses 100 and 200 may allocate the generated motion information at positions corresponding to the first control point cp0, the second control point cp1, and the third control point cp2 from the generated motion information map as the first control point motion information, the second control point motion information, and the third control point motion information, and may obtain the motion information of each sub-block of the current block by using an affine 6-parameter model. In the example of FIG. 42(c), apparatuses 100 and 200 may obtain the sub-block motion vector of each sub-block based on the affine 6-parameter model by using Equation 2. In Equation 2, (cp0x,cp0y) represents the first control point motion vector obtained from the generated motion information map, (cp1x,cp1y) represents the second control point motion vector obtained from the generated motion information map, (cp2x,cp2y) represents the third control point motion vector obtained from the generated motion information map, (mvx,mvy) represents the motion vector at the sub-block position (x,y) of the current block, and w and h represent the width and height of the current block.
[0472] [Equation 2]
[0473]
[0474] Apparatuses 100 and 200 may perform motion compensation for each sub-block by using the generated motion information and the motion vector of each sub-block obtained based on the affine 6-parameter model.
[0475] The above embodiments of the present disclosure may be written as a computer-executable program, and the written computer-executable program may be stored in a medium.
[0476] The medium can continuously store computer-executable programs or temporarily store computer-executable programs for execution or download. In addition, the medium can be any of various recording media or storage media that combine a single piece or multiple pieces of hardware, and the medium is not limited to the medium directly connected to the computer system but can be distributed on a network. Examples of the medium include magnetic media (e.g., hard disks, floppy disks, or magnetic tapes), optical media (e.g., compact disk-read-only memory (CD-ROM) or digital versatile disk (DVD)), magneto-optical media (e.g., floppy optical disks), and ROM, random-access memory (RAM), and flash memory configured to store program instructions. The machine-readable storage medium can be provided as a non-transitory storage medium. The "non-transitory storage medium" is a tangible device and only means that it does not contain signals (e.g., electromagnetic waves). This term does not distinguish between the case of semi-permanently storing data in the storage medium and the case of temporarily storing data. For example, the non-transitory recording medium can include a buffer for temporarily storing data.
[0477] Other examples of the medium include recording media and storage media managed by an app store that distributes applications or by websites, servers, etc. that supply or distribute various other types of software.
[0478] According to an embodiment, the methods according to various disclosed embodiments can be provided by being included in a computer program product. The computer program product as a commodity can be traded between a seller and a buyer. The computer program product is distributed in the form of a device-readable storage medium (e.g., compact disk-read-only memory (CD-ROM)), or can be distributed directly and online (e.g., downloaded or uploaded) between two user devices (e.g., smartphones). In the case of online distribution, at least a part of the computer program product (e.g., a downloadable app) can be at least temporarily stored in a device-readable storage medium (such as the memory of a manufacturer server, an app store server, or a relay server), or can be temporarily generated. TM ) or between two user devices (e.g., smartphones) directly and online (e.g., downloaded or uploaded). In the case of online distribution, at least a part of the computer program product (e.g., a downloadable app) can be at least temporarily stored in a device-readable storage medium (such as the memory of a manufacturer server, an app store server, or a relay server), or can be temporarily generated.
[0479] Although one or more embodiments of the present disclosure have been described with reference to the accompanying drawings, those of ordinary skill in the art will understand that various changes can be made in form and detail without departing from the spirit and scope defined by the appended claims.
[0480] Industrial Applicability
[0481] The present disclosure can be used in various image processing devices, including image decoding devices and image encoding devices.
Claims
1. An image decoding method performed by a device, the method comprising: generating reference motion information for blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; obtaining motion information for a current block included in the current picture based on the generated reference motion information and at least one syntax information obtained from a bitstream; and reconstructing the current block based on the obtained motion information.
2. The method according to claim 1, wherein generating reference motion information for blocks included in the current picture comprises: generating reference motion information for blocks included in the current picture by inputting at least one of pixel data or motion information of at least one reconstructed picture into an artificial neural network.
3. The method according to claim 1, wherein the motion information includes at least one of first motion vector information, second motion vector information, first prediction list utilization information, second prediction list utilization information, or prediction mode information.
4. The method according to claim 1, further comprising aggregating at least one of the formats of the generated reference motion information, scaled reference motion information, or transformed reference motion information.
5. The method according to claim 1, wherein generating reference motion information for blocks included in the current picture comprises: downsampling at least one reconstructed picture; and generating reference motion information for blocks included in the current picture based on at least one of pixel data or motion information of the downsampled reconstructed picture.
6. The method according to claim 1, wherein obtaining motion information for the current block comprises: when a time interval between a reconstructed picture and the current picture is equal to or greater than a specific value, excluding the generated reference motion information from candidates for motion information of the current block.
7. The method according to claim 1, wherein generating reference motion information for blocks included in the current picture comprises: when applying reference picture resampling, downsampling the reconstructed picture based on a low-resolution reference picture.
8. The method according to claim 1, further comprising generating motion information for additional pictures by performing interpolation or extrapolation based on motion information of at least one reconstructed picture.
9. The method according to claim 1, wherein at least one syntax information includes information indicating whether the generated reference motion information is used to obtain motion information for the current block.
10. The method according to claim 1, wherein obtaining motion information for the current block comprises: constructing a merge candidate list including the generated reference motion information; and obtaining motion information for the current block based on the constructed merge candidate list.
11. The method according to claim 1, wherein the generated reference motion information includes motion information at a position corresponding to an x coordinate and a y coordinate, the x coordinate being obtained by adding half the width of the current block to the top-left sample position of the current block in the generated reference motion information map, and the y coordinate being obtained by adding half the height of the current block to the top-left sample position of the current block.
12. The method according to claim 1, wherein Obtaining motion information of a current block includes dividing the current block into a plurality of sub-blocks and assigning the generated reference motion information to the plurality of sub-blocks, and Reconstructing the current block includes performing motion compensation on the sub-blocks based on the generated reference motion information of the sub-blocks.
13. The method according to claim 1, wherein Obtaining motion information of a current block includes obtaining a plurality of control point motion vectors of the current block based on the generated reference motion information, and Reconstructing the current block includes performing motion compensation on the sub-blocks of the current block by using the obtained control point motion vectors.
14. An image encoding method performed by an apparatus, comprising: Generating reference motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; Determining motion information of a current block included in the current picture; and Encoding the motion information of the current block into a bitstream based on the generated reference motion information.
15. A non-transitory computer-readable storage medium storing a bitstream encoded by an image encoding method, The image encoding method comprising: Generating reference motion information of blocks included in a current picture based on at least one of pixel data or motion information of at least one reconstructed picture; Determining motion information of a current block included in the current picture; and Encoding the motion information of the current block into a bitstream based on the generated reference motion information.