Image encoding and decoding method and image decoding device
By selecting prediction candidates and motion vector differences in image encoding and decoding, the motion information export process is optimized, and the problem of high computational complexity in traditional methods is solved, and the performance and efficiency of image encoding and decoding are improved.
Patent Information
- Application Number
- CN202210683144.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-06-05
- Filing Date
- 2016-06-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2036-06-07
AI Technical Summary
Traditional image encoding and decoding methods have high computational complexity in inter- and intra-prediction, resulting in performance degradation.
By selecting prediction candidates from the reference screen of the current screen, using motion vector difference and accuracy information to derive the motion information of the current block, and combining intra prediction and inter prediction, the motion vector selection process is optimized.
Improves the performance and efficiency of image encoding and decoding, simplifies the calculation process, and reduces complexity.
Smart Images

Figure CN115086653B_ABST
Abstract
Description
[0001] This patent application is a divisional application of the following invention patent applications:
[0002] Application number: 201680045584.1
[0003] Application date: June 7, 2016
[0004] Invention Title: Image encoding and decoding method and image decoding device Technical Field
[0005] The present invention relates to an image encoding and decoding method. More specifically, the present invention relates to an image encoding and decoding method and an image decoding device, whereby the method and the device select a prediction candidate from a reference block of a reference picture including a current picture, and use the selected prediction candidate to derive motion information for the current block when encoding and decoding an image. Background Art
[0006] With the popularization of the Internet and mobile terminals and the development of information and communication technology, the use of multimedia data is rapidly increasing. Therefore, it is necessary to improve the performance and efficiency of image processing systems by using image prediction within various systems to provide various services or perform operations on them.
[0007] Meanwhile, in conventional image encoding and decoding methods, the motion vector of the current block is estimated by predicting the motion information of a neighboring block of the current block in at least one reference picture after or before the current picture using an inter-frame prediction method, or by obtaining the motion information in a reference block within the current picture using an intra-frame prediction method.
[0008] However, in the conventional inter prediction method, a prediction block is generated by using a temporal prediction mode between pictures, so its calculation becomes complicated, and intra prediction also becomes complicated.
[0009] Therefore, there is a need to improve image encoding or image decoding performance in conventional image encoding and decoding methods. Summary of the invention
[0010] Technical issues
[0011] Therefore, the present invention is proposed in view of the above-mentioned problems arising in the prior art, and an object of the present invention is to provide an image encoding and decoding method and an image decoding device, wherein the method and the device select a prediction candidate from prediction candidates of a reference picture including a current picture, and use a motion vector candidate selection method to derive motion information of a current block during image encoding and decoding.
[0012] Another object of the present invention is to provide an image encoding and decoding method and an image decoding device, wherein the method and the device use motion vector differences when selecting prediction candidates from reference blocks of a reference picture including a current picture to derive motion information of a current block during image encoding and decoding.
[0013] Another object of the present invention is to provide an image encoding and decoding method and an image decoding device, wherein the method and the device use motion vector accuracy when selecting prediction candidates from reference blocks of a reference picture including a current picture to derive motion information of a current block during image encoding and decoding.
[0014] The present invention provides an image decoding method, comprising: decoding reference picture information indicating whether a current picture including a current block can be used as a reference picture and precision information indicating the precision of a motion vector, wherein the reference picture information is decoded from a picture parameter set; generating a reference picture list, wherein the reference picture list includes pictures decoded before the current picture, and when the reference picture information indicates that the current picture can be used as a reference picture, the reference picture list also includes the current picture; determining a reference picture of the current block from the reference picture list; determining the precision of a motion vector of the current block based on the precision information and the determined reference picture; obtaining a motion vector of the current block based on the motion vector precision; and deriving a prediction sample of the current block based on the motion vector, wherein the precision of the motion vector of the current block is determined based on whether the determined reference picture is the current picture.
[0015] The present invention provides an image encoding method, comprising: determining whether a current picture including a current block can be used as a reference picture and the motion vector accuracy of the current block; generating a reference picture list, wherein the reference picture list includes reference pictures encoded before the current picture, and when the reference picture information indicates that the current picture can be used as a reference picture, the reference picture list also includes the current picture; determining the reference picture of the current block from the reference picture list; obtaining the motion vector of the current block based on the motion vector accuracy; and encoding the reference picture information and accuracy information of the current block based on whether the current picture can be used as a reference picture and the motion vector accuracy of the current block, wherein the reference picture information is encoded in a picture parameter set, wherein the motion vector accuracy of the current block is determined based on whether the determined reference picture is the current picture.
[0016] The present invention also provides a bitstream storage method, comprising: generating a bitstream, the bitstream including reference picture information indicating whether a current picture including a current block can be used as a reference picture and precision information indicating the precision of a motion vector; and storing the bitstream, wherein the reference picture of the current block is determined from a reference picture list; the precision of the motion vector of the current block is determined based on the precision information and the determined reference picture; the motion vector of the current block is obtained based on the motion vector precision; and the prediction sample of the current block is derived based on the motion vector, the reference picture list includes reference pictures decoded before the current picture, when the reference picture information indicates that the current picture can be used as a reference picture, the reference picture list also includes the current picture, the reference picture information is decoded from a picture parameter set, and the precision of the motion vector of the current block is determined based on whether the determined reference picture is the current picture.
[0017] Technical Solutions
[0018] In order to achieve the above-mentioned purpose, in one aspect of the present invention, there is provided an image encoding method, which is an image encoding method for configuring reference pixels when performing intra-frame prediction, the method comprising: obtaining reference pixels of a current block from neighboring blocks when performing intra-frame prediction of the current block; adaptively performing filtering on the reference pixels; generating a prediction block of the current block by using the reference pixels to which the filtering is adaptively applied as input values according to a prediction mode of the current block; and applying an adaptive post-processing filter to the prediction block.
[0019] Here, when obtaining the reference pixel, the reference pixel of the current block may be obtained from a neighboring block.
[0020] Here, obtaining the reference pixel may be determined according to whether a neighboring block is available.
[0021] Here, whether the neighbor block is available may be determined by the position of the neighbor block or a specific flag (constrained_intra_pred_flag) or both. In one embodiment, when the neighbor block is available, the specific flag may have a value of 1. This may indicate that when the prediction mode of the neighbor block is an inter mode, the reference pixels of the corresponding block may be used to predict the current block.
[0022] Here, a specific flag (constrained_intra_pred_flag) may be determined according to a prediction mode of a neighbor block, and the prediction mode may be one of intra prediction and inter prediction.
[0023] Here, when the specific flag (constrained_intra_pred_flag) is 0, whether the neighbor block is available becomes "true" regardless of the prediction mode of the neighbor block, and when the specific flag is 1, when the prediction mode of the neighbor block is intra-frame prediction, whether the neighbor block is available becomes "true", and when the prediction mode of the neighbor block is inter-frame prediction, whether the neighbor block is available becomes "false".
[0024] Here, the inter prediction may generate a prediction block by referring to at least one reference picture.
[0025] Here, the reference picture may be managed by using a reference picture list 0 (List0) and a reference picture list 1 (List1), and at least one of a previous picture, a subsequent picture, and a current picture may be included in List0 and List1.
[0026] Here, for list 0 and list 1, it may be adaptively determined whether to include the current picture in the reference picture list.
[0027] Here, information determining whether to include the current picture in the reference picture list may be included in a sequence, a reference picture parameter set, or the like.
[0028] To achieve the above-mentioned purpose, in another aspect of the present invention, an image decoding method is provided, wherein the method is an image decoding method executed in a computing device, and the method comprises: obtaining a flag indicating whether the reference pixels of a neighboring block are available in a sequence or a picture unit from an input bit stream; when performing intra-frame prediction according to the flag, determining whether the reference pixels of the neighboring block are available; when the flag is 0, regardless of the prediction mode of the neighboring block, using the reference pixels of the neighboring block to predict the current block; when the flag is 1, when the prediction mode of the neighboring block is intra-frame prediction, using the reference pixels of the neighboring block to predict the current block; and when the prediction mode of the neighboring block is inter-frame prediction, not using the reference pixels of the neighboring block to predict the current block.
[0029] Here, the inter prediction may generate a prediction block based on performing block matching in a reference picture.
[0030] Here, reference pictures may be managed by using List 0 in a P picture and using List 0 and List 1 in a B picture.
[0031] Here, in inter prediction, list 0 may include the current picture.
[0032] Here, in inter prediction, list 1 may include the current picture.
[0033] Here, whether to include the current picture in List 0 and List 1 may be determined based on a flag transmitted from a sequence parameter.
[0034] Here, whether to include the current picture in List 0 and List 1 may be determined based on a flag transmitted from the picture parameter.
[0035] In order to achieve the above-mentioned purpose, in another aspect of the present invention, a motion vector candidate selection method is provided, the method comprising: configuring a spatial motion vector candidate (a first candidate); determining whether there is a reference picture of the current block in the current picture; and when there is a reference picture of the current block in the current picture, adding a spatial motion vector candidate (a second candidate) of another block of the current picture that is encoded before the current block.
[0036] Here, the motion vector candidate selection method may further include adding a temporal motion vector candidate (third candidate) when a reference picture of the current block does not exist in the current picture.
[0037] Here, the motion vector candidate selection method may further include configuring a combined list candidate including the first candidate, the second candidate, and the third candidate after adding the spatial motion vector candidate and the temporal motion vector candidate.
[0038] Here, the motion vector candidate selection method may further include: after configuring the combined list candidates, determining whether the current picture is a reference picture; and when the current picture is a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, adding fixed candidates with preset fixed coordinates.
[0039] Here, the motion vector candidate selection method may further include adding a fixed candidate having a (0, 0) coordinate when the current picture is not a reference picture and the number of motion vector candidates within the combined list candidates is less than a preset number.
[0040] Here, the other blocks of the current picture may be blocks facing the current block, having a neighbor block of the current block between the current block and the current block, and may include blocks encoded before the current block in the current picture. Another block of the current picture may be a block encoded by performing inter-frame prediction before the current block.
[0041] In order to achieve the above-mentioned purpose, in another aspect of the present invention, an image encoding method is provided, the method comprising: configuring spatial motion vector candidates (first candidates); determining whether there is a reference picture of the current block in the current picture; when there is a reference picture of the current block in the current picture, adding a spatial motion vector candidate (second candidate) of another block of the current picture encoded before the current block; when there is no reference picture of the current block in the current picture, adding a temporal motion vector candidate (third candidate); and performing reference pixel filtering based on a motion vector candidate including any one of the first candidate, the second candidate and the third candidate.
[0042] Here, the image encoding method may further include: before performing reference pixel filtering, determining whether the current picture is a reference picture; when the current picture is a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, adding a fixed candidate with preset fixed coordinates; and when the current picture is not a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, adding a fixed candidate with (0, 0) coordinates.
[0043] Here, the image encoding method may further include: generating a prediction block by performing intra prediction after performing reference pixel filtering; and encoding a prediction mode of the generated prediction block.
[0044] In order to achieve the above-mentioned purpose, in another aspect of the present invention, an image encoding method is provided, the method comprising: configuring spatial motion vector candidates (first candidates); determining whether there is a reference picture of the current block in the current picture; when there is a reference picture of the current block in the current picture, adding a spatial motion vector candidate (second candidate) of another block of the current picture encoded before the current block; when there is no reference picture of the current block in the current picture, adding a temporal motion vector candidate (third candidate); and performing motion estimation based on motion vector candidates including any one of the second candidate, the third candidate and the first candidate.
[0045] Here, the image encoding method may further include: before performing motion estimation; determining whether the current picture is a reference picture; when the current picture is a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, adding a fixed candidate with preset fixed coordinates; and when the current picture is not a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, adding a fixed candidate with (0, 0) coordinates.
[0046] Here, the image encoding method may further include, after performing the motion estimation, performing the interpolation.
[0047] In order to achieve the above-mentioned purpose, in another aspect of the present invention, there is provided an image decoding method, the method comprising: entropy decoding a coded picture; performing dequantization on a decoded picture; performing inverse transformation on the dequantized picture; selecting motion information prediction candidates for the inverse transformed image based on header information of the decoded picture; and decoding the inverse transformed image based on image information obtained based on the motion information prediction candidates.
[0048] Here, in selecting motion information prediction candidates, motion prediction of the current block can be performed based on a candidate group including spatial motion vector candidates in neighboring blocks of the current picture within an inverse transformed image (first candidates) and spatial motion vector candidates from other blocks of the current picture encoded before the current block (second candidates).
[0049] Here, in selecting the motion information prediction candidate, motion prediction of the current block may be performed based on the candidate group to which the temporal motion vector candidate (third candidate) is further added.
[0050] Here, in selecting motion information prediction candidates, a combined list candidate group including a first candidate, a second candidate, and a third candidate can be configured, and it can be determined whether the current picture of the current block is a reference picture, and when the current picture is a reference picture and the number of motion vector candidates in the combined list candidates is less than a preset number, a fixed candidate with preset fixed coordinates can be added.
[0051] Here, in the selection of motion information prediction candidates in the image decoding method, when the current picture is not a reference picture and the number of motion vector candidates within the combined list candidates is less than a preset number, a fixed candidate with (0, 0) coordinates is added.
[0052] Here, the other blocks of the current picture may be blocks facing the current block, having a neighbor block of the current block therebetween, and blocks included in the current picture that are encoded by performing inter-frame prediction before the current block.
[0053] In order to achieve the above-mentioned purpose, in another aspect of the present invention, an image encoding method is provided, wherein the method is an image encoding method for generating a predicted image from an original image by predicting motion information, the method comprising: configuring a motion information prediction candidate group; changing the motion vector of the candidate block belonging to the candidate group according to the precision unit of the motion vector of the current block; and calculating the difference by subtracting the motion vector of the candidate block from the motion vector of the current block according to the precision unit.
[0054] In order to achieve the above-mentioned purpose, in another aspect of the present invention, an image decoding method is provided, wherein the method is an image decoding method that generates a reconstructed image by entropy decoding a coded image, dequantizing the entropy decoded image, and performing an inverse transform on the dequantized image, the method comprising: configuring a motion information prediction candidate group of the reconstructed image based on header information of the entropy decoded image; changing the motion vector of the candidate block belonging to the candidate group according to the precision unit of the motion vector of the current block; and calculating a differential value by subtracting the motion vector of the candidate block from the motion vector of the current block according to the precision unit.
[0055] Here, when changing the motion vector of the candidate block, the motion vector may be scaled according to a first distance between the current picture where the current block is placed and the reference picture, and a second distance between the picture of the candidate block and the reference picture corresponding to the candidate block.
[0056] Here, the image encoding method may further include: after calculating the differential value, determining an interpolation accuracy of each reference picture based on an average distance of the first distance and the second distance.
[0057] Here, when the reference picture of the current block is the same as the reference picture of the candidate block, the change of the motion vector of the candidate block may be omitted.
[0058] Here, when changing the motion vector of the candidate block, the motion vector of the neighbor block or adjacent block may be changed to a motion vector precision unit of the current block according to the motion vector precision of the current block.
[0059] Here, the neighbor block may be included with the current block and placed at another block. In addition, the neighbor block may be a block whose motion vector is searched by performing inter-frame prediction before the current block.
[0060] In order to achieve the above-mentioned purpose, in another aspect of the present invention, there is provided a memory image decoding device including a stored program or program code, wherein the program or program code is used to generate a reconstructed image by entropy decoding a coded image, dequantize the entropy decoded image, and perform an inverse transform on the dequantized image; and a processor, which is connected to the memory and executes the program, and the processor executes the program to: configure a motion information prediction candidate group of the reconstructed image based on header information of the entropy decoded image; change the motion vector of the candidate block belonging to the candidate group according to the precision unit of the motion vector of the current block; and calculate the difference by subtracting the motion vector of the candidate block from the motion vector of the current block according to the precision unit.
[0061] Here, when the processor changes the motion vector of the candidate block, the processor may scale the motion vector according to a first distance between the current picture where the current block is placed and the reference picture, and a second distance between the picture of the candidate block and the reference picture corresponding to the candidate block.
[0062] Here, when the processor calculates the differential value, the processor may determine the interpolation accuracy of each reference picture based on an average distance of the first distance and the second distance.
[0063] Here, when the reference picture of the current block is the same as the reference picture of the candidate block, the processor may omit the change of the motion vector of the candidate block.
[0064] Here, when the processor changes the motion vector of the candidate block, the processor may change the motion vector of the neighbor block or adjacent block to a motion vector precision unit of the current block according to the motion vector precision of the current block.
[0065] Here, the neighbor block may be included with the current block and placed at another block, and include a block for searching a motion vector by performing inter prediction before the current block.
[0066] To achieve the above-mentioned purpose, in another aspect of the present invention, an image encoding method is provided, the method comprising: when the interpolation accuracy of a reference picture is a first value, searching for the accuracy of a motion vector of a first neighbor block of a current block that refers to the reference picture, the accuracy being equal to the first value or a second value greater than the first value; searching for the accuracy of a motion vector of a second neighbor block of the current block, the accuracy being a third value greater than the second value; and encoding first information of the motion vectors for the first block and the second block and matching information for the accuracy of the motion vectors.
[0067] To achieve the above-mentioned purpose, in another aspect of the present invention, an image decoding method is provided, the method comprising: when the interpolation accuracy of a reference picture has a first value, searching for the accuracy of a motion vector of a first neighbor block of a current block that refers to the reference picture, the accuracy being equal to the first value or a second value greater than the first value; searching for the accuracy of a motion vector of a second neighbor block of the current block, the accuracy being a third value greater than the second value; and image decoding based on matching information of the accuracy of the motion vector and information of the motion vectors for the first block and the second block.
[0068] Here, information of the current picture may be added to the end of reference picture list 0 and reference picture list 1.
[0069] Here, the first value may be a suitable fraction. In addition, the second value or the third value may be an integer.
[0070] Here, when the second value has an occurrence frequency greater than that of the third value within the index including the matching information, the second value may have a shorter number of binary bits than that of the third value.
[0071] In this case, the third value may have the value zero, ie, the precision is 0, when the third value has the highest occurrence frequency within the index.
[0072] Here, the first neighbor block or the second neighbor block may be included in a block spatially different from the current block and placed on a block spatially different from the current block. The first neighbor block or the second neighbor block may be a block encoded by performing inter-frame prediction before the current block.
[0073] Here, the reference picture of the first neighbor block or the second neighbor block may be the current picture. The motion vector of the first neighbor block or the second neighbor block may be searched by performing inter-frame prediction.
[0074] To achieve the above-mentioned purpose, in another aspect of the present invention, an image decoding device is provided, comprising: a memory storing a program or program code for image decoding; and a processor connected to the memory, wherein the processor, through the program: when the interpolation accuracy of the reference picture is a first value, searches for the accuracy of the motion vector of the first neighbor block of the current block referring to the reference picture, the accuracy being equal to the first value or a second value greater than the first value; searches for the accuracy of the motion vector of the second neighbor block of the current block, the accuracy being a third value greater than the second value; and decodes the image based on matching information of the accuracy of the motion vector and information of the motion vectors for the first block and the second block.
[0075] Beneficial effects
[0076] According to the image encoding and decoding method and the image decoding device according to the embodiments of the present invention as described above, motion vector candidates for image encoding and decoding can be effectively selected in various systems configured with an image processing system or including such an image processing system, thereby improving the performance and efficiency of the device or system.
[0077] In addition, since motion vector candidates or motion information prediction candidates are effectively selected, the performance and efficiency of an image encoding device, an image decoding device, or an image processing system can be improved.
[0078] Specifically, the motion vector can be selected by using the motion vector difference, and scaling or precision adjustment can be applied in various forms according to the precision of the block or picture. In addition, the performance and efficiency of encoding and decoding can be improved by selecting the best candidate from the applicable candidate group, calculating the difference with the motion vector of the current block, and encoding the calculated difference.
[0079] In addition, in particular, by selecting a motion vector using the motion vector accuracy, a reference block can be copied and used as a prediction block within the current picture. Therefore, the performance and efficiency of image encoding and decoding can be improved.
[0080] In addition, encoding and decoding performance and efficiency may be improved by extending the precision of a motion vector using intra block copying or block matching and by including the current picture in reference picture list 0 (List 0) and reference picture list 1 (List 1) of the motion vector. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 It is a view for illustrating a system using the image encoding device or the image decoding device or both of the present invention.
[0082] Figure 2 is a block diagram of an image encoding apparatus according to an embodiment of the present invention.
[0083] Figure 3 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.
[0084] Figure 4 is an example view illustrating inter prediction of a P slice in an image encoding and decoding method according to an embodiment of the present invention.
[0085] Figure 5 is an example view illustrating inter prediction of a B slice in an image encoding and decoding method according to an embodiment of the present invention.
[0086] Figure 6 is an exemplary view for illustrating generation of a prediction block in one direction in an image encoding and decoding method according to an embodiment of the present invention.
[0087] Figure 7 is an example view of configuring a reference picture list in an image encoding and decoding method according to an embodiment of the present invention.
[0088] Figure 8 is another example view of performing inter prediction from a reference picture list in an image encoding and decoding method according to an embodiment of the present invention.
[0089] Fig. 9 is an exemplary view for illustrating intra prediction in an image encoding method according to an embodiment of the present invention.
[0090] Fig.10 is an exemplary view for illustrating a prediction principle in a P slice or a B slice in the image encoding method according to an embodiment of the present invention.
[0091] Fig.11 is used to show Fig.10 Example views of interpolation performed in an image encoding method.
[0092] Fig.12 It is a view for illustrating a main process of an image encoding method according to an embodiment of the present invention in syntax of a coding unit.
[0093] Fig.13 is used to show that when Fig.12 An example diagram of an example of supporting symmetric partitioning or asymmetric partitioning like inter-frame prediction when performing block matching in the current picture used to generate a prediction block.
[0094] Fig.14 is used to show that inter prediction supports 2N×2N and N×N as Fig. 9Example diagram of intra-frame prediction.
[0095] Fig.15 1 is a view for illustrating a process of applying a one-dimensional horizontal filter to pixels existing at positions a, b, and c (assumed to be x) of an image in the image decoding method according to an embodiment of the present invention.
[0096] Fig.16 is an example diagram of a current block and neighbor blocks according to a comparative example.
[0097] Fig.17 is an example diagram of a current block and neighbor blocks according to another comparative example.
[0098] Fig.18 is an example diagram of a current block and neighbor blocks according to yet another comparative example.
[0099] Fig.19 is an example diagram of a current block and neighbor blocks that can be selected in an image encoding method according to an embodiment of the present invention.
[0100] Fig. 20 This is an example diagram for illustrating a situation in which, in an image encoding method according to an embodiment of the present invention, when a temporal distance between a reference picture of a current block and a reference picture of a candidate block is equal to or greater than a predetermined distance, the current block is excluded from a candidate group, and when the temporal distance is less than a predetermined distance, the candidate block is included in the candidate group after scaling is performed according to the distance.
[0101] Fig.21 is an exemplary diagram for illustrating a case where a current picture is added to a prediction candidate group when a picture reference through a current block is a picture different from the current picture in an image encoding method according to an embodiment of the present invention.
[0102] Fig. 22 is an exemplary diagram for illustrating a case where a current picture is added to a prediction candidate group when a picture referenced by a current block is a current picture in an image encoding method according to an embodiment of the present invention.
[0103] Fig.23 is a flowchart of an image encoding method according to another embodiment of the present invention.
[0104] Fig.24 It is a diagram for illustrating when the motion vector accuracy changes in block units.
[0105] Fig.25 is a view for illustrating a case where the accuracy of a motion vector of a block is determined according to the interpolation accuracy of a reference picture in the image encoding method according to an embodiment of the present invention.
[0106] Fig.26is a flowchart of an image encoding and decoding method using a motion vector difference according to an embodiment of the present invention.
[0107] Figure 27 to Figure 32 is a view for illustrating a process in an image encoding method according to an embodiment of the present invention of calculating a motion vector difference in various cases when interpolation accuracy is determined in block units.
[0108] Figure 33 to Figure 36 is a view for illustrating a process of expressing the accuracy of a motion vector difference in an image encoding method according to an embodiment of the present invention.
[0109] Fig.37 An example of a reference structure of a random access mode in an image encoding and decoding method according to an embodiment of the present invention is shown.
[0110] Fig.38 is a view for illustrating that a single picture can have at least two interpolation precisons in the image encoding method according to an embodiment of the present invention.
[0111] Fig.39 It shows that the current screen is Fig.38 A view of the reference picture list for an I picture in .
[0112] Fig.40 It shows that the current screen is Fig.38 A diagram of a reference picture list for a P picture in FIG.
[0113] Fig.41 is a view showing a reference picture list when a current picture is B(2) in an image encoding and decoding method according to an embodiment of the present invention.
[0114] Fig.42 is a view showing a reference picture list when a current picture is B(5) in an image encoding and decoding method according to an embodiment of the present invention.
[0115] Fig.43 is a view for illustrating a process of determining the precision of a motion vector of each block according to the interpolation precision of a reference picture in an image encoding and decoding method according to an embodiment of the present invention.
[0116] Fig.44 is a view for illustrating a process of adaptively determining the precision of a motion vector of each block when the interpolation precision of each reference picture is constant in the image encoding / decoding method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0117] The preferred embodiments of the present invention will be explained below with reference to the accompanying drawings. Although the present invention may have various modifications and configurations, certain embodiments have been shown and explained herein. However, this should not be construed as limiting the present invention to any specific disclosed configuration, but should be understood to include all modifications, equivalents or substitutions that may be included under the concept and technical scope of the present invention.
[0118] It should be understood that, although the terms "first", "second", etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the present invention. As used herein, the term "and / or" includes any and all combinations of one or more of the related listed items.
[0119] It should be understood that when an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. Conversely, when an element is referred to as being "directly connected" or "directly coupled" to another element, there are no intermediate elements.
[0120] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the present invention. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that when used herein, the terms "comprise", "include" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0121] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention belongs. It should be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and will not be interpreted in an ideal or excessive formal sense unless explicitly defined as such herein.
[0122] Typically, a video may be configured with a series of pictures, and each picture may be divided into predetermined regions such as frames or blocks. In addition, various sizes or terms such as coding tree unit (CTU), coding unit (CU), prediction unit (PU), and transform unit (TU) may be used to refer to the divided regions. Each unit may be configured with a single luminance block and two chrominance blocks, and the unit may be configured differently according to the color format. In addition, the size of the luminance block and the chrominance block may be determined according to the color format. For example, in the case of 4:2:0, the chrominance block may have a horizontal and vertical length of 1 / 2 of the horizontal and vertical length of the luminance block. For the above units and terms, reference may be made to the terms of conventional HEVC (high efficiency video coding) or H.264 / AVC (advanced video coding).
[0123] In addition, when encoding or decoding a current block or a current pixel, a picture, a block or a pixel reference is referred to as a reference picture, a reference block or a reference pixel. In addition, those skilled in the art will appreciate that the term "picture" described below may be replaced by other terms having equivalent meanings such as image, frame, etc.
[0124] Hereinafter, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. In order to facilitate a thorough understanding of the present invention, the same reference numerals in the accompanying drawings represent the same components, and repeated description of the same components will be omitted.
[0125] Figure 1 It is a view for illustrating a system using the image encoding device or the image decoding device or both of the present invention.
[0126] refer to Figure 1 The system using the image encoding device or the image decoding device or both may be a user terminal 11 such as a personal computer (PC), a notebook computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a portable game console (PlayStation Portable, PSP), a wireless communication terminal, a smart phone, a television, etc., or may be a server terminal 12 such as an application server and a service server. The above system may be referred to as a computing device.
[0127] In addition, the computing device may include various devices, including a communication device such as a communication modem for communicating between various devices and wired and wireless communication networks, a memory 18 storing various programs and data for encoding or decoding an image or for performing intra-frame or inter-frame prediction thereof, and a processor 14 for computing and controlling by executing a program.
[0128] In addition, in the computing device, the image may be encoded into a bit stream by the image encoding device, and the bit stream may be transmitted to the image decoding device in real time or non-real time by using a wired or wired communication network such as the Internet, a near-field wireless communication network, a wireless LAN network, a WiBro network, a mobile communication network, etc., or by using various communication interfaces such as a cable, a universal serial bus (USB), and the transmitted stream may be decoded and played in the image decoding device as a reconstructed image. In addition, the image encoded into a bit stream by the image encoding device may be transmitted from the encoding device to the decoding device by using a computer-readable recording medium.
[0129] Figure 2 is a block diagram of an image encoding apparatus according to an embodiment of the present invention. Figure 3 is a block diagram of an image decoding apparatus according to an embodiment of the present invention.
[0130] like Figure 2 As shown, the image encoding apparatus 20 according to an embodiment may include a prediction unit 200, a subtractor 205, a transform unit 210, a quantization unit 215, a dequantization unit 220, an inverse transform unit 225, an adder 230, a filtering unit 235, a decoded picture buffer 240, and an entropy encoding unit 245. In addition, the image encoding apparatus 20 may further include a partition unit 190.
[0131] In addition, if Figure 3 As shown, the image decoding device 30 according to an embodiment may include an entropy decoding unit 305, a prediction unit 310, a dequantization unit 315, an inverse transform unit 320, an adder 325, a filtering unit 330 and a decoded picture buffer 335.
[0132] The above-mentioned image encoding device 20 and image decoding device 30 may be independent devices. However, according to an embodiment, a single image encoding and decoding device may be formed. Here, the prediction unit 200, the dequantization unit 220, the inverse transform unit 225, the adder 230, the filtering unit 235, and the decoded picture buffer 240 of the image encoding device 20 may be respectively and substantially the same as the prediction unit 310, the dequantization unit 315, the inverse transform unit 320, the adder 325, the filtering unit 330, and the memory 335 of the image decoding device 30, so that their structures are the same or may be implemented to perform the same functions. In addition, the entropy encoding unit 245 may correspond to the entropy decoding unit 305 by performing its functions in reverse. Therefore, in the following detailed description of the technical elements and their operating principles, the overlapping description of the corresponding technical elements will be omitted.
[0133] In addition, the image decoding device corresponds to a computing device in which the image encoding method performed by the image encoding device is applied to the image decoding method, so in the following description, the image encoding device will be mainly described.
[0134] The computing device may include a memory storing a program or software module for implementing the image encoding method or the image decoding method or both, and a processor associated with the memory to execute the program. In addition, the image encoding device may be referred to as an encoder, and the image decoding device may be referred to as a decoder.
[0135] Each configuration element of the image encoding device of the present embodiment will be described in detail.
[0136] The partition unit 190 may partition the input image into blocks having a predetermined size (M×N). Here, M or N is any natural number equal to or greater than 1.
[0137] In detail, the partition unit 190 may be configured with a picture partition unit and a block partition unit. The size of the block form may be determined according to the characteristics and resolution of the image. The size or form of the block supported by the picture partition unit may be an M×N square form (256×256, 128×128, 64×64, 32×32, 16×16, 8×8, 4×4, etc.) represented by an exponential power of 2 in horizontal and vertical lengths or may be an M×N rectangular form. For example, in the case of an 8k UHD image with high resolution, the input image may be divided into a 256×256 block size, in the case of an HD image, the input image may be divided into a 128×128 block size, and in the case of a WVGA image, the input image may be divided into a 16×16 block size.
[0138] The information of the block size or form may be set in a sequence unit, a picture unit, a slice unit, etc. In addition, the information therefor may be transmitted to a decoder. In other words, the information may be set in a sequence parameter set, a picture parameter set, a slice header, or a combination thereof.
[0139] Here, a sequence may be a unit configured with several scenes. In addition, a picture is a term indicating a luminance component (Y) or both luminance and chrominance components (Y, Cb, and Cr) in a single scene or picture, and a single frame or a single domain may be the range of a single picture according to circumstances.
[0140] A slice may refer to a single independent slice segment or multiple dependent slice segments present in the same access unit. An access unit represents a set of network abstraction layer (NAL) units, and the network abstraction layer is associated with a single coded picture. A NAL unit is a video compression bitstream configured with a network-friendly syntax structure in the H.264 / AVC and HEVC standards. Typically, a single slice unit is configured with a single NAL unit. In the system standard, a NAL unit or a NAL set that makes up a single frame is considered a single access unit.
[0141] Describing the picture partition unit again, the information of the block size or form (M×N) can be configured with an explicit flag. In detail, the explicit flag may include block form information, length information when the block is square, or each length or the difference between horizontal and vertical lengths when the block is rectangular.
[0142] For example, when M and N are configured as exponential powers of k (assuming k is 2) (M = 2 m , N = 2 n ), various methods such as unary binarization method, truncated unary binarization method, etc. can be used to encode the information of m and n and transmit it to the decoder.
[0143] In addition, when the minimum available partition size (Minblksize) supported by the picture partition unit is I×J (for ease of explanation, assume that I=J. When I is 2 i , and J is 2 j When M and N are different, the difference between m and n (|mn|) can be transmitted. In addition, when the maximum available partition size (Maxblksize) supported by the picture partition unit is I×J (for ease of explanation, assume that I=J. When I is 2 i , and J is 2 j When , you can send im or nj information.
[0144] In the implicit case, for example, when the syntax of the relevant information exists but is not checked in the encoder / decoder, the encoder or decoder can follow the basic preset setting. For example, when the relevant syntax is not checked while checking the block form information, the block form can be set to the square form as the basic set.
[0145] In addition, when checking the block size information, in more detail, as in the above example, by checking the block size information using the difference value with the minimum available partition size (Minblksize), the syntax associated with the difference value may be checked. However, when it is impossible to check the syntax associated with the difference value, it may be obtained from a preset basic setting value associated with the minimum available partition size (Minblksize).
[0146] Therefore, the block size or form of the picture partition unit may be explicitly transmitted from the encoder or the decoder or both by using relevant information, or may be implicitly determined according to the characteristics and resolution of the image.
[0147] As above, the block divided and determined by the picture partition unit can be used as a basic coding unit. In addition, the block divided and determined by the picture partition unit can be the smallest unit constituting a higher-level unit such as a picture, a slice, a tile, etc., or can be a maximum unit such as a coding block, a prediction block, a transform block, a quantization block, an entropy block, an in-loop filter block, etc. However, some blocks are not limited to this, and exceptions are possible. For example, an in-loop filter block can be applied to a unit larger than the above block size.
[0148] The block partition unit performs partitioning of coding blocks, prediction blocks, transform blocks, quantization blocks, entropy blocks, in-loop filtering blocks, etc. The partition unit 190 can perform its function by being included in each configuration. For example, the transform unit 210 may include a transform block partition unit, and the quantization unit 215 may include a quantization block partition unit. The initial block size or form of the block partition unit may be determined by a previous step or by a partition result of a higher-level block.
[0149] For example, in the case of a coding block, a block obtained by a picture partitioning unit as a previous process may be set as an initial block. Alternatively, in the case of a prediction block, a block obtained by dividing a coding block of a higher level than the prediction block may be set as an initial block. Alternatively, in the case of a transform block, a block obtained by dividing a coding block of a higher level than the transform block may be set as an initial block.
[0150] The conditions for determining the size or form of the initial block are not fixed, and there may be cases where some of them are changed or excluded. In addition, the splitting operation of the current level (whether it is possible to split, the block form that can be split, etc.) may be affected according to at least one combination of the splitting state of the previous process or higher-level block (e.g., the size of the coding block, the form of the coding block) and the setting conditions of the current level (e.g., the supported size or form of the transform block).
[0151] The block partition unit can support a quadtree-based partitioning method. In other words, a block can be partitioned into four blocks with vertical and horizontal lengths of 1 / 2 of the original block. This means that the partitioning can be repeatedly performed until the available partition depth limit (dep_k, k means the number of available partition times, when the block has an available partition depth limit (dep_k), the block size is (M>>k, N>>k)).
[0152] In addition, a binary tree-based partitioning method can be supported. This means that a block can be divided into two blocks, wherein at least one of the horizontal length or the vertical length is 1 / 2 of the original block. Quadtree partitioning and binary tree partitioning can be symmetrical partitioning or asymmetrical partitioning. Which partitioning to use can be determined according to the setting of the encoder / decoder. In the image encoding method according to the present embodiment, it is mainly described with symmetrical partitioning.
[0153] Whether to divide each block can be indicated by using a division flag (div_flag). When the corresponding value is 1, the division is performed, and when the corresponding value is 0, the division is not performed. Alternatively, when the corresponding value is 1, the division is performed and additional division is available, and when the corresponding value is 0, the division is not performed and additional division is not available. The above flags can consider whether to perform division by using conditions such as the minimum available division size, the available division depth limit, etc., and may not consider whether to perform division separately.
[0154] The partition flag can be used for quadtree partitioning and can also be used for binary tree partitioning. In binary tree partitioning, the partition direction can be determined according to the partition depth, coding mode, prediction mode, size, form, type (which can be one of the coding type, prediction type, transform type, quantization type, entropy type or in-loop filter, or can be one of the brightness type or chroma type), slice type, available partition depth limit, and the minimum / maximum available partition size of the block, or according to a combination thereof. In addition, according to the partition flag or the corresponding partition direction or both, in other words, the block can be divided into 1 / 2 in width or 1 / 2 in length.
[0155] For example, assuming that when a block has an M×N (M>N) size and M is greater than N so that length partitioning is supported, and the current partition depth (dep_curr) is less than the available partition depth limit, so that additional partitioning is available, 1 bit is allocated to the above partition flag. When the corresponding value is 1, the partition is performed, otherwise there is no more available partitioning.
[0156] A single partition depth can be used for quadtree partition and binary tree partition, or each partition depth can be used for quadtree partition and binary tree partition. In addition, a single available partition depth limit can be used for quadtree partition and binary tree partition, or each available partition depth limit can be used for quadtree partition and binary tree partition.
[0157] As another example, when the block has an M×N (M>N) size and N is equal to the preset minimum available partition size so that horizontal partitioning is not supported, 1 bit is allocated to the above partition flag. When the corresponding value is 1, vertical partitioning is performed, otherwise no partitioning is performed.
[0158] In addition, for horizontal division and vertical division, respective flags (div_h_flag and div_h_flag) can be supported. According to the above flags, binary division can be supported. Whether to perform horizontal or vertical division of each block can be indicated by a horizontal division flag (div_h_flag) or a vertical division flag (div_v_flag). When the horizontal division flag (div_h_flag) or the vertical division flag (div_v_flag) is 1, horizontal or vertical division is performed, otherwise, horizontal or vertical division is not performed.
[0159] In addition, when each flag is 1, horizontal or vertical splitting is performed and additional horizontal or vertical splitting is available, and when each flag is 0, horizontal or vertical splitting is not performed and no additional horizontal or vertical splitting is available. The flag may consider whether to perform splitting by using conditions such as the minimum available splitting size, available splitting depth limit, etc., and may not consider whether to perform additional splitting.
[0160] In addition, a flag (div_flag / h_v_flag) for horizontal division or vertical division may be supported, and binary division may be supported according to the above flag. The division flag (div_flag) may indicate whether horizontal or vertical division is performed, and the flag (h_v_flag) may indicate whether the division is performed in the horizontal direction or in the vertical direction. When the division flag (div_flag) is 1, division is performed, and horizontal or vertical division is performed according to the division direction flag (h_v_flag). When the division flag (div_flag) is 0, horizontal or vertical division is not performed.
[0161] In addition, when the corresponding value is 1, horizontal or vertical division is performed according to the division direction flag (h_v_flag) and additional horizontal or vertical division is available. When the corresponding value is 0, horizontal or vertical division is not performed, and there is no horizontal or vertical division available. The above flags may consider whether to perform division by using conditions such as the minimum available division size, the available division depth limit, etc., and whether to perform additional division may not be considered.
[0162] Such a division flag can be used for each of the horizontal and vertical divisions, and binary tree division can be supported according to the above flag. In addition, when the division direction is predetermined, one of the two division flags can be used as the above example, or two division flags can be used.
[0163] For example, when all the above flags are used, the block can be divided into any of M×N, M / 2×N, M×N / 2, and M / 2×N / 2. Here, the flags can be encoded as 00, 10, 01, and 11 in the order of horizontal division flag or vertical division flag (div_h_flag / div_v_flag).
[0164] The above case is an example in which the division flag is set to be used in an overlapping manner, and the division flag can be set to be used without being overlapped. For example, the block can be divided in the form of M×N, M / 2×N, and M×N / 2. Here, the above flags can be encoded as 00, 01, 10 in the order of horizontal or vertical division flags. Alternatively, the above flags can be encoded as 00, 10, and 11 in the order of division flags (div_flag) and horizontal-vertical flags (h_v_flag) (a flag indicating whether the division is in the horizontal direction or the vertical direction). Here, the overlapping flag may mean that horizontal division and vertical division are performed simultaneously.
[0165] Any of the quadtree partitioning and binary tree partitioning as described above can be used in the encoder or decoder or both according to the settings of the encoder or decoder or both, or can be used in combination. For example, according to the block size or form, the quadtree partitioning or binary tree can be determined. In other words, when the block form is M×N and M is greater than N so that horizontal partitioning is performed, and when the block form is M×N and N is greater than M so that vertical partitioning is performed, binary tree partitioning can be supported. In other words, when the block form is M×N and N is equal to M, quadtree partitioning can be supported.
[0166] As another example, when the M×M block size is equal to or greater than the block partition boundary value (thrblksize), binary tree partitioning may be supported, otherwise quadtree partitioning may be supported.
[0167] As another example, when M or N of the M×N block is equal to or less than the first maximum available partition size (Maxblksize1) and equal to or greater than the first minimum available partition size (Minblksize1), quadtree partitioning may be supported. When M or N of the M×N block is equal to or less than the second maximum available partition size (Maxblksize2) and equal to or greater than the second minimum available partition size (Minblksize2), binary tree partitioning may be supported.
[0168] When the first partition support range and the second partition support range defined as the maximum available partition size and the minimum available partition size overlap each other, the priority order in the first and second partitions can be assigned according to the settings of the encoder and the decoder. In this embodiment, the first partition can be a quadtree partition, and the second partition can be a binary tree partition.
[0169] For example, when the first minimum available partition size (Minblksize1) is 16, the second maximum available partition size (Maxblksize2) is 64, and the block has a size of 64×64 before being partitioned, since the block belongs to the first partition support range and the second partition support range, both quadtree partitioning and binary tree partitioning are available. When a higher priority is assigned to the first partition (quadtree partitioning in this embodiment) according to a preset setting, and the partition flag (div_flag) is 1, quadtree partitioning is performed, and additional quadtree partitioning is available. When the partition flag (div_flag) is 0, quadtree partitioning is not performed, and quadtree partitioning is no longer performed.
[0170] The above flags may consider whether to perform partitioning by using conditions such as the minimum available partition size, the available partition depth limit, etc., and may not consider whether to perform additional partitioning. When the partition flag (div_flag) is 1, since the size is greater than the first minimum available partition size (Minblksize1), the block with a size of 32×32 is divided into four blocks, so quadtree partitioning can continue to be performed. When the partition flag (div_flag) is 0, since the current block size 64×64 belongs to the second partition support range, additional quadtree partitioning is not performed, and binary tree partitioning can be performed.
[0171] When the division flag (in the order of div_flag / h_v_flag) is 0, division is no longer performed. When the flag is 10 or 11, horizontal division or vertical division can be performed. When the block has a size of 32×32 before being divided and the division flag (div_flag) is 0 so that quadtree division is no longer performed, and the second maximum available division size (Maxblksize2) is 16, since the current block size 32×32 does not belong to the second division support range, division is no longer supported.
[0172] In the above description, the priority order of the division methods may be determined according to at least one of the slice type, the coding mode, and the luminance / chrominance components, or according to a combination thereof.
[0173] As another example, various settings may be supported according to luma and chroma components. For example, the structure of quadtree or binary tree partitioning determined in the luma component may be used as is in the chroma component without encoding / decoding additional information. Alternatively, when independent partitioning for luma and chroma components is supported, quadtree and binary tree partitioning may be supported for the luma component, and quadtree partitioning may be supported for the chroma component.
[0174] In addition, quadtree partitioning and binary tree partitioning can be supported for brightness and chrominance components, and the partition support range may or may not be set the same or proportionally for brightness and chrominance components. For example, when the color format is 4:2:0, the partition support range of the chrominance component may be N / 2 of the partition support range of the brightness component.
[0175] As another example, different settings may be set according to the slice type. For example, in an I slice, quadtree partitioning may be supported, in a P slice, binary tree partitioning may be supported, and in a B slice, quadtree partitioning and binary tree partitioning may be supported.
[0176] Quadtree partitioning and binary tree partitioning can be set and supported according to various conditions as in the above examples. The above examples are not specific to the above cases, but may include cases where the conditions are opposite to each other. Alternatively, at least one condition or a combination thereof of the above examples may be included. Alternatively, in other cases, modifications are also available. The above available partition depth limit may be determined according to at least one of the partitioning method (quadtree partitioning, binary tree partitioning), slice type, luminance / chrominance component, and coding mode, or according to a combination thereof.
[0177] In addition, the partition support range can be determined according to at least one of the partition method (quadtree partition, binary tree partition), slice type, luminance / chrominance component and coding mode or according to a combination thereof. The information for the partition support range can be represented by using the maximum and minimum values of the partition support range. When the information is configured with an explicit flag, the flag can represent the length information of each maximum and minimum value, or the difference between the minimum and maximum values.
[0178] For example, when the maximum value and the minimum value are configured with an exponential power of k (here, assuming that k is 2), the exponential information of the maximum value and the minimum value can be transmitted to the decoding device by encoding the exponential information using various binarization methods. Alternatively, the exponential difference of the maximum value and the minimum value can be transmitted. Here, the transmitted information can be the exponential information of the minimum value and the information of the exponential difference.
[0179] According to the above description, information related to a flag may be generated and transmitted in a sequence unit, a picture unit, a slice unit, a tile unit, a block unit, etc.
[0180] The block partition information may be represented by using the partition flag described in the above example or by using a combination of quadtree and binary tree partition flags. The partition flag may be transmitted to a decoding device by encoding information thereof using various methods such as a unary binarization method, a truncated unary binarization method, etc. The bitstream structure of the partition flag for representing the partition information of the block may be selected from at least one scanning method.
[0181] For example, the split flag may be configured in the bitstream based on the split depth order (from dep0 to dep_k), or the split flag may be configured in the bitstream based on whether the split is performed. The method based on the split depth order is a method of obtaining split information of the current depth level based on the initial block and obtaining split information in the next level of depth, and the method based on whether the split is performed is a method of obtaining additional split information in the split block based on the initial block first. Other scanning methods not described in the above examples may be included and selected.
[0182] In addition, according to an embodiment, the block partition unit may generate index information of a block candidate group having a predefined form instead of the described partition flag, and indicate the generated index information. The form of the block candidate group may be, for example, a form that a block may have before partitioning, and may include M×N, M / 2×N, M×N / 2, M / 4×N, 3M / 4×N, M×N / 4, M×3N / 4, M / 2×N / 2 forms.
[0183] When the candidate group of the partition block is determined as described above, the index information in the partition block form may be encoded using various methods such as a fixed length binarization method, a unary truncated binarization method, a truncated binarization method, etc. As the above-mentioned partition flag, the partition block candidate group may be determined according to at least one of a partition depth, a coding mode, a prediction mode, a size, a form, a type, and a slice type of a block, an available partition depth limit, and a minimum / maximum size of an available partition, or according to a combination thereof.
[0184] For the following description, it is assumed that the first candidate list (list 1) is (M×N, M×N / 2), the second candidate list (list 2) is (M×N, M / 2×N, M×N / 2, and M / 2×N / 2), the third candidate list (list 3) is (M×N, M / 2×N, M×N / 2), and the fourth candidate list (list 4) is (M×N, M / 2×N, M×N / 2, M / 4×N, 3M / 4×N, M×N / 4, M×3N / 4, M / 2×N / 2). For example, based on M×N, when M=N, the partition block candidates of the second candidate list can be supported, and when M≠N, the partition block candidates of the third candidate list can be supported.
[0185] As another example, when M or N of M×N is equal to or greater than the boundary value (blk_th), the partition block candidates of the second candidate list may be supported, otherwise, the partition block candidates of the fourth candidate list may be supported. In addition, when M or N is equal to or greater than the first boundary value (blk_th_1), the partition block candidates of the first candidate list may be supported. When M or N is less than the first boundary value (blk_th_1) but equal to or greater than the second boundary value (blk_th_2), the partition block candidates of the second candidate list may be supported. When M or N is less than the second boundary value (blk_th_2), the partition block candidates of the fourth candidate list may be supported.
[0186] As another example, when the encoding mode is intra prediction, the partition block candidates of the second candidate list may be supported, and when the encoding mode is inter prediction, the partition block candidates of the fourth candidate list may be supported.
[0187] Although the above-mentioned partition block candidates are supported, the bit configuration may be the same or different depending on the binarization method of each block. For example, as the above-mentioned partition flag, when the support of the partition block candidate is limited according to the block size or form, the bit configuration may vary according to the binarization method of the corresponding block candidate. For example, when M>N, according to the block form of horizontal partitioning, in other words, M×N, M×N / 2 and M / 2×N / 2 can be supported, and the binary bits of the index may be different from each other in M×N / 2 of the partition block candidate group (M×N, M / 2×N, M×N / 2, M / 2×N / 2) and in M×N / 2 according to the current conditions.
[0188] The information of block partitioning and block form may be represented by using one of a partition flag and a partition index method according to a block type used in, for example, a coding block, a prediction block, a transform block, a quantization block, an entropy block, an in-loop filtering block, etc. In addition, a block size limit and an available partition depth limit for supporting partitioning and block form may vary according to each block type.
[0189] After the coding block is determined, encoding and decoding of the block unit can be performed according to the process of determining the prediction block, determining the transform block, determining the quantization block, determining the entropy block, determining the in-loop filter, etc. The order of the above encoding and decoding processes is not fixed, and some orders can be changed or excluded. The size and form of each block can be determined according to the encoding cost of each candidate size and the form of the block, and the image data of each determined block and the division information such as the determined size and form of each determined block can be encoded.
[0190] The prediction unit 200 may be implemented by using a prediction module as a software module, or may be implemented by generating a prediction block of a block to be encoded using an intra prediction method or an inter prediction method. Here, the prediction block is a block that closely matches the block to be encoded in terms of pixel difference, and may be determined by using various methods including SAD (sum of absolute difference) and SSD (sum of square difference). In addition, various syntaxes that can be used when decoding an image block may be generated. The prediction block may be classified into an intra block and an inter block according to the encoding mode.
[0191] Intra prediction is a prediction method using spatial correlation, and involves a method of predicting a current block by using reference pixels of a reconstructed block that was previously encoded or decoded. In other words, the luminance values reconfigured by intra prediction and reconstruction are used as reference pixels in the encoder and decoder. Intra prediction can be effective for a planar area with continuity and an area with a predetermined direction, and since intra prediction uses spatial correlation, random access is ensured, and intra prediction can be used to prevent error diffusion.
[0192] Inter-frame prediction involves a compression method that uses temporal correlation to remove overlapping data by referring to at least one image that was previously encoded or can be encoded in a future picture. In other words, inter-frame prediction can generate a prediction signal with high similarity by referring to at least one of the previous or subsequent pictures. An encoder using inter-frame prediction can search for a block with a high correlation with the currently encoded block in the reference picture, and transmit the position information and residual signal of the selected block to the decoder. The decoder can generate the same prediction block as the image of the encoder by using the selected information of the transmitted image, and configure the reconstructed image by compensating the transmitted residual signal.
[0193] Figure 4 is an example diagram illustrating inter-frame prediction of a P slice in an image encoding and decoding method according to an embodiment of the present invention. Figure 5 is an example diagram illustrating inter prediction of a B slice in an image encoding and decoding method according to an embodiment of the present invention.
[0194] In the image encoding method of the present embodiment, since inter-frame prediction generates a prediction block from a previously encoded picture having high temporal correlation, encoding efficiency can be increased. Current (t) may refer to a current picture to be encoded and based on a temporal stream or POC (picture order count) of an image picture, including a first reference picture having a first temporal distance (t-1) before the POC of the current picture, and a second reference picture having a second temporal distance (t-2) before the first temporal distance.
[0195] In other words, Figure 4 As shown, the inter-frame prediction that can be used in the image encoding method of the present embodiment can perform motion estimation, which finds the optimal prediction block with high correlation from the already encoded reference pictures (t-1 and t-2) by performing block matching between the reference block of the reference picture (t-1 and t-2) and the current block (current (t)) of the current picture. In order to perform accurate estimation as needed, the final prediction block can be found by finding the optimal prediction block through interpolation based on a structure in which at least one sub-pixel is arranged between two adjacent pixels, and by performing motion compensation on the optimal prediction block.
[0196] In addition, if Figure 5 As shown, the inter-frame prediction that can be used in the image encoding method of this embodiment can generate a prediction block from reference pictures (t-1 and t+1) that have been encoded and exist in two directions in time based on the current picture (current (t)). In addition, two prediction blocks can be generated from one reference picture.
[0197] When an image is encoded by inter-frame prediction, information of a motion vector for an optimal prediction block and information of a reference picture may be encoded. In the present embodiment, when a prediction block is generated in one direction or in two directions, a prediction block may be generated from a corresponding reference picture list by configuring the reference picture list differently. Generally, a reference picture existing temporally before a current picture may be managed by being allocated to list 0 (L0), and a reference picture existing temporally after a current picture may be managed by being allocated to list 1 (L1).
[0198] When reference picture list 0 is configured and several available reference pictures are not filled, reference pictures existing after the current picture may be allocated. Similarly, when reference picture list 1 is configured and several available reference pictures are not filled, reference pictures existing after the current picture may be allocated.
[0199] Figure 6 is an exemplary diagram for illustrating generation of a prediction block in one direction in an image encoding and decoding method according to an embodiment of the present invention.
[0200] refer to Figure 6 According to the image encoding and decoding method of this embodiment, the prediction block can be traditionally found from the encoded reference pictures (t-1 and t-2), and in addition, the prediction block can be found from the area that has been encoded in the current picture (current (t)).
[0201] In other words, the image encoding and decoding method according to the present embodiment can be implemented to generate a prediction block from previously encoded pictures (t-1 and t-2) with high temporal correlation, and also to find a prediction block with high spatial correlation. Finding such a prediction block with high spatial correlation can correspond to finding a prediction block by using an intra-frame prediction method. In order to perform block matching from an area that has been encoded in the current picture, the image encoding method of the present embodiment can configure the syntax of information related to the prediction candidate by combining with the intra-frame prediction mode.
[0202] For example, when n (n is an arbitrary natural number) intra prediction modes are supported, n+1 modes can be supported by adding one mode to the intra prediction candidate group, and the intra prediction mode can be encoded using M bits that satisfy 2M-1≤n+1≤2M. In addition, selection from a candidate group of prediction modes with high probability, such as the most probable mode (MPM) of HEVC, can be achieved. In addition, it is also possible to preferentially encode a process higher than the encoding of the prediction mode.
[0203] When generating a prediction block by performing block matching in the current picture, the image encoding method can configure the syntax of the relevant information by combining the inter-frame prediction mode. As additional relevant prediction mode information, information related to motion or displacement can be used. The information related to motion or displacement may include information of the best vector candidate among multiple vector candidates, the difference between the best vector candidate and the actual vector, the reference direction, the reference picture information, etc.
[0204] Figure 7 is an example diagram of configuring a reference picture list in an image encoding and decoding method according to an embodiment of the present invention. Figure 8 is another example diagram of performing inter prediction from a reference picture list in an image encoding and decoding method according to an embodiment of the present invention.
[0205] refer to Figure 7 According to the image encoding method of this embodiment, inter-frame prediction can be performed on the current block of the current picture (current (t)) from the first reference picture list (reference list 0, L0) and the second reference picture list (reference list 1, L1).
[0206] refer to Figure 7 and Figure 8 , reference picture list 0 may be configured with reference pictures before the current picture (t), and t-1 and t-2 indicate reference pictures having a first temporal distance (t-1) and a second temporal distance (t-2) having (one or more) POCs before the POC of the current picture (t). In addition, reference picture list 1 may be configured with reference pictures after the current picture (t), and t+1 and t+2 indicate reference pictures having (one or more) POCs after the POC of the current picture (t).
[0207] The above example of configuring the reference picture list is described with an example of configuring the reference picture list using a reference picture having a temporal distance of 1 (based on POC in this example). However, it may be configured as a reference picture having a different temporal distance. In other words, this means that the index difference between the reference pictures and the temporal distance difference between the reference pictures may not be proportional. In addition, the list configuration order may not be configured based on the temporal distance. The description of the above features will be confirmed by an example of configuring a reference picture list to be described later.
[0208] Depending on the slice type (I, P, or B), prediction can be performed from the reference pictures present in the list. In addition, when a prediction block is generated by performing block matching in the current picture (current (t)), encoding can be performed using an inter-frame prediction method by adding the current picture to the reference picture list (reference list 0 or reference list 1 or both).
[0209] like Figure 8 As shown, the current picture (t) may be added to the reference picture list 0 (reference list 0), or the current picture (current (t)) may be added to the reference picture list 1 (reference list 1). In other words, the reference picture list 0 may be configured by adding a reference picture having a temporal distance (t) before the current picture (t) to the reference picture. Then, the reference picture list 1 may be configured by adding a reference picture having a temporal distance (t) after the current picture (t) to the reference picture.
[0210] For example, when configuring reference picture list 0, a reference picture before the current picture may be assigned to reference picture list 0, and then the current picture (t) may be assigned. When configuring reference picture list 1, a reference picture after the current picture may be assigned to reference picture list 1, and then the current picture (t) may be assigned. Alternatively, when configuring reference picture list 0, the current picture (t) may be assigned, and then the reference picture before the current picture may be assigned, and when configuring reference picture list 1, the current picture (t) may be assigned, and then the reference picture after the current picture may be assigned.
[0211] In addition, when configuring the reference picture list 0, the reference picture before the current picture may be allocated, the reference picture after the current picture may be allocated, and then the current picture (t) may be allocated. Similarly, when configuring the reference picture list 1, the reference picture after the current picture may be allocated, the reference picture before the current picture may be allocated, and then the current picture (t) may be allocated. The above examples are not limited to the above cases, but may include cases where the conditions are opposite to each other. Examples of other cases are also possible.
[0212] Whether to include the current picture in each reference picture list (for example, not added to any list, added to list 0, added to list 1, or added to list 0 and 1) can be set in the same manner in the encoder / decoder, and information thereon can be transmitted in sequence units, picture units, slice units, etc. Information thereon can be encoded by using a method such as a fixed-length binarization method, a unary truncated binarization method, a truncated binarization method, etc.
[0213] In the image encoding and decoding method of this embodiment, different from Figure 7 A method for selecting a prediction block by performing block matching in a current picture (t), configuring a reference picture list including information of the selected prediction block, and the reference picture list is used for encoding and decoding an image.
[0214] When configuring the reference picture list, the order and rules of each list, the number of available reference pictures, etc. can be set differently. The above features can be determined according to whether the current picture is included in the list (whether the current picture is included as a reference picture in inter-frame prediction), the slice type, the list reconfiguration parameter (which can be applied to each of list 0 and list 1, and can be applied to both list 0 and list 1), the position within the group of pictures (group of picture, GOP) and the temporal layer information (temporal id), or according to a combination thereof. The information for it can be explicitly transmitted in sequence units, picture units, etc.
[0215] For example, in the case of a P slice, the reference picture list 0 may follow the list configuration rule A regardless of whether the current picture is included in the list. In the case of a B slice, the reference picture list 0 in the current picture may follow the list configuration rule C, and the reference picture list 1 may follow the list configuration rule C, the reference picture list 0 not including the current picture may follow the list configuration rule D, and the reference picture list 1 may follow the list configuration rule E. Among the rules, rules B and D may be identical to each other, and rules C and E may be identical to each other. The list configuration rules may be configured in the same or modified manner as described in the reference picture list configuration example described above.
[0216] As another example, when the current picture is included in the list, the first number of available reference pictures may be set, or when the current picture is not included in the list, the second number of available reference pictures may be set. The first number of available reference pictures and the second number of available reference pictures may be the same or different. Basically, the difference between the first number of available reference pictures and the second number of available reference pictures may be set to 1.
[0217] As another example, when the current picture is included in the list and the list reconfiguration parameter is applied, all reference pictures may be included in the list reconfiguration candidate group in slice A, and some reference pictures may be included in the list reconfiguration candidate group in slice B. Here, slice A or B may be distinguished by whether the current picture is included in the list, temporal layer information, and the position within the GOP. Whether to be included in the candidate group may be determined by the POC of the reference picture or the index of the reference picture, the reference prediction direction (before / after the current picture), and whether the current picture is in the list.
[0218] According to the above configuration, when motion prediction of an I slice is performed, since a reference block encoded by inter prediction in a current picture can be used, inter prediction may be available or may be used.
[0219] In addition, when configuring the reference picture list, the index allocation or list configuration order may vary according to the slice type. In the case of an I slice, as an example of reference picture list configuration, a smaller index (e.g., such as idx=0, 1, 2) may be used by increasing the priority of the current picture (current (t)). In addition, the amount of bits used when encoding an image may be reduced by using a binarization method (fixed length binarization method, unary truncated binarization method, truncated binarization method, etc.) that uses the number of available reference pictures of the corresponding reference picture list as the maximum value.
[0220] In addition, in the case of a P or B slice, when the probability of selecting a reference picture of a current block as a prediction candidate by performing block matching in the current picture is determined to be lower than the probability of selecting a prediction candidate by using other reference pictures, by lowering the priority of performing block matching in the current picture and using a higher index (e.g., such as idx=C, C-1), the amount of bits when encoding the image can be reduced by using various binarization methods that use the number of available reference pictures of the corresponding reference picture list as a maximum value.
[0221] In the above example, the setting of the priority of the current picture can be configured in the same or modified method as described in the example of the reference picture list configuration. In addition, according to the slice type (e.g., I slice), the information of the reference picture can be omitted by not configuring the reference picture list. For example, the prediction block can be generated by using the conventional inter prediction, but the inter prediction information can be represented by the motion information in the inter prediction mode excluding the reference picture information.
[0222] Whether a method for performing block matching in the current picture is supported may be determined according to the slice type. For example, the method may be set to support block matching in the current block in an I slice, but not in a P slice or a B slice. Modifications to other examples are also possible.
[0223] In addition, whether a method of performing block matching in the current picture is supported may be determined in picture units, slice units, tile units, etc., or may be determined according to a position within a GOP, temporal layer information (temporal ID), etc. The setting information may be transmitted while encoding, or may be transmitted from an encoder to a decoder in sequence units, picture units, slice units, etc.
[0224] In addition, although the setting information or syntax related to the above appears in a higher-level unit and the operation related to the setting is turned on, and the same setting information or syntax as the above appears in a lower-level unit, the setting information of the lower-level unit may have a higher priority than the setting information of the higher-level unit. For example, when the same or similar setting information is processed in a sequence unit, a picture unit, and a slice unit, the picture unit may have a higher priority than the sequence unit, and the slice unit may have a higher priority than the picture unit.
[0225] Fig. 9 is an exemplary diagram for illustrating intra prediction in an image encoding method according to an embodiment of the present invention.
[0226] refer to Fig. 9 According to the present embodiment, the intra-frame prediction method may include a series of steps of reference sample filling, reference sample filtering, intra-frame prediction and boundary filtering.
[0227] Reference pixel padding may be an example of reference pixel configuration, reference pixel filtering may be performed by a reference pixel filtering unit, intra prediction may include generating a prediction block and encoding a prediction mode, and boundary filtering may be an example of post-processing filtering.
[0228] In other words, the intra-frame prediction performed in the image encoding method of the present embodiment may include configuring reference pixels, filtering reference pixels, generating prediction blocks, encoding prediction modes, and post-processing filtering. One or part of the above steps may be omitted according to, for example, block size, block form, block position, prediction mode, prediction method, quantization parameter, etc. Alternatively, other steps may be added, or the steps may be changed in an order different from the above order.
[0229] The above-mentioned configuration of reference pixels, filtering of reference pixels, generation of prediction blocks, encoding of prediction modes, and post-processing filtering can be implemented in the form of being executed by a processor connected to a memory storing a software module. Therefore, in the following description, for the convenience of description, as a functional unit or configuration unit or both that performs its functions generated by a combination of a software module that implements each step and a processor that executes the step, each of the reference pixel configuration unit, the reference pixel filtering unit, the prediction block generation unit, the prediction mode encoding unit, and the post-processing filtering unit is referred to as an execution unit of each step.
[0230] Describing each configuration in more detail, the reference pixel configuration unit configuration will be used to predict the reference pixels of the current block by performing reference pixel filling. When the reference pixel does not exist or is unavailable, the reference pixel filling can be used to fill the reference pixel by copying the pixel value of the pixel closest to the available pixel. When copying the pixel value, a reconstructed picture buffer or a decoded picture buffer (DPB) can be used.
[0231] In other words, intra prediction performs prediction by using reference pixels of blocks that have been encoded before the current picture. To this end, in the configuration of reference pixels, neighboring pixels of neighboring blocks of the current block can be mainly used as reference pixels, in other words, left, upper left, lower left, upper side, and upper right side blocks can be mainly used as reference pixels.
[0232] However, the candidate group of neighbor blocks for reference pixels is an example obtained according to raster scanning or z-scanning of the coding order of the blocks. When the candidate group of neighbor blocks is obtained according to a coding order scanning method such as reverse z-scanning, in addition to the above side blocks, neighbor pixels of the right, lower right and lower side blocks can be used as reference pixels.
[0233] In addition, according to an embodiment, in addition to neighboring pixels, additional pixels may be replaced or may be used in combination with conventional reference pixels according to the configuration of each step of intra prediction.
[0234] In addition, when prediction is performed by a direction mode in an intra prediction mode, reference pixels may be generated by performing linear interpolation using reference pixels of integer units. Modes for performing prediction using reference pixels present at positions of integer units include multiple modes such as vertical direction, horizontal direction, 45 degrees, and 135 degrees. For the above prediction modes, generating reference pixels of real numbers may not be necessary.
[0235] Reference pixels interpolated in a prediction mode having other directions than the above prediction modes may have interpolation accuracy of an exponential power of 1 / 2 such as 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, and 1 / 64, or may have accuracy of a multiple of 1 / 2.
[0236] This is because the interpolation accuracy can be determined according to the number of supported prediction modes or the prediction direction of the prediction mode. For pictures, slices, tiles, and blocks, constant interpolation accuracy can be supported. Alternatively, adaptive interpolation accuracy can be supported according to block size, block form, prediction direction of supported modes, etc. Here, the prediction direction of the mode can be presented as tilt information or angle information of the direction indicated by the mode based on a specific line (e.g., the positive <+> x-axis on the coordinate plane).
[0237] As an interpolation method, linear interpolation can be performed by using neighboring integer pixels. However, other interpolation methods can also be supported. For interpolation, at least one filter type and several taps can be supported, for example, a 6-tap Wiener filter, an 8-tap Kalman filter. Which interpolation to perform can be determined based on block size, prediction direction, etc. In addition, information thereon can be transmitted in sequence units, picture units, slice units, block units, etc.
[0238] The reference pixel filtering unit can perform filtering on the reference pixels after configuring the reference pixels, thereby reducing the degradation generated during encoding and improving the prediction efficiency. The reference pixel filtering unit can implicitly or explicitly determine the filter type and whether to apply filtering according to the block size, block form and prediction mode. In other words, when a tap filter is applied, the filter coefficients can be determined differently according to the filter type. For example, a 3-tap filter such as [1, 2, 1] / 4 and [1, 6, 1] / 8 can be used.
[0239] In addition, the reference pixel filtering unit can determine whether to apply filtering by determining whether to send additional bits. For example, in the implicit case, the reference pixel filtering unit can determine whether to apply filtering based on the characteristics (distribution, standard deviation, etc.) of the pixels in the neighbor reference block.
[0240] In addition, the reference pixel filtering unit can determine whether to apply filtering when the relevant flag satisfies a preset concealment condition such as a residual coefficient, an intra-frame prediction mode, etc. The number of taps can be set, for example, 3 taps for a small block (blk) such as [1, 2, 1] / 4, and 5 taps for a large block (blk) such as [2, 3, 6, 3, 2] / 16. The number of times of application can be determined by whether filtering is performed, whether filtering is performed once, and whether filtering is performed twice.
[0241] In addition, the reference pixel filtering unit may apply filtering substantially to the neighbor reference pixels closest to the current block. In addition to the closest neighbor reference pixels, the reference pixel filtering unit may consider applying filtering to additional reference pixels. For example, when replacing the closest neighbor reference pixels, filtering may be applied to the additional reference pixels, or filtering may be applied by combining the closest neighbor reference pixels with the additional reference pixels.
[0242] The filtering may be applied constantly or adaptively. This may be determined based on the size of the current block or the size of the neighboring block, the coding mode of the current block or the neighboring block, the block boundary characteristics of the current block and the neighboring block (e.g., whether the boundary is the boundary of the coding unit or the boundary of the transform unit, etc.), the prediction mode or direction of the current block or the neighboring block, the prediction method of the current block or the neighboring block, at least one of the quantization parameters, or a combination thereof.
[0243] The determination thereof can be set identically in the encoder / decoder (implicitly), or can be determined by considering the encoding cost (explicitly). Basically, the filter applied is a low-pass filter, and the number of taps, filter coefficients, whether to encode the filter flag, the number of filter application times, etc. can be determined by the above factors. In addition, the information thereof can be set in sequence units, picture units, slice units, and block units, and can be transmitted to the decoder.
[0244] When performing intra prediction, the prediction block generation unit may generate a prediction block by using an extrapolation method using reference pixels, an interpolation method such as using an average value of reference pixels (DC mode), and a planar mode, or a method of copying reference pixels.
[0245] In the method of duplicating reference pixels, at least one predicted pixel may be generated by duplicating a single reference pixel, or at least one predicted pixel may be generated by duplicating at least one reference pixel. The number of duplicated reference pixels may be equal to or less than the number of duplicated predicted pixels.
[0246] In addition, according to the prediction method, the prediction method can be classified into a directional prediction method and a non-directional prediction method. Specifically, the directional prediction method is classified into a linear directional method and a curved directional method. The linear directional method uses an extrapolation method and generates pixels of a prediction block by using reference pixels located on a prediction direction line. The curved directional method uses an extrapolation method and generates pixels of a prediction block by using reference pixels located on a prediction direction line. However, considering the detailed direction of the block (e.g., an edge), part of the prediction direction of the pixel unit can be changed. In the image encoding and decoding method of the present embodiment, the linear directional method is mainly described as a directional prediction mode.
[0247] In addition, in the directional prediction method, the intervals between adjacent prediction modes may be uniform or non-uniform, and this may be determined according to the block size or form. For example, when the current block is divided into blocks of size M×N by a block partition unit, and M is equal to N, the intervals between the prediction modes may be uniform, and when M is not equal to N, the intervals between the prediction modes may be non-uniform.
[0248] As another example, when M is greater than N, in the vertical direction mode, a precise interval may be set between prediction modes close to the vertical mode (90 degrees), and a wide interval may be allocated between prediction modes far from the vertical mode. When N is greater than M, in the horizontal direction mode, a precise interval may be set between prediction modes close to the horizontal mode (180 degrees), and a wide interval may be allocated between prediction modes far from the horizontal mode.
[0249] The above examples are not limited to the above cases, and may include cases where the conditions are opposite to each other, and examples of other cases may also be modified. Here, the interval between prediction modes may be calculated based on a numerical value representing the direction of each mode. The direction of the prediction mode may be quantified in the tilt information or angle information of the direction.
[0250] In addition, in addition to the above method, other methods may be included to generate a prediction block by using spatial correlation. For example, a reference block using inter prediction such as motion search and compensation using a current picture as a reference picture may be generated as a prediction block.
[0251] When generating a prediction block, the prediction block can be generated by using reference pixels according to the above prediction method. In other words, the prediction block can be generated by using a directional prediction or non-directional prediction method such as extrapolation, interpolation, replication and averaging of conventional intra-frame prediction according to the above prediction method, and the prediction block can be generated by using an inter-frame prediction method or by using other additional methods.
[0252] The intra prediction method can be supported under the same setting of the encoder / decoder and can be determined according to the slice type, block size, and block form. The intra prediction method can be supported according to at least one of the described prediction methods or according to a combination thereof. The intra prediction mode can be configured according to the supported prediction method. The number of supported intra prediction modes can be determined according to the prediction method, slice type, block size, block form, etc. Information thereon can be set and transmitted in sequence units, picture units, slice units, block units, etc.
[0253] The encoding of the prediction mode may determine a mode optimized in terms of encoding cost according to each prediction mode as the prediction mode of the current block.
[0254] In one embodiment, in order to reduce the bit amount of the prediction mode, the prediction mode encoding unit can use the mode of at least one neighboring block to predict the mode of the current block. A mode with a higher probability (most_probable_mode, MPM) that is the same as the mode of the current block can be included in the candidate group. The mode of the neighboring block can be included in the above candidate group. For example, the prediction modes of the upper left, lower left, upper side, and upper right side blocks of the current block can be included in the above candidate group.
[0255] The candidate group of the prediction mode may be configured according to at least one of the position of the neighbor block, the priority of the neighbor block, the priority in the partition block, the size or form of the neighbor block, the preset specific mode, and the prediction mode of the luminance block (in the case of the chrominance block) or according to a combination thereof. Information thereon may be transmitted in sequence units, picture units, slice units, block units, etc.
[0256] For example, when a block adjacent to a current block is divided into at least two blocks, it may be determined under the same setting of an encoder / decoder which one of the two divided blocks has a mode to be included in the mode prediction candidates of the current block.
[0257] In addition, for example, when a left block (M×M) among neighbor blocks of the current block is configured with three partition blocks by quadtree partitioning of a block partition unit so that the left block includes M / 2×M / 2, M / 4×M / 4, and M / 4×M / 4 blocks in a top-to-bottom direction, based on the block size, the prediction mode of the M / 2×M / 2 block may be included as a mode prediction candidate for the current block.
[0258] As another example, when the upper block (N×N) of the neighbor block of the current block is configured with three partition blocks through binary tree partitioning of the block partition unit so that the upper block includes N / 4×N, N / 4×N and N / 2×N blocks from left to right, according to a preset order (priority is assigned from left to right), the prediction mode of the first N / 4×N block on the left is included as a mode prediction candidate for the current block.
[0259] As another example, when the prediction mode of the neighbor block of the current block is a directional prediction mode, prediction modes adjacent to the corresponding mode (tilt information or angle information of the direction of the above mode) may be included in the mode prediction candidate group of the current block.
[0260] In addition, a preset mode (plane, DC, vertical, horizontal, etc.) may be preferentially included according to the prediction mode configuration of the neighboring block or a combination thereof. In addition, a prediction mode with a high frequency of occurrence in the prediction mode of the neighboring block may be preferably included. The above priority may mean the probability of being included in the mode prediction candidate group of the current block, and the probability of being assigned a higher priority or index in the above candidate group configuration (in other words, the probability of being assigned fewer bits in the binarization process).
[0261] As another example, when the maximum number of modes within the mode prediction candidate group of the current block is k, the left block is configured as m blocks having a vertical length shorter than the vertical length of the current block, and the upper block is configured as n blocks having a horizontal length higher than the horizontal length of the current block, the candidate group is filled according to a preset order (from left to right, from top to bottom) when the sum (m+n) of the divided blocks of the neighboring block is greater than k. When the sum (m+n) of the divided blocks of the neighboring block is greater than the maximum number k, the prediction mode of the neighboring block (the left block and the upper block) and the prediction mode of the blocks of other neighboring blocks (e.g., lower left, upper left, upper right, etc.) other than the neighboring block may be included in the mode prediction candidate group of the current block. The above examples are not limited to the above cases, but may include cases where the conditions are opposite to each other, and examples of other cases are also possible.
[0262] Therefore, the candidate blocks for predicting the mode of the current block are not limited to blocks at specific positions. Prediction mode information of at least one block located at the left, upper left, lower left, upper side, and upper right blocks can be used. Considering the various features as the above examples, the prediction mode of the current block can form a candidate group.
[0263] The prediction mode encoding unit can classify the prediction mode into a mode candidate group (referred to as the first candidate group in this example) having a high probability of being the same as the mode of the current block and another mode candidate group (referred to as the second candidate group in this example). The prediction mode encoding process can be changed depending on which of the two candidate groups the prediction mode of the current block belongs to.
[0264] All prediction modes can be configured by summing the prediction modes of the first candidate group and the prediction modes of the second candidate group. The number of prediction modes of the first candidate group and the number of prediction modes of the second candidate group can be determined according to at least one of a slice type, a block size, and a block form, or according to a combination thereof. Depending on the candidate group, the same binarization method or different binarization methods can be applied.
[0265] For example, a fixed length binarization method can be applied to the first candidate group, and a unary truncated binarization method can be applied to the second candidate group. In the above description, two candidate groups are used as examples. The candidate group can be extended to a first mode candidate group with a high probability of being the same as the mode of the current block, a second mode candidate group with a high probability of being identical to the mode of the current block, and another mode candidate group, etc., and variations thereof are also possible.
[0266] Considering that the correlation between the reference pixels adjacent to the boundary between the current block and the neighboring block and the pixels within the adjacent current block is high, the post-processing filtering performed by the post-processing filtering unit may replace several predicted pixels of the previously generated prediction block with values generated by performing filtering on at least one reference pixel and at least one predicted pixel adjacent to the boundary. Alternatively, the post-processing filtering unit may also replace the predicted pixels with values generated by applying feature quantitative values (e.g., differences between pixel values, tilt information, etc.) between the reference pixels adjacent to the boundary of the block to the filtering. In addition to the above method, other methods having similar purposes (correcting several predicted pixels of a prediction block by using reference pixels) may also be added thereto.
[0267] In the post-processing filtering unit, the filter type and whether to apply filtering can be determined implicitly or explicitly. The reference pixel, the position and number of the current pixel, and the type of applied prediction mode used in the post-processing filtering unit can be set in the encoder / decoder. The information related to the above can be transmitted in units of sequence units, picture units, slice units, etc.
[0268] In addition, in the post-processing filtering, additional post-processing such as boundary filtering may be performed after the prediction block is generated. In addition, considering the characteristics of the pixels of the reference block, a post-processing filter similar to the above boundary filtering may be performed on the current block. The current block is reconstructed by summing the residual signal obtained by performing the transform / quantization, the inverse process of the residual signal after obtaining it, and the prediction signal.
[0269] Finally, a prediction block is selected or obtained by using the above process. Information related to the above may be information of a prediction mode and may be transmitted to the transform unit 210 so that a residual signal is encoded after obtaining the prediction block.
[0270] Fig.10 is an exemplary diagram for illustrating a prediction principle in a P slice or a B slice in an image encoding method according to an embodiment of the present invention. Fig.11 is used to show Fig.10 An example diagram of interpolation performed in an image encoding method.
[0271] Referring to 10, the image encoding method according to the present embodiment may include a motion estimation step (motion estimation module) and an interpolation step. The motion vector, reference picture index, and reference direction information generated in the motion estimation step may be transmitted to the interpolation step. In motion estimation and interpolation, the value stored in the reconstructed picture buffer (decoded picture buffer, DPB) may be used.
[0272] In other words, the image coding device may perform motion estimation to find a block similar to the current block from a previously encoded picture. In addition, in order to perform a more accurate prediction than a real number unit, the image coding device may perform interpolation of a reference picture. Finally, the image coding device may obtain a predicted block through a predictor (predictor). Information related to the above may be a motion vector, a reference picture index (or reference index), a reference direction, etc. The image coding device may then encode the residual signal.
[0273] In this embodiment, since intra prediction is performed in a P slice or a B slice, the image encoding method can support inter prediction and intra prediction. Fig.11 This can be achieved by a combination of methods.
[0274] like Fig.11 As shown, the image encoding method according to this embodiment may include: reference sample filling, reference pixel filtering, intra-frame prediction, boundary filtering, motion estimation and interpolation.
[0275] In an image coding device, when block matching is supported in the current picture, the prediction method in the I slice can be Fig.11 The configuration shown instead of Fig. 9In other words, the image encoding device can use the prediction mode in the I slice and information such as motion vector, reference picture index, reference direction, etc. generated in the P slice or B slice to generate a prediction block. However, since the reference picture is the current picture, some information may be omitted. In one embodiment, when the reference picture is the current picture, the reference picture index and reference direction may be omitted.
[0276] In addition, in an image encoding device, when interpolation is applied, due to the characteristics of the image, that is, due to the artificial characteristics of the image caused by computer graphics, block matching may not need to be performed until the real number unit, so whether to perform block matching can be set in the encoder, and can be set in sequence units, picture units, slice units, etc.
[0277] For example, the image encoding device may not perform interpolation on the reference picture used for inter-frame prediction according to the setting of the encoder. Various settings are available, such as performing interpolation when performing block matching in the current picture. In other words, in the image encoding device of the present embodiment, it is possible to set whether to perform interpolation of the reference picture. Here, it is possible to determine whether to perform interpolation of all or part of the reference pictures constituting the reference picture list.
[0278] In one embodiment, the image encoding device may not perform interpolation on some blocks when block matching does not need to be up to real units because the image has artificial features in which reference blocks exist, and the image encoding device may perform interpolation when block matching needs to be up to real units because the image is a natural image.
[0279] In addition, in the image encoding device, it can be set whether to apply block matching to a reference picture for performing interpolation in block units. For example, when a natural image and an artificial image are combined, interpolation can be performed on the reference picture. When an optimal motion vector is obtained by searching a portion of an artificial image, the motion vector can be represented in a predetermined unit (here, an integer unit). In addition, selectively, when an optimal motion vector is obtained by searching a portion of a natural image, the motion vector can be represented in another predetermined unit (here, a 1 / 4 unit).
[0280] Fig.12 It is a view for illustrating a main process of an image encoding method according to an embodiment of the present invention in syntax of a coding unit.
[0281] refer to Fig.12, curr_pic_BM_enabled_flag may represent a flag indicating whether block matching is allowed in the current picture, and may be defined and transmitted in sequence units and picture units. Here, generating a prediction block by performing block matching in the current picture may mean a case where an operation is performed by inter prediction. In addition, it may be assumed that cu_skip_flag, which is an inter method of not encoding a residual signal, is a flag supported in a P slice or a B slice other than an I slice. Here, when curr_pic_BM_enabled_flag is ON, in an I slice, block matching (BM) may be supported in an inter prediction mode.
[0282] In the image encoding method according to the present embodiment, when a prediction block is generated by performing block matching in the current picture, skipping can be supported, and in addition to block matching, skipping of other inter-frame methods can also be supported. In addition, skipping may not be supported in an I slice according to conditions. Whether skipping is supported can be determined according to the setting of the encoder.
[0283] In one embodiment, when skipping is supported in an I slice, by using a specific flag of if(cu_skip_flag), the prediction block can be directly reconstructed as a reconstructed block by performing block matching instead of encoding the residual signal in association with prediction_unit() as a prediction unit. In addition, the image encoding apparatus can classify a method of using a prediction block by performing block matching in a current picture as an inter-frame prediction method, and distinguish the classified methods by using a specific flag of pred_mode_flag.
[0284] In addition, the image encoding device according to the present embodiment can set the prediction mode to the inter-frame prediction mode (MODE_INTER) when pred_mode_flag is 0, and set it to the intra-frame prediction mode (MODE_INTRA) when pred_mode_flag is 1. The above method may be similar to the conventional intra-frame method. In order to distinguish it from the conventional structure, it may be classified as an inter-frame method or an intra-frame method in an I slice. In other words, the image encoding device of the present embodiment may not use the temporal correlation in the I slice, but may use the structure of the temporal correlation. part_mode means information on the block size and block form of the blocks divided in the coding unit.
[0285] Some examples of syntax used in this embodiment are provided below.
[0286] The syntax of sps_curr_pic_ref_enabled_flag in the sequence parameters may be a flag indicating whether IBC is used.
[0287] The syntax of pps_curr_pic_ref_enabled_flag in the picture parameters may be a flag indicating whether IBC is used in picture units.
[0288] In addition, whether to increase NumPicTotalCurr of the current picture is determined according to the ON / OFF of the syntax of pps_curr_pic_ref_enabled_flag.
[0289] When the reference picture list 0 is configured, according to pps_curr_pic_ref_enabled_flag, it may be determined whether the current picture is included in the reference picture.
[0290] When the reference picture list 1 is configured, according to pps_curr_pic_ref_enabled_flag, it can be determined whether to add the reference picture list of the current picture.
[0291] According to the above configuration, according to the above flag, it can be determined whether the current picture is included in the reference picture list during inter prediction.
[0292] Fig.13 is used to show that when Fig.12 An example diagram of an example of supporting symmetric partitioning or asymmetric partitioning like inter-frame prediction when performing block matching in the current picture used to generate a prediction block.
[0293] refer to Fig.13 When a prediction block is generated by performing block matching in the current picture, the image encoding method according to the present embodiment can support symmetric partitioning such as 2N×2N, 2N×N, N×2N and N×Nm, or asymmetric partitioning such as nL×2N, nR×2N, 2N×nU and 2N×nD.
[0294] Fig.14 is used to show that inter prediction supports 2Nx2N and NxN as Fig. 9 Various sizes and forms of blocks can be determined according to the division method of the block partition unit.
[0295] refer to Fig.14 , the image encoding method according to the present embodiment can support 2N×2N and N×N blocks as prediction block forms used in conventional intra prediction. This is an example showing that the block partition unit supports a square form by using a quadtree partitioning method or a partitioning method according to a predefined predetermined block candidate group. In addition, in intra prediction, another block form can be supported by adding a rectangular form to a binary tree partitioning method or a predefined predetermined block candidate group. The setting thereof can be performed in the encoder.
[0296] In addition, in the encoder, it is possible to set whether to apply skip when performing block matching (ref_idx=curr) on the current picture during intra prediction, whether to apply skip to conventional intra prediction, and whether to apply new intra prediction to a prediction block having another (other) partition form (reference Fig.13 ). The information thereof can be transmitted in sequence units, picture units, slice units, etc.
[0297] Subtractor 205 (reference Figure 2 ) can generate a residual block by subtracting the pixel values of the prediction block generated by the prediction unit 200 from the pixel values of the current block to be encoded, and by deriving pixel difference values.
[0298] Transformation unit 210 (reference Figure 2 ) receives a residual block as a difference between the current block and the prediction block generated by intra prediction or inter prediction from the subtractor 205, and transforms the received residual block into the frequency domain. Each pixel of the residual block is transformed into a transformation coefficient of the transformation block through a transformation process. The size and form of the transformation block may be equal to or smaller than the size and form of the coding unit. In addition, the size and form of the transformation block may be equal to or smaller than the size and form of the prediction unit. The image encoding device may perform the transformation by grouping a plurality of prediction units.
[0299] The size or form of the transform block can be determined by the block partition unit, and can support transforming to a square form or a rectangular form according to the block partitioning. The settings related to the transforms supported in the encoder / decoder (the size and form of the supported transform blocks) can affect the operation of the block partitioning.
[0300] The size and form of each transform block may be determined according to the encoding cost of each candidate for the size and form of the transform block, and information such as division of image data of each determined transform block, and the size and form of each determined transform block, etc. may be encoded.
[0301] The transformation can be performed by a one-dimensional transformation matrix. For example, each transformation matrix can be adaptively used in a discrete cosine transform (DCT) unit, a discrete sine transform (DST) unit, a horizontal unit, and a vertical unit. Adaptive use can be determined based on, for example, block size, block form, block type (luminance / chrominance), coding mode, prediction mode information, quantization parameter, coding information of neighboring blocks, etc.
[0302] For example, in the case of intra prediction, when the prediction mode is horizontal, a DCT-based transform matrix may be used in the vertical direction, and a DST-based transform matrix may be used in the horizontal direction. In addition, when the prediction mode is vertical, a DCT-based transform matrix may be used in the horizontal direction, and a DST-based transform matrix may be used in the vertical direction.
[0303] The transformation matrix is not limited to the above description. Information about it can be determined by using an implicit or explicit method, and can be determined according to at least one of block size, block form, coding mode, prediction mode, quantization parameter, coding information of neighbor blocks, or a combination thereof. Information can be transmitted in sequence units, picture units, slice units, block units, etc.
[0304] Here, when the explicit method is used and at least two transformation matrices are included as candidate groups for horizontal and vertical directions, information on which transformation matrix is used for each direction may be transmitted. Alternatively, when two transformation matrices for each horizontal and vertical direction are grouped into a pair and at least two pairs are included in the candidate group, information on which transformation matrices are used for the horizontal and vertical directions may be transmitted.
[0305] In addition, taking into account the characteristics of the image, the entire or part of the transform can be omitted. For example, the transform of any one or both of the horizontal and vertical components can be omitted. When intra-frame prediction or inter-frame prediction is not performed well, so that a large difference is generated between the current block and the predicted block, in other words, when the residual component is large and the transform is performed, in such a case, the loss can become larger when encoding is performed. The omission of the transform can be determined based on at least one of the coding mode, prediction mode, block size, block form, block type (luminance / chrominance), quantization parameter, and coding information of neighboring blocks, or a combination thereof. According to the above conditions, the omission of the transform can be expressed by using an implicit or explicit method. Information thereon can be transmitted in units of sequence units, picture units, slice units, etc.
[0306] Quantization unit 215 (reference Figure 2 ) performs quantization on the transformed residual component. The quantization parameter may be determined in block units, and the quantization parameter may be set in sequence units, picture units, slice units, block units, etc.
[0307] In one embodiment, the quantization unit 215 may predict the current quantization parameter by using one or at least two quantization parameters derived from neighboring blocks of the current block, such as left, upper left, upper side, upper right, and lower left blocks.
[0308] In addition, when there is no quantization parameter predicted from a neighbor block, in other words, when the block is located at a boundary of a picture, a slice, etc., the quantization unit 215 may output or transmit a difference value from a basic parameter transmitted in sequence units, picture units, slice units, etc. When there is a quantization parameter predicted from a neighbor block, the difference value may be transmitted by using the quantization parameter of the corresponding block.
[0309] The priority of the block from which the quantization parameter is derived may be preset, or may be transmitted in sequence units, picture units, slice units, etc. The residual block may be quantized by using a dead zone uniform threshold quantization (DZUTQ) method, a quantization weighting matrix method, or a method modified therefrom. Therefore, at least one quantization method may be included as a candidate, and the method may be determined based on a coding mode, prediction mode information, etc.
[0310] For example, in the quantization unit 215, it can be set whether to apply the quantization weighting matrix to the inter-frame coding unit and the intra-frame coding unit, etc. In addition, different quantization weighting matrices can be applied according to the intra-frame prediction mode. Assuming that the quantization weighting matrix has an M×N size and the block has a size equivalent to the size of the quantization block, the quantization coefficient can be applied differently to each position of each frequency component. In addition, the quantization unit 215 can select one from various traditional quantization methods, and the quantization unit 215 can be used in the encoder / decoder under the same setting. Information thereon can be transmitted in sequence units, picture units, slice units, etc.
[0311] at the same time, Figure 2 and Figure 3 The dequantization units 220 and 315 and the inverse transform units 225 and 320 shown in FIG. 2 can be implemented by inversely performing the processes of the above transform unit 210 and the quantization unit 215. In other words, the dequantization unit 220 can perform dequantization on the quantized transform coefficients generated in the quantization unit 215, and the inverse transform unit 225 can generate a reconstructed residual block by inversely transforming the dequantized transform coefficients.
[0312] exist Figure 2 and 3 The adders 230 and 324 shown in FIG. 2 can generate a reconstructed block by adding the pixel values of the prediction block generated by the prediction unit to the pixel values of the reconstructed residual block. The reconstructed block can be stored in the decoded picture buffers 240 and 335 and provided to the prediction unit and the filtering unit.
[0313] The filtering units 235 and 330 may apply in-loop filters such as a deblocking filter, a sample adaptive offset (SAO), and an adaptive loop filter (ALP) to the reconstructed block. The deblocking filter may perform filtering on the reconstructed block to remove distortion generated in the block boundary during encoding and decoding. SAO is a filtering process that reconstructs the offset difference between the original image and the reconstructed image in pixel units for the residual block. ALF may perform filtering to minimize the difference between the predicted block and the reconstructed block. ALF may perform filtering based on a comparison value between a block reconstructed by a deblocking filter and a current block.
[0314] The entropy encoding unit 245 may entropy encode the transform coefficients quantized by the quantization unit 215. For example, the entropy encoding unit 245 may perform methods such as context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) encoding, and other encoding methods.
[0315] In the entropy coding unit 245, a bit sequence in which the quantization coefficient is encoded and various types of information for decoding the bit sequence may be included in the encoded data. The encoded data may include a block form of encoding, a quantization coefficient and a bit sequence in which the quantization coefficient is encoded, and information required for prediction, etc. In the case of a quantization coefficient, a two-dimensional quantization coefficient may be scanned to one dimension. The quantization coefficient may have different distributions according to image features. In particular, in the case of intra-frame prediction, since the coefficient distribution may have a specific distribution according to the prediction mode, the scanning method may be set differently.
[0316] In addition, the entropy encoding unit 245 may be set differently according to the size of the block to be encoded. At least one of various patterns such as a zigzag pattern, a diagonal pattern, a raster pattern, etc. may be set as a scanning pattern or a candidate thereof. The scanning pattern may be determined by a coding mode, prediction mode information, etc., or may be used in an encoder and a decoder under the same setting. Information thereon may be transmitted in sequence units, picture units, slice units, etc.
[0317] The quantized block (hereinafter, quantized block) input to the entropy encoding unit 245 may have a size equal to or smaller than the size of the transform block. In addition, the quantized block may be divided into at least two sub-blocks. When the quantized block is divided, the scanning pattern in the divided block thereof may be set to be identical to or different from the scanning pattern of the original quantized block.
[0318] For example, when the original quantized block is scanned in a zigzag scan pattern, the zigzag scan pattern may be applied to all sub-blocks. Alternatively, the zigzag scan pattern may be applied to the sub-block located on the upper left side of the block including the DC component, and the diagonal scan pattern may be applied to the remaining blocks. The determination thereof is performed according to the encoding mode, prediction mode information, etc.
[0319] In addition, in the entropy coding unit 245, the start position of the scanning pattern basically starts from the upper left position. However, according to image characteristics, the start position may be at the upper right position, the lower right position, or the lower left position. Information on which one of the at least two candidate groups is selected may be transmitted in sequence units, picture units, slice units, etc. As an encoding method, an entropy encoding method may be used, but it is not limited thereto.
[0320] at the same time, Figure 2 and 3 The dequantization of the dequantization unit 220 and the inverse transformation of the inverse transformation unit 225 shown may be implemented by inversely configuring the quantization of the quantization unit 215 and the transformation of the transformation unit 210 , and by combining the basic filtering units 235 and 330 .
[0321] Next, interpolation that can be used in the image encoding device of the present invention will be briefly described below.
[0322] In order to perform more accurate prediction by performing block matching, interpolation is performed with a resolution of real units more accurate than integer units. For the above interpolation, a discrete cosine transform-based interpolation filter (DCT-IF) method can be used. For interpolation in high-efficiency video coding (HEVC), DCT-IF is used, for example, by generating a reference picture with 1 / 2 and 1 / 4 unit pixels between integers, and generating a prediction block by referring to it by performing block matching.
[0323] Table 1
[0324] Pixel location Filter coefficient (fi) 0 0,0,0,64,0,0,0,0 1 / 4 -1,4,-10,58,17,-5,1,0 1 / 2 -1,4,-11,40,40,-11,4,-1 3 / 4 0,1,-5,17,58,-10,4,-1
[0325] Table 2
[0326] Pixel location Filter coefficient (fi) 0 0,64,0,0 1 / 8 -2,58,10,-2 1 / 4 -4,54,16,-2 3 / 8 -6,46,28,-4 1 / 2 -4,36,36,-4 5 / 8 -4,28,46,-6 3 / 4 -2,16,54,-4 7 / 8 -2,10,58,-2
[0327] Tables 1 and 2 show filter coefficients used in the luminance component and the chrominance component, respectively, an 8-tap DCT-IF filter is used for the luminance component, and a 4-tap DCT-IF filter is used for the chrominance component. For the chrominance component, different filters may be applied according to the color format. In the case of 4:2:0 of YCbCr, the filter shown in Table 2 may be applied. In the case of 4:4:4, the filter shown in Table 1 or other filters may be applied instead of the filter in Table 2. In the case of 4:2:2, the horizontal 1D 4-tap filter of Table 2 and the vertical 1D 8-tap filter of Table 1 may be applied.
[0328] Fig.15 1 is a view for illustrating a process of applying a one-dimensional horizontal filter to pixels existing at positions (assumed to be x) at a, b, and c of an image in an image decoding method according to an embodiment of the present invention.
[0329] like Fig.15 As shown, for sub-pixels at positions a, b, and c (assuming x) between a first pixel (G) and a second pixel (H) adjacent to the first pixel, a 1D horizontal filter may be applied. This may be expressed as the following formula.
[0330] x=(f1*E+f2*F+f3*G+f4*H+f5*I+f6*J+32) / 64
[0331] Then, for sub-pixels located at positions d, h, and n (assuming y), a 1D vertical filter can be applied. This can be expressed as the following formula.
[0332] y=(f1*A+f2*C+f3*G+f4*M+f5*R+f6*T+32) / 64
[0333] For sub-pixels f, g, i, j, k, p, q, and r located at the center, a 2D separable filter can be applied. For example, sub-pixel e can be interpolated by interpolating sub-pixel a and pixels located in the vertical direction and by using the interpolated pixels. Then, when sub-pixel a is interpolated between pixel G and pixel H, 1D horizontal filtering is performed. Then, the value of sub-pixel e can be obtained by performing 1D vertical filtering on the sub-pixel obtained by 1D horizontal filtering. For chrominance signals, similar operations can be performed.
[0334] The above description is part of the interpolation. In addition to the DCT-IF filter, other filters can also be used. The filter type and the number of taps can be applied differently to each real number unit. For example, for a 1 / 2 unit, an 8-tap Kalman filter can be used, for a 1 / 4 unit, a 6-tap Wiener filter can be used, and for a 1 / 8 unit, a 2-tap linear filter can be used. The filter coefficients can be encoded by using fixed coefficients such as DCT-IF or by calculating the filter coefficients. As described above, a single interpolation filter can be used for the picture, or different interpolation filters can be used for each area according to the image characteristics. Alternatively, at least two reference pictures to which multiple interpolation filters are applied can be generated, and one of the reference pictures can be selected.
[0335] Different filters may be applied according to encoding information such as reference picture type, temporal layer, state of reference picture (e.g., whether it is the current picture), etc. The above information may be set and transmitted in sequence units, picture units, slice units, etc.
[0336] Before describing in detail an improved method of motion estimation, motion compensation, and motion prediction that can be used in the image encoding method according to the present embodiment, its basic meaning is defined as follows.
[0337] When predicting the motion of the current block to be encoded, motion estimation refers to the process of dividing an image frame into small blocks and estimating which blocks of a frame (reference frame) that has been encoded before or after in time the small block has been moved. In other words, motion estimation can be the process of finding the block closest to the target block when encoding the current block to be compressed. Block-based motion estimation can be the process of estimating to which position a video object or image processing unit block (macroblock) has moved in time.
[0338] Motion compensation means a process of generating a prediction block of a current block based on motion information (motion vector, reference picture index) of an optimal prediction block during motion estimation by predicting the current image using at least a portion of an area of a previously encoded reference image in order to encode the current image. In other words, motion compensation may be the difference between a current block to be encoded and a reference block found to be closest to the current block, and involves a process of generating an error block.
[0339] Motion prediction means a process of finding a motion vector when encoding is performed for motion compensation. Skipping, temporal prediction, spatial prediction, etc. are used for motion prediction. Skipping means omitting encoding of a corresponding image block when the motion of an image is constant so that the size of the image block predicted in the image encoding device is 0, or when the residual is small enough to be negligible. Temporal prediction may be mainly used for inter-frame prediction, and spatial prediction or temporal prediction may be mainly used for intra-frame prediction.
[0340] The information related to the inter prediction may include information for distinguishing the reference picture list direction (one direction (L0 and L1), two directions), the index of the reference picture in the reference picture list, the motion vector, etc. Since the time correlation is used, when the same or similar characteristics as the motion vector of the current block are used, the motion information can be efficiently encoded.
[0341] Fig.16 is an example diagram of a current block and neighbor blocks according to a comparative example. Fig.17 is an example diagram of a current block and neighbor blocks according to another comparative example. Fig.18 is an example diagram of a current block and neighbor blocks according to yet another comparative example.
[0342] like Fig.16 As shown, Fig.16 An example of a current block (current PU) and neighboring blocks (A, B, C, D, and E) is shown. A block spatially adjacent to the current block, in other words, at least one of the upper left block (A), upper side block (B), upper right block (C), left side block (D), and lower left block (E) may be selected as a candidate, and its information may be used.
[0343] In addition, a block F existing within a reference block (co-located PU) at a position corresponding to the current block in the selected reference picture or its neighbor block (G) may be selected as a candidate, and information thereof may be used. Fig.16 Two candidates are shown, and in addition to the block (F) located in the center that is identical to the current block, at least one of the upper left, upper side, upper right, left side, lower left, lower side, lower right, and right side blocks may be selected as a candidate. Here, the position of the blocks included in the candidate group may be determined according to the picture type, the size of the current block, and the correlation with the spatially adjacent candidate blocks. In addition, the selected reference picture may mean a picture having a distance of 1 from a picture existing before or after the current picture.
[0344] In addition, if Fig.17 and Fig.18 As shown, it is possible to use information of a specific block according to the size and position of a current block (current PU). A prediction unit (PU) may be referred to as a basic unit for performing prediction of a prediction unit.
[0345] In the image encoding method of the present embodiment, the area of spatially and temporally adjacent blocks can be extended as follows.
[0346] Fig.19 is an example diagram of a current block and neighbor blocks that can be used in an image encoding method according to an embodiment of the present invention.
[0347] refer to Fig.19 According to the image encoding method of this embodiment, when selecting prediction candidates for motion information, the candidate blocks of motion information can be effectively selected by using the motion information of the coding blocks adjacent to the current block, the blocks located at the coordinates equal to or adjacent to the current block in the reference image (co-located blocks), or the blocks that are not adjacent but located in the same space. Here, the blocks that are not adjacent but located in the same space can be used as candidates for the blocks including at least one other block located between the current block and the corresponding block determined based on the information of the encoding mode, the reference picture index, the preset coordinates, etc.
[0348] In order to reduce the amount of motion information, an image encoding method may encode the motion information by using motion vector copy (MVC) or motion vector prediction (MVP).
[0349] Motion vector replication is a method of deriving motion vectors, and actually uses motion vector prediction values (motion vector predictors) instead of transmitting motion vector differences or reference picture indices. The above features are similar to the merge mode in HEVC. The merge mode combines motion vectors of spatially and temporally adjacent blocks. However, in the motion vector replication of the present embodiment, motion vectors are used for blocks that are not spatially and temporally adjacent (e.g., blocks H, I, and J). Taking the above features into account, motion vector replication has a different concept from the merge mode, and may have a higher concept than the merge mode. In addition, the goal of motion vector prediction is to generate the minimum motion vector difference when generating a prediction block.
[0350] Motion vector replication, such as Fig.19 As shown, reference directions, reference picture indices, and motion vector prediction values can be derived from blocks in various candidate groups. Motion vectors can be transmitted by selecting n candidate blocks from each candidate group, and the index information of the best block among the n candidate blocks is encoded. Here, when n is 1, motion vector replication can use the motion information of the block with the highest priority according to a preset criterion. In addition, when n is 2, considering the encoding cost, motion vector replication can determine which is the best for each candidate group.
[0351] The above-mentioned information of n can be fixed in the encoder or decoder, or can be transmitted in sequence units, picture units, slice units, etc. In one embodiment, whether n is 1 or 2 can be determined based on the information of neighboring blocks. Here, when the motion vector obtained from the block of the candidate group is similar to the preset criterion by comparing therebetween, it is determined that the generation of at least two candidate groups is meaningless, and the motion vector prediction value can be generated as a fixed value by calculating the average value or median (mvA, mvB, ..., mvJ). When it is determined that the generation of at least two candidate groups is not meaningless, the optimal motion vector prediction value can be generated by using at least two candidate groups.
[0352] Motion vector prediction can be performed similar to motion vector replication. However, the criteria of priority can be the same or different from motion vector replication. In addition, the number n of candidate blocks can be the same or different. Information about it can be fixed in the encoder or decoder or both, or can be transmitted in sequence units, picture units, slice units, etc.
[0353] Whether to set the reference candidate group may be determined based on information such as the current picture type, temporal layer information (temporal id), etc. The information thereon may be used by being fixed, or the information thereon may be transmitted in sequence units, picture units, or slice units. Here, setting whether to reference the candidate group may include, for example, using spatially adjacent blocks (A, B, C, D, and E), using spatially adjacent blocks (A, B, C, D, and E) and temporally adjacent blocks (H, I, and J), or using spatially adjacent blocks (A, B, C, D, and E) and spatially separated blocks (F and G).
[0354] Next, the setting of “block matching can be applied to the current picture” will be described below.
[0355] First, the case of an I picture is described, for example, since a spatially adjacent block may have the highest priority, and the remaining blocks may be set as a candidate group. In one embodiment, the availability check of the reference block may be performed in the order of E→D→C→B→A→H→I→J. The availability may be determined by the coding mode of the candidate block, motion information, the position of the candidate block, etc. The motion information may include a motion vector, a reference direction, a reference picture index, etc.
[0356] Since the current picture is an I picture, there is motion information when the coding mode is intra prediction (hereinafter, INTER) of the present embodiment. Therefore, according to the priority, it is possible to check whether the coding mode is INTER. For example, when n is 3 and E is encoded by INTER, E is excluded from the candidate group and the next D is checked. When D is encoded by INTER, block matching is performed in the current picture so that D has motion information, and D is added to the candidate group based on the motion information. After that, two n remain. Then, the image encoding device can check the priority again. When 3 candidates are filled as above, the generation of the candidate group is stopped.
[0357] The availability is not only determined by the coding mode, but also by the picture boundary, slice boundary, tile boundary, etc. In the case of the boundary, the availability is checked as unavailable. In addition, when it is determined that the block is equal to or similar to the candidate that has been filled, the corresponding block is excluded from the candidate, and the reference pixel configuration unit checks the availability of the next candidate.
[0358] Here, INTER is different from conventional inter prediction (inter). In other words, the INTER mode of the present embodiment can use the structure of inter prediction (inter) and generate a prediction block in the current picture, so the INTER mode is different from the inter prediction that generates a prediction block in the reference picture. In other words, in the encoding mode of the present embodiment, a method of performing block matching in the current picture can be applied by being classified into the INTER mode and intra (intra, equivalent to conventional intra).
[0359] At the same time, the availability of the first filled candidate may not be checked. The availability may be checked according to the situation. In other words, when the motion vector range of the candidate block is equal to or greater than a predetermined criterion or exceeds a picture boundary, a slice boundary, a tile boundary, etc., the availability may not be checked.
[0360] Hereinafter, when encoding is performed, the availability check of the reference blocks in the order of E→D→C→B→A→H→I→J described will be described with reference to various embodiments (refer to Fig.19 ), and obtain a predetermined number (eg, n=3) of candidate blocks.
[0361] <Example 1>
[0362] E(uncoded)→D→C→B→A→H→I→J: exclude E
[0363] E (remove) → D (interframe) → C → B → A → H → I → J: including D
[0364] E (remove) → D (include) → C (interframe) → B → A → H → I → J: including C
[0365] (After checking for similarity, block C is determined to be different)
[0366] E(remove)→D(include)→C(include)→B(intra)→A→H→I→J: exclude BE(remove)→D(include)→C(include)→B(remove)→A(inter)→H→I→J: include A (after similarity check, block A is checked as different)
[0367] As above, three blocks (D, C, and A) may be selected.
[0368] <Example 2>
[0369] E(interframe)→D→C→B→A→H→I→J: including E
[0370] E (inclusive) → D (interframe) → C → B → A → H → I → J: including D
[0371] (After the similarity check, block D is checked to be identical or dissimilar to the current block).
[0372] E (inclusive) → D (inclusive) → C (interframe) → B → A → H → I → J: exclude C
[0373] (After the similarity check, block C is checked to be identical or similar).
[0374] E (included) → D (included) → C (removed) → B (boundary) → A → H → I → J: exclude BE (included) → D (included) → C (removed) → B (removed) → A (intra-frame) → H → I → J: exclude A
[0375] E(include)→D(include)→C(remove)→B(remove)→A(remove)→H(intra)→I→J: including H (after similarity check, block H is checked as different.)
[0376] As above, three blocks (E, D, and H) can be selected.
[0377] <Example 3>
[0378] E(intraframe)→D→C→B→A→H→I→J: exclude E
[0379] E (remove) → D (intra-frame) → C → B → A → H → I → J: exclude D
[0380] E (remove) → D (remove) → C (intra-frame) → B → A → H → I → J: exclude C
[0381] E(remove)→D(remove)→C(remove)→B(boundary)→A→H→I→J: exclude BE(remove)→D(remove)→C(remove)→B(remove)→A(intra-frame)→H→I→J: exclude A
[0382] E (remove) → D (remove) → C (remove) → B (remove) → A (remove) → H (intra-frame) → I → J: exclude H
[0383] E (remove) → D (remove) → C (remove) → B (remove) → A (remove) → H (remove) → I (interframe) → J: including I
[0384] E(v)→D(remove)→C(remove)→B(remove)→A(remove)→H(remove)→I(include)→J(inter): including J (after similarity check, block J is checked as different.)
[0385] E (remove) → D (remove) → C (remove) → B (remove) → A (remove) → H (remove) → I (include) → J (include): including (-a, 0)
[0386] In the above availability check of eight reference blocks, two reference blocks are selected as candidate blocks. In order to fill several (n=3) candidate blocks, a preset fixed coordinate (-a, 0) may be added as a candidate.
[0387] <Example 4>
[0388] E(intraframe)→D→C→B→A→H→I→J: exclude E
[0389] E (remove) → D (intra-frame) → C → B → A → H → I → J: exclude D
[0390] E (remove) → D (remove) → C (intra-frame) → B → A → H → I → J: exclude C
[0391] E(remove)→D(remove)→C(remove)→B(boundary)→A→H→I→J: exclude BE(remove)→D(remove)→C(remove)→B(remove)→A(interframe)→H→I→J: include A
[0392] E (remove) → D (remove) → C (remove) → B (remove) → A (include) → H (intra) → I → J: exclude H
[0393] E (remove) → D (remove) → C (remove) → B (remove) → A (include) → H (remove) → I (intra) → J: exclude I
[0394] E (remove) → D (remove) → C (remove) → B (remove) → A (include) → H (remove) → I (remove) → J (intra): exclude J
[0395] E (remove) → D (remove) → C (remove) → B (remove) → A (include) → H (remove) → I (remove) → J (remove) → (-a, 0): including (-a, 0)
[0396] E (remove) → D (remove) → C (remove) → B (remove) → A (include) → H (remove) → I (remove) → J (remove) → (-a, 0) → (0, -b): including (0, -b)
[0397] Since one reference block (A) is selected as a candidate block in the above availability check of eight reference blocks, in order to fill a preset number (n=3) of candidate blocks, preset two fixed coordinates (-a, 0) and (0, -b) may be added as candidates.
[0398] Next, Example 3 and Example 4 will be described.
[0399] In Embodiment 3, when n candidates are not found in the preset candidate group, a fixed candidate with fixed coordinates can be added as a motion vector candidate. In other words, in the inter-frame prediction mode of HEVC, a zero vector can be added. However, when block matching is performed in the current picture as in the present invention, since the (0, 0) vector does not exist in the current picture, a preset virtual zero vector (-a, 0) is added as in Embodiment 3.
[0400] In addition, Embodiment 4 shows that when several motion vector candidates are not filled by performing availability check after adding a virtual zero vector, another preset virtual zero vector (0, -b) can be added to the candidate group. Here, considering the length size or width size of the block, a of (-a, 0) and b of (0, -b) can be set differently.
[0401] In addition, in the above description, the reference block is excluded by performing the availability check, and it can be set not to perform the availability check. Here, since there is no need to transmit information about which one must be transmitted, when n is assumed to be 2, it can be automatically determined whether to use the average of the two values or to use one of the two values according to the preset setting. Here, the index information of which of the two values is used can be omitted. In other words, the index of the motion vector prediction value can be omitted.
[0402] In the above description, the availability check is performed by configuring the entire candidate group in a single group. The present invention is not limited to this, and candidates can be selected by dividing the candidate group into at least two groups and performing an availability check on each candidate group. Here, assuming that the number of candidate blocks (n) is 2, the maximum number of candidates for each group can be equal, or can be set independently. In one embodiment, such as the following embodiment, one candidate for group 1_1 and one candidate for group 1_2 can be set. In addition, the priority of each group can be set independently. In addition, after performing an availability check on group 1_1 and group 1_2, when the criterion for the number of motion vector candidates is not met, in order to fill the number of motion vector candidates, group 2 can be applied. Group 2 may include fixed candidates with fixed coordinates.
[0403] For example, the order of availability checks or the priority order of group 1_1, group 1_2, and group 2 and the group distinction criteria are described in the following order.
[0404] Group 1_1 = {A, B, C, I, J}, (C→B→A→I→J), based on the upper block of the current block Group 1_2 = {D, E, H}, (E→D→H), based on the left and lower left blocks of the current block
[0405] Group 2 = {fixed candidate}, (-a, 0) → (0, -b), (-a, 0)
[0406] In summary, the image encoding device of this embodiment can spatially search for motion vector copy candidates or motion vector prediction candidates, and generate a candidate group in a preset fixed candidate order according to the searched candidate results.
[0407] Hereinafter, motion vector copying (MVC) and motion vector prediction (MVP) will be described separately because the involvement of the scaling process is different.
[0408] Motion vector prediction (MVP) will be described first below.
[0409] In the case of a P picture or a B picture
[0410] The description will be made by adding the temporal candidates (F and G) to the above candidates (A, B, C, D, E, F, G, H, I, and J). In this embodiment, it is assumed that the candidates are searched spatially, searched temporally, and searched by configuring a combination list, and fixed candidates are searched. The above operations are performed in state order.
[0411] First, the priorities of the candidates are determined, and availability checks are performed according to the priorities. Assume that the number of motion vector candidates (n) is set to 2, and their priorities are as indicated in parentheses.
[0412] For example, when searching spatially, it can be classified into the following groups.
[0413] Group 1_1 = {A, B, C, I, J}, (C → B → A → I → J)
[0414] Group 1_2 = {D, E, H}, (D → E → H)
[0415] In this example, group 1_1 of the two groups includes directly above, left and right blocks based on the current block, and group 1_2 includes directly adjacent left, non-directly adjacent left and left bottom blocks based on the current block.
[0416] As another example, candidate blocks of a motion vector may be spatially searched by being classified into three groups. The three groups may be classified as follows.
[0417] Group 1_1 = {A, B, C}, (C → B → A)
[0418] Group 1_2 = {D, E}, (D→E)
[0419] Group 1_3 = {H, I, J}, (J→I→H)
[0420] In this embodiment, among the three groups, group 1_1 includes directly upper adjacent, upper left, and upper right blocks based on the current block, group 1_2 includes directly adjacent left and directly adjacent lower left blocks based on the current block, and group 1_3 includes non-adjacent blocks having a distance of at least one block from the current block.
[0421] As another example, candidate blocks of motion vectors may be searched spatially by classifying using other methods. The three groups may be classified as follows.
[0422] Group 1_1 = {B}
[0423] Group 1_2 = {D}
[0424] Group 1_3 = {A, C, E}, (E→C→A)
[0425] In this example, among the three groups, group 1_1 includes blocks located in a vertical direction based on the current block, group 1_2 includes neighboring blocks located in a horizontal direction based on the current block, and group 1_3 includes remaining neighboring blocks based on the current block.
[0426] As described above, in a P picture or a B picture, since there are various kinds of reference information such as a reference direction, a reference picture, etc., a candidate group can be set according to the above information. A candidate block based on a current block having a reference picture different from the reference picture of the current block can be included in the candidate group. Alternatively, considering the temporal distance (counted picture, POC) between the reference picture of the current block and the reference picture of the candidate block, the vector of the corresponding block can be scaled and added to the candidate group. In addition, the scaled candidate group can be added according to which picture the reference picture of the current block is. In addition, when the temporal distance between the reference picture of the current block and the reference picture of the candidate block exceeds a predetermined distance, the scaled block is excluded from the candidate group. Alternatively, when the distance is equal to or less than a predetermined distance, the scaled block can be added to the candidate group.
[0427] The similarity check described above is the process of comparing and determining how similar the newly added motion vector is to the motion vectors already included in the prediction candidate set. According to the definition, it can be set to true when the x, y components match exactly, or it can be set to true when the components have a difference equal to or less than a predetermined threshold range.
[0428] Fig. 20 It is an example diagram for illustrating a situation in an image encoding method according to an embodiment of the present invention in which a reference picture of a current block is excluded from a candidate group when a temporal distance between a reference picture of a candidate block is equal to or greater than a predetermined distance, and is included in a candidate group after scaling is performed according to the distance when the temporal distance is less than a predetermined distance. Fig.21 is an exemplary diagram for illustrating a case in which, when a picture reference through a current block is a picture different from the current picture, a current picture is added to a prediction candidate group in an image encoding method according to an embodiment of the present invention. Fig. 22 is an exemplary diagram for illustrating a case in which, when a picture referenced by a current block is a current picture, a current picture is added to a prediction candidate group in an image encoding method according to an embodiment of the present invention.
[0429] In the image encoding method according to the present embodiment, motion information prediction candidates can be selected from reference pictures (rf1, rf2, rf3, rf4 and rf5) encoded before a predetermined time (t-1, t-2, t-3, t-4, and t-5) based on the time (t) of the current picture, and from candidate blocks within the current picture encoded before the current block of the current picture.
[0430] refer to Fig. 20 According to the image encoding method of this embodiment, the motion information prediction candidate of the current block can be selected from the reference pictures (rf1, rf2, rf3, rf4 and rf5) encoded temporally before the current picture (current (t)).
[0431] In one embodiment, when the current block (B_t) is encoded by referring to the second reference picture (rf1), for the upper left block (A) referring to the picture having a temporal distance equal to or greater than 3 based on the second reference picture (rf1), scaling is not allowed, so that the corresponding motion vector (mvA_x', mvA_y') is not included in the candidate group. However, since the blocks have a temporal distance less than 3 based on (t-2), the motion vectors of the remaining blocks can be included in the candidate group by being scaled. A subset of the motion vectors included in the candidate group is as follows.
[0432] MVS={(mvB_x', mvB_y'), (mvC_x', mvC_y'), (mvD_x', mvD_y'), (mvE_x', mvE_y')}
[0433] In other words, the criterion for whether to include a motion vector in a candidate group may be based on a picture referenced by the current block, or may be based on the current picture. Therefore, whether to add a motion vector to a candidate group may be determined by the distance between pictures, and information thereof, for example, a threshold value or information of a picture as a criterion may be transmitted by a sequence, a picture, a slice, etc. Fig. 20 In , the solid lines show the motion vectors of the respective blocks, and the dotted lines show the blocks appropriately scaled for the picture referenced by the current block. The above features can be expressed as the following formula 1.
[0434] [Formula 1]
[0435] Scaling factor = (POCcurr - POCcurr_ref) / (POCcan - POCcan_ref)
[0436] In Formula 1, POCcurr means the POC of the current picture, POCcurr_ref means the POC of the picture referenced by the current block, POCcan means the POC of the candidate block, and POCcan_ref means the POC of the picture referenced by the candidate block. Here, the POC of the candidate block may represent a POC identical to the current picture when the candidate block is a spatially adjacent candidate block, and may represent a POC different from the current picture when the candidate block is a block (co-located block) corresponding to a temporally identical position within a frame.
[0437] For example, Fig.21As shown, when the current block (B_t) and the reference picture are different so that scaling is performed, the motion vector obtained by scaling the corresponding motion vector, for example, {(mvA_x', mvA_y'), (mvC_x', mvC_y'), (mvD_x', mvD_y'), and (mvE_x', mvE_y')} may be included in the candidate group. It is unlikely that the motion vector is included in the candidate group even when the reference picture index is different as in HEVC. The above feature indicates that there may be a restriction depending on which reference picture is used as the reference picture.
[0438] In addition, if Fig.21 As shown, when the picture referenced by the current block (B_t) is another picture (t-1, t-2, t-3) different from the current picture (t), that is, in the case of normal inter-frame prediction (Inter), and the current picture is a picture different from the current picture as the current block (B_t), except when the picture of the candidate group indicates the current picture (t), the corresponding motion vectors {(mvA_x, mvA_y), (mvC_x, mvC_y), (mvD_x, mvD_y), (mvE_x, mvE_y)} may be included in the candidate group.
[0439] In addition, if Fig. 22 As shown, when the picture referenced by the current block (B_t) is the current picture (t), the motion vector {(mvB_x, vB_y), (mvC_x, mvC_y)} within the current picture (t) can be added to the candidate group. In other words, when the same reference picture is indicated, that is, when the same picture is indicated by comparing the reference picture index, the corresponding motion vector can be added to the candidate group. This is because, since block matching is performed in the current picture (t), motion information used in conventional inter-frame prediction (inter) with reference to another picture is not required. In addition, in addition to performing block matching in the current picture (t), when the reference picture indexes are different from each other as in H.264 / AVC, it can be achieved by excluding them from the candidate group.
[0440] This embodiment shows an example condition where the reference picture indicates the current picture. However, it can be extended to the case where the reference picture does not indicate the current picture. For example, a setting such as "blocks using pictures farther away than the picture indicated by the current block are excluded" can be used. Fig.16 In , some blocks (A, B, and E) can be excluded from the candidate set because the temporal distance is farther than t-2.
[0441] In this embodiment, although the current block and the reference picture are different, the reference picture can be included in the candidate group by performing scaling. Assume that the number of motion vector candidates (n) is 2. The following group 1_1 and group 1_2 show the case where the reference picture of the current block is not the current picture.
[0442] Group 1_1 includes adjacent blocks or non-adjacent but upper side blocks (A, B, C, I, and J) based on the current block, and the priority of availability check of the blocks is as described below.
[0443] Priority order: C→B→A→I→J→C_s→B_s→A_s→I_s→J_s
[0444] Group 1_2 includes adjacent or non-adjacent but left and lower left blocks (D, E, H) based on the current block, and the priority of the availability check of the blocks is as follows.
[0445] Priority order: D→E→H→D_s→E_s→H_s
[0446] In the above order, blocks having symbols with a subscript (s) refer to blocks that have been scaled.
[0447] Meanwhile, in a modified example of the present embodiment, first, an availability check is performed on some blocks (A, B, C, I, and J) of group 1_1 or some blocks (D, E, and H) of group 1_2 in a predetermined priority order. Then, when it is determined that the block is unavailable even if the block is inter-coded but its reference picture is different from the reference picture of the current block, the block is scaled according to the distance between the reference picture of the candidate block and the current picture (t), and then the availability check can be additionally performed on the scaled blocks (C_s, B_s, A_s, I_s, J_s, D_s, E_s, and H_s) according to the predetermined priority. Based on the above availability check, the image encoding device can determine two candidates, one candidate from each group, and use the motion vector of the optimal candidate block of the candidate group as the motion vector prediction value of the current block.
[0448] Meanwhile, when not even a single candidate block is obtained from a single group, for example, group 1_1, two candidate blocks may be obtained from another group 1_2. In other words, when two candidates are not filled from spatially adjacent blocks, candidate blocks may be obtained from a temporal candidate group.
[0449] In one embodiment, when the reference picture of the current block is the current picture, a block (co-located block) located at an equivalent position in a temporally adjacent reference picture and a block obtained by scaling the above blocks may be used. The above blocks (G, F, G_s, and F_s) may have the following priority order (G→F→G_s→F_s).
[0450] Similar to processing spatially adjacent candidate blocks, an availability check is performed according to the priority. When the distance between the picture of the candidate block and the reference picture of the candidate block is different from the distance between the current picture and the reference picture of the current block, an availability check can be performed on the candidate block that has been scaled. When the reference picture of the current block is the current picture, the above process is not performed.
[0451] Combination List
[0452] When motion information of each reference picture appearing in the reference picture list (L0 and L1) exists, the reference picture list (L0 and L1) is obtained by performing bidirectional prediction on the current block. The availability check is performed according to the preset candidate group priority. Here, the candidates encoded by bidirectional prediction are checked first.
[0453] When the respective reference pictures are different, scaling is performed. Among the previously obtained candidate blocks, when searching is performed spatially and temporally, when a block obtained by bidirectional prediction is added to the candidate group, and when the number of blocks does not exceed the maximum number of candidates, a block encoded by unidirectional prediction is added to the preliminary candidate group. Candidates for bidirectional prediction can be obtained by combining the above candidates.
[0454] [Table 3]
[0455]
[0456]
[0457] (a)
[0458] Candidate index L0 L1 0 <![CDATA[mvA1,ref1]]> <![CDATA[mvA2,ref0 <!-- 37 -->]]> 1 <![CDATA[mvB1',ref1]]> <![CDATA[mvB2',ref0]]> 2 mvC',ref1 mvD,ref0 3 mvC',ref1 mvE',ref0 4
[0459] (b)
[0460] First, assume that the motion information of the bidirectional prediction of the current block is referenced from reference picture 1 in L0 and reference picture 0 in L1. In Table 3(a), the motion information of the block included as the first candidate is (mvA1, ref1) and (mvA2, ref0), and the motion information of the block included as the second candidate is (mvB1', ref1) and (mvB2', ref0). Here, the apostrophe symbol (') means a scaled vector. After completing the spatial and temporal search, when the number of candidates is 2, here, assuming that n is 5, the blocks unidirectionally predicted in the previous process can be added to the preliminary candidates according to the preset priority.
[0461] In Table 3(a), candidates are not populated as many as the maximum number of candidates, so new candidates can be added by combining unidirectional candidates obtained by performing scaling using the remaining motion vectors of mvC, mvD, and mvE.
[0462] In Table 3(b), each motion information of a block obtained by unidirectional prediction can be scaled according to the reference picture of the current block. Here, an example of configuring a new combination by using a unidirectional candidate has been described. However, a new combination can be configured by using the motion information of each bidirectional reference picture (L0 and L1) that has been added. Here, the above configuration is not performed for unidirectional prediction. In addition, when the reference picture of the current block is the current picture, the above configuration is not performed.
[0463] Fixed candidates
[0464] When a candidate block having a maximum of n candidates (in this embodiment, it is assumed that n is 2) is not configured by the above process, a fixed candidate having a preset fixed coordinate can be added. Fixed candidates having fixed coordinates such as (0, 0), (-a, 0), (-2*a, 0), and (0, -b) can be used, and the number of fixed candidates can be set according to the maximum number of candidates.
[0465] The fixed coordinates may be set as above, or the fixed candidate may be set by calculating the average, weighted average, median of at least two motion vectors included in the candidate group by the above process until now. When n is 5, and three candidates {(mvA_x, mvA_y), (mvB_x, mvB_y), (mvC_x, mvC_y)} are added until now, the remaining two candidates may be added by the fixed candidate according to the priority, and the fixed candidate may be selected from the fixed candidate group including the fixed candidates with the preset priority. The fixed candidate group may include fixed candidates such as {(mvA_x+mvB_x) / 2, (mvA_y+mvB_y) / 2), ((mvA_x+mvB_x+mvC_x) / 3, (mvA_y+mvB_y+mvC_y) / 3), (median(mvA_x, mvB_x, mvC_x), median(mvA_y, mvB_y, mvC_y)), etc.
[0466] In addition, fixed candidates may be set differently according to the reference picture of the current block. For example, when the current picture is a reference picture, the fixed candidates may be set to (-a, 0), (0, -b), (-2*a, 0), etc. When the current picture is not a reference picture, the fixed candidates may be set to (0, 0), (-a, 0), (average (mvA_x, ...), average (mvA_y, ...)), etc. Information thereon may also be preset in an encoder or a decoder, or information thereon may be transmitted in sequence units, picture units, slice units, etc.
[0467] Hereinafter, an embodiment to which fixed candidates are applied will be described by way of example. In the following embodiment, it is assumed that n is 3.
[0468] <Example 1: n=3>
[0469] When the reference picture of the current block is the current picture, and the reference picture of the current block (B_t) is the reference picture (rf1 and rf0). The numbers 0, 1, 2, 3, 4, and 5 representing the reference picture and rf0, rf1, rf2, rf3, rf4, and rf5 of other pictures are applied to the present embodiment as examples, and no meaning is specified.
[0470] (Spatial search: E→D→A→B→C→E_s→D_s→A_s→B_s→C_s)
[0471] E (inter frame, rf0) → D → A → B → C → E_s → D_s → A_s → B_s → C_s: exclude E (because the reference picture of the current block is the current picture)
[0472] E (remove) → D (interframe, rf2) → A → B → C → E_s → D_s → A_s → B_s → C_s: exclude D
[0473] E (removal) → D (removal) → A (inter-frame, rf1) → B → C → E_s → D_s → A_s → B_s → C_s: including A
[0474] E (removed) → D (removed) → A (included) → B (inter-frame, rf1) → C → E_s → D_s → A_s → B_s → C_s: including B (determined to be different after similarity check)
[0475] E (remove) → D (remove) → A (include) → B (include) → C (intra) → E_s → D_s → A_s → B_s → C_s: exclude C
[0476] E (removed) → D (removed) → A (included) → B (included) → C (removed) → E_s → D_s → A_s → B_s → C_s: Excluding E_s (since the reference picture of the current block is the current picture, no scaling is required.)
[0477] E (removed) → D (removed) → A (included) → B (included) → C (removed) → E_s (removed) → D_s → A_s → B_s → C_s: including D_s (after checking, it is determined that the similarity is different. In addition, since the reference picture is different, it is excluded in D. Here, since it has been scaled, it does not need to be compared with the reference picture.)
[0478] <Example 2: n=3>
[0479] When the reference picture of the current block is the current picture, and the reference picture of the current block is a single reference picture (rf0).
[0480] (Spatial search: E→D→A→B→C→E_s→D_s→A_s→B_s→C_s)
[0481] E(interframe, rf1)→D→A→B→C→E_s→D_s→A_s→B_s→C_s: exclude E
[0482] E (removal) → D (intra-frame) A → B → C → E_s → D_s → A_s → B_s → C_s: exclude DE (removal) → D (removal) → A (inter-frame, rf0) → B → C → E_s → D_s → A_s → B_s → C_s: include A
[0483] E (remove) → D (remove) → A (include) → B (inter-frame, rf0) → C → E_s → D_s → A_s → B_s → C_s: including B (according to the result of similarity check)
[0484] E (remove) → D (remove) → A (include) → B (include) → C (intra) → E_s → D_s → A_s → B_s → C_s: exclude C
[0485] E (remove) → D (remove) → A (include) → B (include) → C (remove) → E_s → D_s → A_s → B_s → C_s: exclude C_s from E_s
[0486] Since three candidates have not been populated, fixed candidates may be applied, or additional candidate groups such as H, I, J, etc. may be checked.
[0487] <Example 3: n=3>
[0488] When the reference picture of the current block is not the current picture, the reference picture of the current block is a single reference picture (rf1), and a condition is added to exclude reference pictures having a distance equal to or greater than 2 from the reference picture of the current block from the candidate group. Here, when the number x of (tx) of the reference pictures (rfx-1) means the distance from the current picture (t), scaling may not be supported from (t-2).
[0489] Spatial search: E→D→A→B→C→E_s→D_s→A_s→B_s→C_s
[0490] E(intra)→D→A→B→C→E_s→D_s→A_s→B_s→C_s: exclude E
[0491] E (remove) → D (interframe, rf2) → A → B → C → E_s → D_s → A_s → B_s → C_s: exclude D
[0492] E (removal) → D (removal) → A (inter-frame, rf1) → B → C → E_s → D_s → A_s → B_s → C_s: including A
[0493] E (remove) → D (remove) → A (include) → B (interframe, rf3) → C → E_s → D_s → A_s → B_s → C_s: exclude B
[0494] E (remove) → D (remove) → A (include) → B (remove) → C (interframe, rf2) → E_s → D_s → A_s → B_s → C_s: exclude C
[0495] E (remove) → D (remove) → A (include) → B (remove) → C (remove) → E_s (intra-frame) → D_s → A_s → B_s → C_s: exclude E_s
[0496] E(remove) → D(remove) → A(include) → B(remove) → C(remove) → E_s(remove) → D_s → A_s → B_s → C_s: includes D_s (similarity check, using scaled motion vectors)
[0497] E (remove) → D (remove) → A (include) → B (remove) → C (remove) → E_s (remove) → D_s (include) → A_s → B_s → C_s: exclude A_s (already included in A)
[0498] E (remove) → D (remove) → A (include) → B (remove) → C (remove) → E_s (remove) → D_s (include) → A_s (remove) → B_s → C_s: exclude B_s (because the number of reference pictures is equal to or greater than 2)
[0499] E(removed) → D(removed) → A(included) → B(removed) → C(removed) → E_s(removed) → D_s(included) → A_s(removed) → B_s(removed) → C_s: including C_s (similarity check, using scaled motion vectors)
[0500] <Example 4: n=3>
[0501] When the reference picture of the current block is not the current picture, and the current picture is a specific reference picture (rf1).
[0502] (Spatial search: Select two motion vector candidates through spatial search)
[0503] (Time search: G→F→G_s→F_s)
[0504] G(intra)→F→G_s→F_s: exclude G
[0505] G(remove)→F(inter, rf3)→G_s→F_s: exclude F (here, assume that the scaling factor is not 1. In other words, the distance between the current block of the current picture and the reference picture is different from the distance between the picture of the co-located block and the reference picture of the corresponding block).
[0506] G(remove) → F(remove) → G_s → F_s: exclude G_s (because G is intra-frame)
[0507] G(remove)→F(remove)→G_s(remove)→F_s: including F_s (after checking, the similarity is determined to be different. Here, since the motion vector is a scaled vector, there is no need to compare with the reference picture)
[0508] Hereinafter, motion vector copy (MVC) will be described in detail.
[0509] Description of P picture or B picture
[0510] In this embodiment, it is assumed that temporal candidates (F and G) are also included. The candidate group includes A, B, C, D, E, F, G, H, I, and J. Although the search order is not determined, it is assumed here that MVC candidates are searched spatially, searched temporally, and searched by configuring a combination list, and then fixed candidates are added.
[0511] In other words, in the above section, similar to the above embodiment, the search is performed in an arbitrarily defined order rather than by using a preset order. The priority order is determined and the availability check is performed according to the priority. Assuming n is 5, the priority is as indicated in the brackets.
[0512] In the following description, the parts different from the described motion vector prediction (MVP) will be described. The above description can be continued to describe MVP, but the part described by scaling is excluded. For spatial candidates, availability check can be performed while scaling is omitted. However, similar to MVC, reference picture type, distance to the reference picture of the current picture or current block, etc. can be excluded from the candidate group.
[0513] When a combination list exists, candidates for bidirectional prediction can be configured by combining the candidates added so far as shown in Table 4 below.
[0514] Table 4
[0515] Candidate index L0 L1 0 mvA,ref0 1 mvB,ref1 2 mvC,ref0 3 4
[0516] (a)
[0517]
[0518]
[0519] (b)
[0520] As shown in Table 4(a), a new candidate can be added to the motion vector candidate group by combining the candidate using reference list L0 and the candidate using reference list L1. When the preset number of motion vectors is not filled, as shown in Table 4(b), a new candidate can be added by combining the candidate next to L0 and the candidate using L1.
[0521] The above-mentioned motion vector candidate selection method will be described below as an example. Assume that n is 3 in this example.
[0522] <Example 5: n=3>
[0523] When the reference picture of the current block is not the current picture, the reference picture ref of the current block is 1. (Spatial search: B→D→C→E→A→J→I→H)
[0524] B (interframe, rf3) → D → C → E → A → J → I → H: including B
[0525] B (inclusive) → D (inter-frame, rf1) → C → E → A → J → I → H: including D (similarity check)
[0526] B (inclusive) → D (inclusive) → C (intra) → E → A → J → I → H: exclude C
[0527] B (include) → D (include) → C (remove) → E (interframe, rf0) → A → J → I → H: exclude E
[0528] B (include) → D (include) → C (remove) → E (remove) → A (inter-frame, rf3) → J → I → H: including A (similarity check)
[0529] <Example 6: n=3>
[0530] An example in which the reference picture of the current block is the current picture is as follows.
[0531] (Spatial search: B→D→C→E→A→J→I→H)
[0532] B (inter-frame, rf1) → D → C → E → A → J → I → H: exclude B
[0533] B (remove) → D (intra-frame) → C → E → A → J → I → H: exclude D
[0534] B (remove) → D (remove) → C (interframe, rf0) → E → A → J → I → H: including C
[0535] B (remove) → D (remove) → C (include) → E (inter-frame, rf0) → A → J → I → H: including E (similarity check)
[0536] B (remove) → D (remove) → C (include) → E (include) → A (interframe, rf2) → J → I → H: exclude A
[0537] B (remove) → D (remove) → C (include) → E (include) → A (remove) → J (inter-frame, rf0) → I → H: including J (similarity check)
[0538] The encoding may be performed according to a mode such as MVP, MVC, etc., which finds the optimal candidate for motion information as above.
[0539] In the case of skip mode, encoding can be performed by using MVC. In other words, information for the optimal motion vector candidate can be encoded after processing the skip flag. When the number of candidates is 1, the above processing can be omitted. Encoding can be performed by performing transformation and quantization on the residual component which is the difference between the current block and the predicted block instead of encoding the motion vector difference separately.
[0540] In the case of not skip mode, it is checked whether the motion information is processed by performing MVC according to the priority order. When the motion information is processed, the information of the candidate group of the optimal motion vector can be encoded. When the motion information is not processed by performing MVP, the motion information can be processed by performing MVP. When performing MVP, the information of the optimal motion vector candidate can be encoded. Here, when the number of candidates is 1, the processing of the motion information can be omitted. In addition, information about the difference with the motion vector, reference direction, reference picture index, etc. of the current block can be encoded, the residual component can be obtained, and the encoding can be performed by performing transformation and quantization.
[0541] Hereinafter, codec-related parts, such as entropy and post-processing filtering, will be omitted to avoid redundancy with the above description.
[0542] The motion vector selection method used in the above image encoding method will be briefly described below.
[0543] Fig.23 is a flowchart of an image encoding method according to another embodiment of the present invention.
[0544] refer to Fig.23, the method of selecting a motion vector in the image encoding method according to the present embodiment basically includes: step S231 of configuring a spatial motion vector candidate; step S232 of determining whether a reference picture of a current block (blk) exists in the current picture; and step S233 of adding a spatial motion vector candidate of the current picture when a reference picture of the current block (blk) exists in the current picture (yes, Y). Step S233 involves checking additional spatial motion vector candidates and adding the checked spatial motion vector candidates.
[0545] In other words, in this embodiment, the method of adding a motion vector to a candidate group includes: configuring motion vectors of spatially adjacent blocks in the candidate group by performing MVC or MVP, and adding motion information of blocks existing in the same current picture when the reference picture of the current block is the current picture. Fig.19 The I, J, and H blocks of the above may correspond to the above case. The above block may be a block that is not directly adjacent to the current block but is most recently encoded by INTER. The above block means that the candidate group is configured by blocks located spatially different from blocks obtained by MVC, MVP, etc. Therefore, the above block can be represented by using "added" or "attached".
[0546] When the reference picture of the current block is not the current picture, a candidate group may be set from blocks of temporally adjacent pictures. Then, by configuring a combination list, a candidate for bidirectional prediction may be configured by a combination of candidates added so far. In addition, a fixed candidate may be included according to the reference picture of the current picture. In addition, when the current picture is a reference picture, a fixed candidate with preset fixed coordinates may be added to the candidate group of the motion vector. Otherwise, a fixed candidate with (0, 0) coordinates may be included in the candidate group of the motion vector.
[0547] Meanwhile, in step S232 , when the reference picture of the current block (blk) does not exist within the current picture (No, N), step S234 of searching for temporal motion vector candidates and adding the searched temporal motion vector candidates may be performed.
[0548] Then, in step S235, the image encoding apparatus performing the motion vector selection method may configure a combination list candidate including the motion vector candidates configured in at least one of steps S231, S233, and S234.
[0549] Then, the image encoding apparatus checks whether the current picture is a reference picture in step S236, and adds fixed candidates with preset fixed coordinates in step S237 when the current picture is a reference picture and the number of combined list candidates is less than the number of preset motion vector candidates.
[0550] Meanwhile, in step S236 , when the current picture is not a reference picture, the image encoding apparatus adds a fixed candidate having coordinates of (0, 0) to the candidate group in step S238 .
[0551] When the reference pixel is configured by the above motion vector selection method, it can be performed Fig. 9 In one embodiment, when the reference pixel is configured by the motion vector selection method and motion estimation is performed based on the above, it is possible to perform Fig.15 interpolation.
[0552] Fig.24 It is a diagram for illustrating when the motion vector accuracy changes in block units. Fig.18 1 is a view for illustrating a case where the accuracy of a motion vector of a block is determined according to the interpolation accuracy of a reference picture in the image encoding method according to an embodiment of the present invention.
[0553] exist Fig.24 In , it is assumed that the quantity indicated in the parentheses for each frame, that is, (2), is the interpolated precision depth information. In other words, Fig.24 In , the interpolation accuracy of each picture including the current picture (t) and the reference pictures (t-1, t-2 and t-3) is constant, but the motion vector accuracy is adaptively determined in units of blocks. Fig.25 In , the motion vector accuracy is determined based on the interpolation accuracy of the reference picture. Fig.25 In the depth information of the interpolation accuracy
[0554] The motion vector prediction value may be generated for each of the above two cases based on the above-mentioned “motion information prediction candidate selection” or “motion vector candidate selection”.
[0555] In this embodiment, the maximum number of candidates for predicting a motion vector may be set to 3. Here, the candidate blocks may be limited to the left, upper, and upper right blocks of the current block. The candidate blocks may be set with spatially and temporally adjacent blocks as well as spatially non-adjacent blocks.
[0556] Fig.26 is a flowchart of an image encoding and decoding method using a motion vector difference according to an embodiment of the present invention.
[0557] Referring to 26 , the image decoding method according to the present embodiment may configure a motion information prediction candidate group by using a prediction unit, and calculate a differential value using a motion vector of a current block.
[0558] Describing in more detail, first, in step S262, the motion vectors of the blocks belonging to the candidate group may be changed according to the precision unit of the motion vector of the current block.
[0559] Then, the motion vector scaling step S264 may be performed according to the distance between the motion vector of the current block and the reference picture, that is, the distance between the current picture and the reference picture and the distance between the picture of the block belonging to the candidate group and the reference picture of the corresponding block.
[0560] Then, in step S266, the prediction unit may obtain a motion vector difference between the current block and the corresponding block based on the motion vector scaled in a single picture.
[0561] A method of obtaining a motion vector difference will be described in detail by using the above steps.
[0562] Figure 27 to Figure 32 is a view for illustrating a process of calculating a motion vector difference in various cases when interpolation accuracy is determined in units of blocks in the image encoding method according to an embodiment of the present invention.
[0563] When determining the interpolation accuracy in units of screens
[0564] like Fig. 27 As shown, it is assumed that three blocks adjacent to the current block are candidate blocks for encoding motion information. Here, it is assumed that the interpolation accuracy of the current picture (t) is an integer (Int), the interpolation accuracy of the first reference picture (t-1) is 1 / 4, the interpolation accuracy of the second reference picture (t-2) is 1 / 4, and the interpolation accuracy of the third reference picture (t-3) is 1 / 2.
[0565] exist Fig. 27 In the example, the block of (A1) can be represented as the block of (A2) according to each motion vector precision.
[0566] When a block indicating a reference picture different from the reference picture of the current block is set to be used as a candidate by performing scaling considering the candidates between the reference pictures, the block of (A2) may be scaled considering the distance between the reference picture of the candidate block of (A3) and the reference picture of the current block. Blocks of the same reference picture are not scaled. When scaling is performed, when the above distance is less than the distance between the current picture and the reference picture, at least one of rounding, rounding up, and rounding down may be selected and applied. Then, the motion vector of the candidate block may be adjusted considering the motion vector accuracy of the current block such as the block of (A4). In the present embodiment, it is assumed that the upper right block is selected as the optimal candidate.
[0567] exist Fig. 27In the example, based on the current block located at the lower center, since the accuracy of the upper right block is 1 / 2 unit, in order to change the unit to 1 / 4 unit, the unit is multiplied by 2. However, in an alternative case, for example, when the accuracy is adjusted from 1 / 8 unit to 1 / 4 unit, the best candidate (MVcan) can be selected from the candidate group, and the difference with the motion vector (MVx) of the current block can be calculated. The calculated difference can then be encoded. This can be expressed as the following formula.
[0568] MVD=MVx-MVcan→(2 / 4, 1 / 4)
[0569] In the above-described embodiment, scaling is performed first and then the accuracy is adjusted, but it is not limited thereto. The accuracy may be adjusted first and then scaling may be performed.
[0570] Assuming that the interpolation accuracy is determined in units of blocks, the reference picture of the current block and the reference picture of the candidate block are equal of.
[0571] like Fig.28 As shown, the three blocks adjacent to the current block are candidate blocks for encoding motion information. Here, it is assumed that all interpolation precisions of the reference picture are 1 / 8. The block of (B1) can be represented as the block of (B2) according to each motion vector precision. Since the reference picture of the current block and the reference picture of the candidate block are identical, scaling is omitted.
[0572] Then, the corresponding motion vector of the candidate block of (B2) may be adjusted to the block of (B3) according to the motion vector accuracy of the current block.
[0573] Then, the best candidate (MVcan) is selected from the candidate group, the difference with the motion vector (MVx) of the current block is calculated, and the calculated difference is encoded. Fig.28 In , based on the current block located at the lower center, it is assumed that its upper side block is selected as the optimal candidate.
[0574] MVD=MVx-MVcan→(1 / 2,-1 / 2)
[0575] In the case where the interpolation accuracy is determined in units of blocks, and the reference picture of the current block and the reference picture of the candidate block The picture is different.
[0576] In this embodiment, it is assumed that the interpolation precision of the reference pictures is the same. In one embodiment, the interpolation precision of the reference pictures may be 1 / 8.
[0577] refer to Fig.29 , the block of (C1) can be represented as the block of (C2) according to each motion vector accuracy. In addition, since the reference picture of the current block and the reference picture of the reference picture are different, the block of (C3) can be obtained by performing scaling.
[0578] Then, in block (C3), the motion vector of the candidate block may be adjusted to that of block (C4) in consideration of the motion vector accuracy of the current block. Then, the best candidate (MVcan) may be selected from the candidate group, the difference (MVD) between the motion vector (MVx) of the current block may be calculated, and the calculated difference (MVD) may be encoded. Fig.29 In , based on the current block located at the lower center, it is assumed that its upper right block is selected as the best candidate. This can be expressed as the following formula.
[0579] MVD=MVx-MVcan→(1 / 4, 3 / 4)
[0580] In the case where the interpolation accuracy is determined in units of pictures, and temporally located candidate blocks are included.
[0581] In this embodiment, it is assumed that the interpolation accuracy of the reference pictures is equal. In one embodiment, it is assumed that the interpolation accuracy of the reference pictures is 1 / 2. Fig.30 As shown, the block of (D1) can be represented as the block of (D2) according to each motion vector precision. In the case of using the block of the current picture as the reference picture, since it is not a normal motion prediction, the corresponding block can be excluded (invalidated) from the candidate group.
[0582] Then, since the reference picture of the current block and the reference picture of the candidate block are different, the block of (D3) can be obtained by performing scaling. Here, it is assumed that the picture of the co-located block is a specific reference picture (t-1).
[0583] Then, in the block of (D3), the motion vector of the candidate block can be adjusted to the block of (D4) in consideration of the motion vector accuracy of the current block located at a lower center. Then, the optimal candidate (MVcan) can be selected from the candidate group, the difference (MVD) with the motion vector (MVx) of the current block can be calculated, and the calculated difference (MVD) can be encoded. In this embodiment, it is assumed that the co-located block is selected as the optimal candidate. The above configuration can be expressed as the following formula.
[0584] MVD=MVx-MVcan→(2 / 4, 2 / 4)
[0585] In the case where the interpolation accuracy is determined in units of blocks, and the current block refers to the current picture.
[0586] In this embodiment, it is assumed that the interpolation precision of the reference pictures is the same. In one embodiment, the interpolation precision of the reference pictures may be 1 / 4.
[0587] like Fig.31As shown, the block of (E1) can be represented as the block of (E2) according to each motion vector accuracy. In the case where a block of the current picture located at a lower center is selected as a reference picture, since it is not a normal motion prediction and block matching is performed in the current picture, the upper side block performing the same process is selected as a candidate, and the remaining blocks are excluded (invalid) from the candidate group.
[0588] Since the reference picture of the current block and the reference picture of the candidate block are identical, no scaling is performed. Then, considering the motion vector accuracy of the current block, the motion vector of the candidate block can be adjusted to the block of (E3).
[0589] Then, the best candidate (MVcan) can be selected from the candidate group, the difference (MVD) with the motion vector (MVx) of the current block can be calculated, and the calculated difference (MVD) can be encoded. In this embodiment, it is assumed that the upper block is selected as the best candidate. The above configuration can be expressed as the following formula.
[0590] MVD=MVx-MVcan→(-5 / 2,-1 / 2)
[0591] Then, the image encoding and decoding method according to the present embodiment can use the information of the reference block adjacent to the current block to be encoded. Fig.32 As shown, the coding mode and information of the neighboring block can be used based on the current block.
[0592] exist Fig.32 In FIG. 1 , the upper block (E5) is a block when the current picture is an I picture, and encoding is performed by performing block matching in the current picture to generate a prediction block.
[0593] In the inter-frame prediction of the present embodiment, the method of the present embodiment is applied to the method based on conventional extrapolation. The inter-frame prediction of the present embodiment may be expressed as INTRA, and may include Inter when block matching is performed in the current picture.
[0594] Describing in more detail, in the case of Inter, information such as a motion vector, a reference picture, and the like can be used. When the coding mode of the neighbor block and the coding mode of the current block are equal, when the current block is encoded in the prescribed order, information of the corresponding blocks (E5, E6, and E7) can be used. In addition, for blocks that are not directly adjacent, when a block is added to a candidate block by checking the coding mode of a block in a block of (E5), and the coding mode is inter, a reference picture (ref) is additionally checked, and the current block can be encoded using a block whose ref is t. When the motion vector accuracy is determined according to the interpolation accuracy of the reference picture, for example, when the interpolation accuracy of the current picture is an integer unit, the motion vector of the Inter-encoded block can be expressed as an integer unit.
[0595] (E6) shows the case where the current picture is a P or B picture. Information of a reference block adjacent to the current block to be encoded can be used. When the motion vector accuracy is determined based on the interpolation accuracy of the reference picture, the motion vector accuracy of each block can be set based on the interpolation accuracy of each reference picture.
[0596] (E7) shows a case where the current picture is a P or B picture, and the current picture is included (used) as a reference picture. The motion vector accuracy of each block can be determined according to the interpolation accuracy of the reference block.
[0597] (E8) shows the case where the current block is encoded using the information of the co-located block. The motion vector accuracy of each block can be determined according to the interpolation accuracy of each reference picture.
[0598] Then, when encoding and decoding, encoding motion information can be performed. However, in the present embodiment, encoding is performed by adaptively determining the accuracy of the motion vector difference. In other words, when the interpolation accuracy of the reference picture is constant, the motion vector accuracy between blocks is constant. In other words, the motion vector accuracy can be determined according to the interpolation accuracy of the reference picture.
[0599] Figure 33 to Figure 36 is a view for illustrating a process of expressing the accuracy of a motion vector difference in an image encoding method according to an embodiment of the present invention.
[0600] In this embodiment, the motion vector of the best candidate block among various candidate blocks is (c, d), and the motion vector of the current block is (c, d). For ease of description, it is assumed that the reference pictures are identical.
[0601] like Fig.33 As shown in (F1), it is assumed that the interpolation accuracy of the reference picture is 1 / 4. The motion vector difference (ca, db) is calculated and transmitted to the decoder. To this end, in this embodiment, the block of (F1) can be represented as a block of (F2) according to each motion vector accuracy. Since the reference pictures are identical and the motion vector accuracy of each block is identical, the motion vector of the block of (F2) can be directly used for encoding.
[0602] In other words, as shown in (F2), when (c, d) as the motion vector of the current block is (21 / 4, 10 / 4) and the optimal candidate block is the left block, (a, b) as the motion vector of the left block is (13 / 4, 6 / 4), so the difference (cd, db) therebetween becomes (8 / 4, 4 / 4). Since the motion vector accuracy is 1 / 4 unit, the accuracy can be processed with binary index (bin index) using 8 and 4 as shown in Table 5 below. However, when the accuracy information of the motion vector difference is transmitted, the accuracy can be processed with a short bit index.
[0603] Table 5
[0604]
[0605] When transmitting information that the motion vector difference of the current block has the precision of integer units, a shorter bin index than the traditional one can be used by using 2 and 1, whereby the traditional bin index 8 and 4 are changed to integer units. When various binarization methods are used, for example, when a unary binarization method is used, in order to transmit the above differential value (8 / 4, 4 / 4), 1111111110+11110 bits must be transmitted. However, when transmitting the differential value corresponding to (2 / 1, 1 / 1), that is, 110+10, and the precision information of the bit, the coding efficiency can be improved by using shorter bits.
[0606] When the precision of the motion vector difference is used as Fig.34 When the method shown is expressed, in the above case, information of 11 (integer) and information of the difference value (2 / 1, 1 / 1) can be transmitted.
[0607] In the above embodiment, the motion vector difference precision is applied for both the x and y components. However, the motion vector difference precision may be applied independently.
[0608] exist Fig.35 In , (G1) can be expressed as (G2) according to each motion vector precision. When the reference pictures are identical and the corresponding motion vector precisions of the blocks are identical, (G2) can be used to encode the motion vector.
[0609] When it is assumed that the optimal candidate block is the left block of the current block, the difference value thereof can be expressed as (8 / 4, 5 / 4). When the precision of the motion vector difference is applied to x and y of the optimal candidate block as expressed above, x can be expressed as 2 / 1 in integer units, and y can be expressed as 5 / 4 in 1 / 4 units.
[0610] When the precision of the motion vector difference is expressed as Fig.34In the following tree in the table, information about 11 (integer) for x, information about 0 (1 / 4) for y, and information about 2 / 1 and 5 / 4 for each difference value can be transmitted (refer to Fig.36 ).
[0611] In other words, in the above case, if Fig.36 As shown, the range of the differential value precision can be set from the maximum precision (1 / 4) to the minimum precision (integer) as a candidate group. However, the present invention is not limited to the above configuration, and the present invention can configure various candidate groups. It can be configured to include precision of at least two units from the maximum precision (e.g., 1 / 8) to the minimum precision (integer).
[0612] Fig.37 An example of a reference structure of a random access mode in an image encoding and decoding method according to an embodiment of the present invention is shown.
[0613] refer to Fig.37 , first, when encoding is performed sequentially from a picture with a lower temporal identifier (temporalID), an I picture and a P picture with a lower ID are encoded respectively, and then a B(2) picture with a higher ID is encoded. Assuming that a picture with an ID equal to or higher than itself is not referenced, a picture with a lower ID may be selected as a reference picture.
[0614] First, based on Fig.37 Described, the interpolation accuracy of the picture used as a reference picture can be used by being constant (for example, any one of an integer, 1 / 2, 1 / 4, 1 / 8). Alternatively, according to an embodiment, the interpolation accuracy of the picture used as a reference picture can be set differently. For example, in the case where the I picture has an ID of 0, when the I picture is used as a reference picture of a certain picture, by calculating the distance between the pictures, the I picture can have distances of 1, 4, and 8 (in the case of pictures B4, B2, and P1). In the case of a P picture, the P picture can have distances of 4, 2, and 1 (B2, B6, and B8). In the case where B2 has an ID of 1, B2 can have distances of 2, 1, 1, and 2 (B3, B5, B7, and B6). In the case where B3 has an ID of 2, it can have distances of 1 and 1 (B4 and B5).
[0615] The interpolation accuracy of each reference picture can be determined based on the average distance from the reference picture. In other words, due to the low motion difference, accurate interpolation is required for closer reference pictures. In other words, for B3 or B6 with a short average distance from the reference picture, interpolation is performed by increasing the interpolation accuracy (e.g., 1 / 8). When the distance from the reference picture is equal to or greater than this value, the interpolation accuracy can be reduced. For example, the interpolation accuracy can be reduced to 1 / 4 (refer to Case 1 of Table 6 below). Alternatively, when accurate interpolation accuracy is required for reference pictures that are far away, the opposite of the above example can be applied.
[0616] In addition, more accurate interpolation can be performed on a picture that is referenced multiple times. For example, interpolation can be performed by increasing the interpolation accuracy of B2. Alternatively, other interpolation accuracies can be applied to pictures that are referenced less frequently (refer to case 2 of Table 6 below).
[0617] In addition, the interpolation accuracy may be applied differently according to the temporal layer. For example, more accurate interpolation may be performed for a picture with ID 0, and the accuracy may be reduced for pictures with other IDs or vice versa (refer to case 3 of Table 6 below).
[0618] Table 6
[0619] show 0 1 2 3 4 5 6 7 8 coding 0 4 3 5 2 7 6 8 1 Case 1 1 / 4 - 1 / 8 - 1 / 4 - 1 / 8 - 1 / 4 Case 2 1 / 4 - 1 / 4 - 1 / 8 - 1 / 4 - 1 / 4 Case 3 1 / 8 - 1 / 4 - 1 / 4 - 1 / 4 - 1 / 8
[0620] Therefore, inter-frame prediction can be performed by setting the interpolation accuracy in picture units. Information thereon may be predefined in an encoder or a decoder, or may be transmitted in sequence units, picture units, and the like.
[0621] Fig.38 is a view for illustrating that a single picture can have at least two interpolation precisons in the image encoding method according to an embodiment of the present invention.
[0622] refer to Fig.38 , Fig.38 An example of performing block matching in the current picture is shown. In other words, Fig.38 The case where I0, P1 and B2 are each referenced to themselves is shown.
[0623] In this embodiment, a picture may not refer to a picture having an ID higher than itself, and may refer to a picture having an ID equal to or lower than itself. Fig.38 It is not shown in the figure, but may be included in the case when the current picture with temporalID of 2 or 3 is selected as the reference picture. The above case means that the current picture can be selected according to the temporalID. In addition to the temporal identifier (temporalID), the reference picture can be determined according to the picture type.
[0624] Fig.38 The situation and Fig.37 The situation is similar, but Fig.38 The case of shows the interpolation accuracy when the current picture is selected as the reference picture when encoding the current picture, and the interpolation accuracy of the general case can be equal. In addition, Fig.38 It is shown that different interpolation precisities can be determined when the current picture is selected as the reference picture. In other words, a single picture can have at least two interpolation precisities.
[0625] So far, a method of determining the interpolation accuracy of a reference picture, and then determining the accuracy of a motion vector has been described. In the following description, a method of determining the accuracy of a motion vector without considering the interpolation accuracy will be described.
[0626] The interpolation precision of the picture may be constant (e.g., 1 / 4), but the motion vector precision of the block may be set in block units of the picture to be encoded. The preset motion vector precision is used as a basic setting. In the present embodiment, it may mean that at least two candidates may be supported. For example, when the interpolation precision of the reference picture is 1 / 4, the motion vector of the block referring to the reference picture must generally be represented as a 1 / 4 unit. However, in the present embodiment, the motion vector may be represented as a usual 1 / 4 unit, additionally as a 1 / 2 unit, or additionally as an integer unit.
[0627] Table 7
[0628]
[0629] Table 7 shows the matching of the precision of the motion vector according to the predetermined constant. For example, to represent 1, the precision may correspond to 1 in integer units, 2 in 1 / 2 units, 4 in 1 / 4 units. To represent 2, the precision may correspond to 2 in integer units, 4 in 1 / 2 units, and 8 in 1 / 4 units.
[0630] When the precision is increased by 2 times based on the current precision of the block referring to the specific picture, the integer 1 becomes 1 / 2, and 1 / 2 becomes 1 / 4, so in order to express the above situation, the number becomes doubled. When the above situation is expressed by using various binarization methods such as a unary binarization method, a truncated Rice binarization method, a k-order exponential Golomb binarization method, etc., the number of bits expressing it increases as the precision increases.
[0631] When the interpolation precision increases, for example, from 1 / 2 to 1 / 4, but the motion vector is found in a lower precision unit (for example, integer or 1 / 2), the encoding efficiency may decrease due to the increase in the number of bits. In order to prevent this, in the present embodiment, the information of the precision unit from which the motion vector is found is expressed in block units, and the information thereon is used in the encoder or the decoder or both. Here, encoding and decoding can be performed efficiently.
[0632] For example, when a motion vector is encoded by using a unary binarization method, and the x component of the vector is 8 / 4 and the y component of the vector is 4 / 4, 1111111110+11110 of binarization bits are required to represent it. However, according to the present embodiment, the above binarization bits can be represented by 1 / 2 units (4 / 2, 2 / 2). Here, the above binarization can be represented as 11110+110. In addition, the above binarization can be represented by integer units (2, 1). Here, the above binarization can be represented as 110+10.
[0633] In addition, information about which precision unit is used when encoding the motion vector (for example, 0 in 1 / 4 unit, 10 in 1 / 2 unit, and 11 in integer unit) is transmitted, and in the above case, the above binarization can be expressed as 11 (precision unit) + 110 + 10, so fewer bits can be used than 1111111110 + 11110 of 1 / 4 unit.
[0634] Therefore, the encoding result of the motion vector of the present embodiment can be used to transmit / receive information representing the precision of the motion vector of the corresponding block in integer units, and thus the transmitted / received encoding bits can be reduced.
[0635] In addition, as an embodiment, when the x and y components of the motion vector are 3 / 4 and 1 / 4, since the components are not represented by integers or 1 / 2 units, the components can be represented as 1110+10 of 1 / 4 units. Here, encoding and decoding according to the present embodiment can be performed by transmitting information representing the motion vector accuracy of the corresponding block in 1 / 4 units. Therefore, the interpolation accuracy in picture units can be constant, while the accuracy of the motion vector is adaptively determined in block units. The 2 maximum accuracy of the motion vector in block units is equivalent to the interpolation accuracy of the reference picture.
[0636] In addition, the precision group supported in this embodiment can be variably configured. For example, assuming that the interpolation precision of the reference picture is 1 / 8, the precision of the above integer, 1 / 2, 1 / 4 and 1 / 8 units can be used. Alternatively, at least two precisions such as (1 / 2, 1 / 4, 1 / 8), (1 / 4, 1 / 8), (integer, 1 / 2, 1 / 8), (1 / 4, 1 / 8), etc. can be configured and used.
[0637] The precision group can be set as a picture unit, a slice unit, etc., or can be determined in consideration of the encoding cost. Information thereof can be transmitted in sequence units, picture units, slice units, etc. In addition, the decoder can select one of the transmitted sets and adaptively use the set when determining the motion vector precision.
[0638] In addition, in the present embodiment, the index for selecting the precision can be expressed according to the number of configured candidate groups by using a fixed-length binarization method, a unary binarization method, etc. Short bits are allocated to the unit with the highest frequency of occurrence or a high chance of occurrence based on the precision of the reference picture, otherwise, long bits can be allocated.
[0639] For example, assuming that the 1 / 8 unit has the highest frequency of occurrence, 0 may be assigned to the 1 / 8 unit, 10 to the 1 / 4 unit, 110 to the 1 / 2 unit, and 111 to the integer unit. In other words, the shortest bit is assigned to the 1 / 8 unit and the longest bit is assigned to the integer unit.
[0640] In addition, when the 1 / 4 unit has the highest frequency of occurrence, the shortest bit is allocated to the 1 / 4 unit. In addition, according to an embodiment, a fixed length can be allocated regardless of the frequency of occurrence. For example, a fixed length can be allocated based on the statistical frequency of occurrence, for example, 00 can be allocated to an integer unit, 01 to a 1 / 2 unit, 10 to a 1 / 4 unit, and 11 to a 1 / 8 unit.
[0641] In addition, in the present embodiment, when the reference picture is a specific picture, that is, when the reference picture is the current picture, a high priority may be assigned to a specific precision. When the reference picture is the current picture, a high priority may be assigned to an integer unit. The shortest bit may be assigned to a basic 1 / 8 unit, and the next short bit may be assigned to an integer unit.
[0642] In other words, 0 to 1 / 8 unit, 10 to integer unit, 110 to 1 / 4 unit, and 11 to 1 / 2 unit may be allocated. In addition, according to an embodiment, the shortest bit may be allocated to an integer unit. For example, 0 to an integer unit, 10 to 1 / 8 unit, 110 to 1 / 4 unit, and 111 to 1 / 2 unit may be allocated.
[0643] Since, in the case of an image including screen contents, a motion search is rarely performed on a corresponding portion in real units, but a motion search is mostly performed on a natural image portion in real units, bits may be allocated as above.
[0644] In other words, the image encoding and decoding method of this embodiment can be effectively applied when the image is configured to include a combination of two areas in which a portion of the image is a screen content area such as a computer-captured image and the other portion of the image is a natural image area.
[0645] In addition, in this embodiment, information of a reference block for predicting motion information of the current block may be used. Here, the reference block used may use information of a first spatially adjacent block. The first block may include at least one of the upper left, upper side, upper right, and lower left blocks based on the current block.
[0646] In addition, information of a reference block (co-located block) existing at a position corresponding to the current block in the selected reference picture may be used. In addition to the co-located block, the reference block may also select at least one of the upper left, upper side, upper right, left side, lower left, lower side, lower right and right blocks as a candidate based on a reference block (center block) located at an equivalent position of the current block. Here, the position included in the candidate group may be determined based on encoding-related parameters such as picture type, size of the current block, mode, motion vector, reference direction, etc., correlation between spatially adjacent candidate blocks, etc. Then, the selected reference picture may mean a picture having a distance of 1 from a picture existing before or after the current picture.
[0647] In addition, in the present embodiment, information of a block that is not adjacent to the current block but located at the same space may be used. The block may include a block including at least one block between the determined current block and the corresponding block based on information of the coding mode, the reference picture index, the preset coordinates, etc. as a candidate. The preset coordinates may be set to a predetermined distance of the length and width of the current block having the upper left coordinates of the current block.
[0648] For example, it is assumed that the accuracy of the motion vector of the current block is represented by referring to the left block and the upper block, and a truncated unary binarization method is used to represent the accuracy of the motion vector. The bit configuration can be 0-10-110-111. When the accuracy of the motion vector of the left block is determined in 1 / 2 units, and the accuracy of the motion vector of the upper block is determined in 1 / 2 units, the motion vector of the current block can be represented in 1 / 2 units with a large change. It is determined, and the accuracy-related information can be binarized by allocating the shortest bit between various units. The various units may include integers, 1 / 2, 1 / 4, 1 / 8 units supported in the device, and can be basically set to 1 / 8 units.
[0649] As another embodiment, when the precision of the left block is 1 / 4 unit, and the precision of the upper side block is 1 / 8 unit, relatively short bits are allocated to 1 / 4 unit and 1 / 8 unit, and relatively long bits are allocated to other remaining units. Here, in addition, when including 1 / 8 unit as a basic unit, the shortest bit can be allocated to 1 / 8 unit. In addition, when not including basic units, such as when the precision of the left block is 1 / 4 unit and the precision of the upper side block is 1 / 2 unit, the shortest bit can be allocated to a fixed position so that the fixed position can be used first. For example, when the left block has the highest priority, the shortest bit can be allocated to a 1 / 4 unit.
[0650] In addition, in the present embodiment, the motion information of the candidate block may be used. In addition to the motion vector, the motion information may include a reference picture index, a reference direction, etc. In one embodiment, when the precision of the left block is a 1 / 4 unit and the precision of the upper block is an integer unit, and when the left block has t-1 as a reference picture, the upper block has t-2 as a reference picture, and the current block is t-2, the highest priority is assigned to the block having the same reference picture, and the shortest bit may be assigned to the integer unit of the precision of the motion vector as the corresponding block.
[0651] The above embodiments are briefly represented as 1) to 4) below.
[0652] 1)(1 / 2,1 / 2)1 / 2-1 / 8-1 / 4-integer
[0653] 2)(1 / 4, 1 / 8)1 / 8-1 / 4-1 / 2-integer
[0654] 3)(1 / 4, 1 / 2)1 / 4-1 / 2-1 / 8-integer
[0655] 4) (1 / 4, integer) integer-1 / 4-1 / 8-integer
[0656] In addition, in the present embodiment, the motion vector accuracy can be adaptively or constantly supported in block units according to information of the current picture type, etc., while the interpolation accuracy is fixed in picture units.
[0657] like Fig.38 As shown, when the current picture is a reference picture, the motion vector accuracy can be adaptively used. Alternatively, in order to find more accurate motion when the distance between reference pictures such as B4, B5, B7, B9 is short, the motion vector accuracy in blocks can be supported to the corresponding picture, and the constant motion vector accuracy can be supported to the remaining pictures.
[0658] In other words, when the temporal layer information (Temporal ID) is 3, one of integer, 1 / 2 and 1 / 4 units is selected for the corresponding ID and the selected unit is supported, and 1 / 4 units are used for the remaining IDs. Alternatively, in the case of a P picture having a distance farther from the reference picture, it is assumed that there are areas with precise motion and areas with imprecise motion, and the motion vector accuracy is adaptively determined in units of blocks. In the case of other pictures, constant motion vector accuracy can be supported.
[0659] In addition, according to the embodiment, at least two sets configured with different precisions can be supported. For example, at least two of the precision candidates such as (1 / 2, 1 / 4), (integer, 1 / 2, 1 / 4), (1 / 2, 1 / 4, 1 / 8), (integer, 1 / 4) can be used (refer to Table 8).
[0660] Table 8
[0661]
[0662] When setting the interpolation accuracy in picture units.
[0663] Fig.39 It shows that the current screen is Fig.38 A view of the reference picture list for an I picture in .
[0664] refer to Fig.39 , a prediction block can be generated by performing block matching in the current picture. The I picture is added to the reference picture (N). Expressed as I*(0) means that the current picture is selected as the reference picture when encoding the current picture.
[0665] When the interpolation precision is set in units of pictures, it corresponds to a case where precision up to integer units is allowed, and thus interpolation is not performed. The image encoding and decoding apparatus may refer to reference picture list 0 (L0).
[0666] In addition, except for the case where the current picture is selected as the reference picture based on the picture type, temporal identifier (temporalID), etc., for example, block matching is allowed when the current picture is an I picture, and block matching is not allowed for other pictures, the reference picture with the asterisk (*) symbol can be omitted.
[0667] Fig.40 It shows that the current screen is Fig.38 A diagram of a reference picture list for a P picture in FIG.
[0668] refer to Fig.40, a P picture may also be added to the reference picture. Expressed as P*(1) means that the P picture is selected as the current picture when the current picture is encoded. Here, I(0) has a different meaning from the above I*(0). In other words, in I*(0), since motion search is performed during encoding, motion search can be performed in a picture to which a post-processing filter such as a deblocking filter is applied or partially applied. However, I(0) is a picture to which a post-processing filter is applied after the encoding process is completed, so I(0) may be a picture with the same POC, but may be a different picture since filtering is applied to it.
[0669] When the previous I picture is selected as the reference picture, the interpolation accuracy of the corresponding picture becomes 1 / 8 units, so the optimal motion vector can be found by performing a motion search up to 1 / 8 units. When the current picture (P*(1)) is selected as the reference picture, the interpolation accuracy of the corresponding picture becomes integer units, so the motion search can be performed in integer units.
[0670] Fig.41 is a view showing a reference picture list when a current picture is B(2) in an image encoding and decoding method according to an embodiment of the present invention. Fig.42 is a view showing a reference picture list when a current picture is B(5) in an image encoding and decoding method according to an embodiment of the present invention.
[0671] refer to Fig.41 , when the current picture is B(2), an I picture (I(0)) and a P picture (P(1)) are added to the reference picture, and a motion search can be performed on the corresponding picture up to the precision unit according to the interpolation precision of each picture. In this embodiment, a motion vector can be searched with a precision of 1 / 4 unit precision in the P picture, and a motion vector can be searched with a precision of 1 / 8 unit in the I picture.
[0672] In addition, refer to Fig.42 When the current picture is B(5), neither I picture nor P picture is included in the reference picture, but the motion vector can be searched by using different accuracies according to the interpolation accuracies of multiple reference pictures {B(2), B(4), B(3)}.
[0673] Fig.43 1 is a view for illustrating a process of determining the precision of a motion vector of each block according to the interpolation precision of a reference picture in an image encoding and decoding method according to an embodiment of the present invention. Fig.44 is a view for illustrating a process of adaptively determining the precision of a motion vector of each block when the interpolation precision of each reference picture is constant in the image encoding / decoding method according to an embodiment of the present invention.
[0674] refer to Fig.43 , when the interpolation precision of the current picture is an integer (Int) unit, and the interpolation precisions of the three reference pictures (t-1, t-2, and t-3) are 1 / 4, 1 / 2, and 1 / 8 units, respectively, and when searching for the motion vector of a neighbor block of the current block referring to the reference picture, the motion vector of the neighbor block can be searched by using the respective precisions of the corresponding interpolation precisions of the corresponding reference pictures.
[0675] refer to Fig.44 , when the interpolation precision of the current picture and the reference picture is equal, such as a 1 / 4 unit, and when searching for the motion vector of the neighbor block of the current block referring to the reference picture, the motion vector of the neighbor block can be searched by using the precision corresponding to each of the predetermined interpolation precisions. However, according to an embodiment, although the precision of the motion vector is determined in block units, the interpolation precision of the reference picture can be set separately.
[0676] In the above embodiment, the accuracy of the motion vector is determined according to the interpolation accuracy of the reference picture. Alternatively, the accuracy of the motion vector is determined in block units, and the interpolation accuracy is fixed. However, the present invention is not limited to the above configuration. It can be configured by combining the interpolation accuracy of the reference picture and the accuracy of the motion vector determined in block units. Here, the maximum accuracy of the motion vector can be determined according to the accuracy of the reference picture.
[0677] According to the above-described embodiment, during inter-frame prediction, interpolation accuracy can be adaptively set in picture or block units. In addition, interpolation accuracy can be adaptively supported according to the time layer in the GOP structure, the average distance from the referenced picture, etc. In addition, when index information is encoded according to the accuracy, the amount of encoded information can be reduced, and the index information can be encoded by using various methods.
[0678] In the above embodiment, when the above motion vector accuracy is used when decoding the encoded image, the image encoding method may be used and replaced by the image decoding method. In addition, the image encoding and decoding method may be performed by at least one device for encoding and decoding, or an image processing device or an image encoding and decoding device configured with a configuration unit that performs functions corresponding to the above devices.
[0679] According to the above-mentioned embodiments, a coding and decoding method with high performance and efficiency can be provided, which can be generally used in international codecs such as MPEG-2, MPEG-4, H.264, etc. or other codecs, and media and image processing industries using these codecs. In addition, in the future, the method of the present invention can be applied to the current high-efficiency image coding method (HEVC), and the image processing field using standard codecs and intra-frame prediction such as H.264 / AVC.
[0680] Although the preferred embodiments of the present invention have been described for illustrative purposes, those skilled in the art will appreciate that various modifications, additions and substitutions are possible, without departing from the scope and spirit of the invention as disclosed in the accompanying claims.
Claims
1. An image decoding method, comprising: Obtaining, from a parameter set, a parameter setting flag indicating whether the current picture can be a reference picture of the current block; decoding precision information indicating precision of the motion vector; Determine the motion vector accuracy of the current block based on the motion vector accuracy information; Determining a motion vector of the current block from motion information of the current block based on the motion vector accuracy of the current block; as well as predicting the current block based on the motion vector of the current block, Wherein, whether the reference picture of the current block is the current picture is determined based on the parameter setting flag, wherein the motion vector accuracy of the current block is determined based on whether the determined reference picture of the current block is the current picture, Wherein, when the motion vector accuracy information of the current block indicates 0, in the case where the reference picture of the current block is the current picture, the motion vector accuracy of the current block is determined to be 1, The motion vector accuracy of the current block is determined to be less accurate than or equal to the maximum motion vector accuracy of the current block, and the maximum motion vector accuracy of the current block is determined based on the reference picture of the current block.
2. The image decoding method according to claim 1, wherein when the determined reference picture is not the current picture, the motion vector accuracy of the current block is determined in integer or decimal units based on the motion vector accuracy information, When the determined reference picture is the current picture, the precision of the motion vector of the current block is determined based on an integer.
3. The image decoding method according to claim 1, wherein the motion vector accuracy information is signaled in a unit higher than a unit of a block.
4. An image encoding method, comprising: Determine the motion vector accuracy of the current block; Determining, based on the motion vector accuracy, a motion vector of the current block from the motion information of the current block; Predicting the current block based on the motion vector of the current block; and encoding the motion vector precision information of the current block, wherein the motion vector precision information of the current block indicates the motion vector precision of the current block, The motion vector accuracy of the current block is determined based on whether the determined reference picture of the current block is the current picture, Wherein, when the motion vector accuracy information of the current block indicates 0, in the case where the reference picture of the current block is the current picture, the motion vector accuracy of the current block is determined to be 1, The motion vector accuracy of the current block is determined to be less accurate than or equal to the maximum motion vector accuracy of the current block, and the maximum motion vector accuracy of the current block is determined based on the reference picture of the current block.
5. A bitstream storage method, comprising: generating a bitstream, the bitstream comprising motion information of a current block and motion vector accuracy information indicating accuracy of a motion vector of the current block; as well as storing the bitstream, in, The motion vector accuracy of the current block is determined based on the motion vector accuracy information of the current block; The motion vector of the current block is determined from motion information of the current block based on the motion vector accuracy of the current block; and The prediction sample of the current block is derived based on the motion vector, The motion vector accuracy of the current block is determined based on whether the determined reference picture of the current block is a current picture, Wherein, when the motion vector accuracy information of the current block indicates 0, in the case where the reference picture of the current block is the current picture, the motion vector accuracy of the current block is determined to be 1, The motion vector accuracy of the current block is determined to be less accurate than or equal to the maximum motion vector accuracy of the current block, and the maximum motion vector accuracy of the current block is determined based on the reference picture of the current block.
Citation Information
Patent Citations
Adaptive motion resolution for video coding
US20110206125A1