Method and apparatus for encoding / decoding image
By deriveing motion information in the video decoder and adopting the template block prediction method, the problem of low video encoding efficiency in the prior art is solved, and higher prediction accuracy and coding efficiency are achieved.
Patent Information
- Application Number
- CN202510154897.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-03-29
- Filing Date
- 2019-03-29
- Publication Date
- 2025-05-02
AI Technical Summary
When existing video encoding technology processes moving images, the channel bandwidth cannot keep up with the rapid growth of multimedia data volume, resulting in insufficiency of encoding.
By deriveing motion information in the decoder, the prediction method of the template block is adopted to determine the reference area based on the pre-reconstructed area spatially adjacent to the current block, and inter/intra prediction is performed.
Improve prediction accuracy and encoding efficiency, and enhance video compression performance.
Smart Images

Figure CN119922301A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application date of March 29, 2019, application number 201980022972.1, and invention name “Method and device for encoding / decoding images”. Technical Field
[0002] The present disclosure relates to video encoding / decoding methods and devices. Background Art
[0003] Recently, the demand for multimedia data such as moving images in the Internet is growing rapidly. However, the development of channel bandwidth may not keep up with the rapid growth of multimedia data volume. Therefore, VCEG (Video Coding Experts Group) of ITU-T and Moving Picture Experts Group (MPEG) of ISO / IEC, which are international standardization organizations, have established the video compression standard HEVC (High Efficiency Video Coding) version 1 in February 2014.
[0004] HEVC defines methods such as intra-frame prediction, inter-frame prediction, transform, quantization, entropy coding, and in-loop filter. Summary of the invention
[0005] Technical issues The present disclosure relates to a video encoding / decoding method and device, and aims to provide an inter-frame / intra-frame prediction method and device.
[0006] The present disclosure relates to a video encoding / decoding method and device, and aims to provide a block partitioning method and device.
[0007] An object of the present disclosure is to provide a method and apparatus for deriving motion information in a decoder.
[0008] Technical Solution According to the video encoding / decoding method and device of the present disclosure, reference information specifying the position of a reference area for predicting a current block may be determined, a reference area as a pre-reconstruction area spatially adjacent to the current block may be determined based on the reference information, and the current block may be predicted based on the reference area.
[0009] In the video encoding / decoding method and apparatus according to the present disclosure, the reference region may be determined as one of a plurality of candidate regions, wherein the plurality of candidate regions may include at least one of a first candidate region, a second candidate region, and a third candidate region.
[0010] In the video encoding / decoding method and device according to the present disclosure, the first candidate area may include a top adjacent area of the current block, the second candidate area may include a left adjacent area of the current block, and the third candidate area may include a top adjacent area and a left adjacent area of the current block.
[0011] In the video encoding / decoding method and apparatus according to the present disclosure, the first candidate region may further include a partial region of an upper right neighboring region of the current block.
[0012] In the video encoding / decoding method and apparatus according to the present disclosure, the second candidate area may further include a partial area of a lower left neighboring area of the current block.
[0013] The video encoding / decoding method and apparatus according to the present disclosure may partition a current block into a plurality of sub-blocks based on encoding information of the current block, and may sequentially reconstruct the plurality of sub-blocks based on a predetermined priority.
[0014] In the video encoding / decoding method and apparatus according to the present disclosure, the encoding information may include first information indicating whether the current block is partitioned.
[0015] In the video encoding / decoding method and apparatus according to the present disclosure, the encoding information may further include second information indicating a partition direction of the current block.
[0016] In the video encoding / decoding method and apparatus according to the present disclosure, the current block may be partitioned into two in a horizontal direction or a vertical direction based on the second information.
[0017] The video encoding / decoding method and apparatus according to the present disclosure may derive a partition boundary point of a block by using a block partitioning method that uses motion boundary points in a reconstruction area around a current block.
[0018] According to the video encoding / decoding method and apparatus of the present disclosure, initial motion information of a current block may be determined, incremental motion information of the current block may be determined, and the initial motion information of the current block may be improved using the incremental motion information, and motion compensation may be performed on the current block using the improved motion information.
[0019] In the video encoding / decoding method and apparatus according to the present disclosure, determining the incremental motion information may include determining a search area for improving motion information, generating a SAD (Sum of Absolute Difference) list from the search area, and updating the incremental motion information based on SAD candidates of the SAD list.
[0020] In the video encoding / decoding method and apparatus according to the present disclosure, the SAD list may specify a SAD candidate at each search position in a search area.
[0021] In the video encoding / decoding method and apparatus according to the present disclosure, the search area may include an area extending N sample lines from a boundary of a reference block, wherein the reference block may be an area indicated by initial motion information of the current block.
[0022] In the video encoding / decoding method and apparatus according to the present disclosure, the SAD candidate may be determined as a SAD value between an L0 block and an L1 block, and the SAD value may be calculated based on some samples in the L0 block and the L1 block.
[0023] In the video encoding / decoding method and apparatus according to the present disclosure, the position of the L0 block may be determined based on the position of the L0 reference block of the current block and a predetermined offset, wherein the offset may include at least one of a non-directional offset and a directional offset.
[0024] In the video encoding / decoding method and apparatus according to the present disclosure, updating of incremental motion information is performed based on a comparison result between a reference SAD candidate and a predetermined threshold, wherein the reference SAD candidate may mean a SAD candidate corresponding to a non-directional offset.
[0025] In the video encoding / decoding method and apparatus according to the present disclosure, when the reference SAD candidate is greater than or equal to a threshold, a SAD candidate having a minimum value among the SAD candidates in the SAD list is identified, and the incremental motion information can be updated based on an offset corresponding to the identified SAD candidate.
[0026] In the video encoding / decoding method and apparatus according to the present disclosure, when the reference SAD candidate is less than a threshold value, the incremental motion information may be updated based on parameters calculated using all or some of the SAD candidates included in the SAD list.
[0027] In the video encoding / decoding method and apparatus according to the present disclosure, improvement of initial motion information may be performed on a sub-block basis in consideration of the size of a current block.
[0028] In the video encoding / decoding method and apparatus according to the present disclosure, improvement of initial motion information may be selectively performed based on at least one of the size of the current block, the distance between the current picture and the reference picture, the inter-prediction mode, the prediction direction, or the resolution of the motion information.
[0029] Technical Effects According to the present disclosure, the method and device can improve the accuracy of prediction via template block-based prediction and improve encoding efficiency.
[0030] According to the present disclosure, encoding efficiency may be improved by performing sub-block based prediction through adaptive block partitioning.
[0031] According to the present disclosure, the method and apparatus may use motion boundary points between different objects within a reconstruction region to more accurately detect partition boundary points of a block, thereby improving compression efficiency.
[0032] According to the present disclosure, the accuracy of inter-frame prediction can be increased by improving motion information. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a simplified flow chart of a video encoding device.
[0034] Figure 2 is a diagram for illustrating a block partition unit of a video encoding apparatus in detail.
[0035] Figure 3 is a diagram for illustrating a prediction unit of a video encoding apparatus in detail.
[0036] Figure 4 is a diagram for illustrating a motion estimator of a prediction unit in a video encoding apparatus.
[0037] Figure 5 is a flowchart illustrating a method of deriving candidate motion information in skip mode and merge mode.
[0038] Figure 6 is a flow chart illustrating a method of deriving candidate motion information in AMVP mode.
[0039] Figure 7 is a diagram showing positions of reconstructed blocks used to derive spatial / temporal candidates in candidate motion information in skip mode, merge mode, and AMVP mode.
[0040] Figure 8 is a diagram illustrating a method of deriving a temporal candidate among candidate motion information in skip mode, merge mode, and AMVP mode.
[0041] Fig. 9 is a diagram illustrating a method of combining bidirectional candidate modes in deriving candidate motion information in skip mode and merge mode.
[0042] Fig.10 is a diagram illustrating a motion estimation pattern used in a motion estimator in a video encoding apparatus or a DMVD motion estimator in a video encoding / decoding apparatus according to one embodiment.
[0043] Fig.11 is a flow chart illustrating a method of encoding partition information and prediction information.
[0044] Fig.12 is a schematic flow chart showing a video decoding device.
[0045] Fig.13 is a diagram for illustrating a block partition unit of a video decoding apparatus.
[0046] Fig.14 is a diagram for illustrating a prediction unit of a video decoding apparatus.
[0047] Fig.15is a flow chart illustrating a method of decoding partition information and prediction information.
[0048] Fig.16 is a diagram for illustrating a block partition unit of a video encoding apparatus according to an embodiment of the present disclosure.
[0049] Fig.17 is a diagram for illustrating a block partition unit of a video decoding device according to an embodiment of the present disclosure.
[0050] Fig.18 is a diagram for illustrating a prediction unit of a video encoding apparatus according to an embodiment of the present disclosure.
[0051] Fig.19 is a diagram for illustrating a prediction unit of a video decoding device according to an embodiment of the present disclosure.
[0052] Fig. 20 is a diagram for illustrating a DMVD motion estimator of a prediction unit in a video encoding / decoding apparatus according to an embodiment of the present disclosure.
[0053] Fig.21 is a diagram for illustrating a DMVD mode according to one embodiment of the present disclosure.
[0054] Fig. 22 is a diagram for illustrating a method of determining a template block in a reconstruction area in a DMVD mode according to an embodiment of the present disclosure.
[0055] Fig.23 2 is a diagram for illustrating a method of searching for a partitioned motion boundary point of a block according to an embodiment of the present disclosure.
[0056] Fig.24 Is used to show the search Fig. 22 Illustration of a fast algorithm for finding moving boundary points in .
[0057] Fig.25 is a diagram for illustrating a DMVD initial motion information detector in a prediction unit in a video encoding / decoding apparatus according to an embodiment of the present disclosure.
[0058] Fig.26 It is shown in Fig.23 Illustration of how the cost value for each line is calculated in the fast algorithm for .
[0059] Fig. 27 2 is a diagram for illustrating an example of a block divided into two or four based on one or two partition motion boundary points for the block according to one embodiment of the present disclosure.
[0060] Fig.28 It is shown based on Fig.23 and Fig.24 An illustration of a method for searching motion information using only a partial area adjacent to a current block in a reconstructed area as a template, with reference to partitioned motion boundary points of the reconstructed area.
[0061] Fig.29 It is shown based on Fig.23 and 24 An illustration of a method of searching motion information by using an entire area adjacent to a current block in a reconstructed area as a template based on partitioned motion boundary points.
[0062] Fig.30 is a flowchart illustrating a method of encoding partition information and prediction information according to one embodiment of the present disclosure.
[0063] Fig.31 is a flowchart illustrating a method of decoding partition information and prediction information according to one embodiment of the present disclosure.
[0064] Fig.32 An example of partitioning a coding block according to an embodiment of the present disclosure is shown.
[0065] Fig.33 is a diagram for illustrating a DMVD motion estimator of a prediction unit in a video encoding / decoding apparatus according to an embodiment of the present disclosure.
[0066] Fig.34 is a flow chart illustrating a method of determining a slope of a motion boundary line according to one embodiment of the present disclosure.
[0067] Fig.35 It shows confirmation Fig.34 An illustration of a method of partitioning a current block by looking at the slope of a motion boundary line in a motion boundary search area.
[0068] Fig.36 is a diagram illustrating a filtering method when a current block is partitioned along a straight line according to one embodiment of the present disclosure.
[0069] Fig.37 is a diagram illustrating a filtering method when a current block is partitioned along a diagonal line according to one embodiment of the present disclosure.
[0070] Fig.38 is a flowchart illustrating a method of encoding prediction information according to an embodiment of the present disclosure.
[0071] Fig.39 is a flowchart illustrating a method of decoding prediction information according to an embodiment of the present disclosure.
[0072] Fig.40A method for performing intra prediction in a prediction unit of a video encoding / decoding device in an embodiment to which the present disclosure is applied is shown.
[0073] Fig.41 The sub-block based intra prediction method in the embodiment to which the present disclosure is applied is shown.
[0074] Fig.42 A prediction method based on inter-component reference in an embodiment to which the present disclosure is applied is shown.
[0075] Fig.43 A method for determining incremental motion information in an embodiment to which the present disclosure is applied is shown. DETAILED DESCRIPTION
[0076] Public best practices According to the video encoding / decoding method and device of the present disclosure, reference information specifying the position of a reference area for predicting a current block may be determined, a reference area as a pre-reconstruction area spatially adjacent to the current block may be determined based on the reference information, and the current block may be predicted based on the reference area.
[0077] In the video encoding / decoding method and apparatus according to the present disclosure, the reference region may be determined as one of a plurality of candidate regions, wherein the plurality of candidate regions may include at least one of a first candidate region, a second candidate region, and a third candidate region.
[0078] In the video encoding / decoding method and device according to the present disclosure, the first candidate area may include a top adjacent area of the current block, the second candidate area may include a left adjacent area of the current block, and the third candidate area may include a top adjacent area and a left adjacent area of the current block.
[0079] In the video encoding / decoding method and apparatus according to the present disclosure, the first candidate region may further include a partial region of an upper right neighboring region of the current block.
[0080] In the video encoding / decoding method and apparatus according to the present disclosure, the second candidate area may further include a partial area of a lower left neighboring area of the current block.
[0081] The video encoding / decoding method and apparatus according to the present disclosure may partition a current block into a plurality of sub-blocks based on encoding information of the current block, and may sequentially reconstruct the plurality of sub-blocks based on a predetermined priority.
[0082] In the video encoding / decoding method and apparatus according to the present disclosure, the encoding information may include first information indicating whether the current block is partitioned.
[0083] In the video encoding / decoding method and apparatus according to the present disclosure, the encoding information may further include second information indicating a partition direction of the current block.
[0084] In the video encoding / decoding method and apparatus according to the present disclosure, the current block may be partitioned into two in a horizontal direction or a vertical direction based on the second information.
[0085] The video encoding / decoding method and apparatus according to the present disclosure may derive a partition boundary point of a block by using a block partitioning method that uses motion boundary points in a reconstruction area around a current block.
[0086] According to the video encoding / decoding method and apparatus of the present disclosure, initial motion information of a current block may be determined, incremental motion information of the current block may be determined, and the initial motion information of the current block may be improved using the incremental motion information, and motion compensation may be performed on the current block using the improved motion information.
[0087] In the video encoding / decoding method and apparatus according to the present disclosure, determining the incremental motion information may include determining a search area for improving motion information, generating a SAD (Sum of Absolute Difference) list from the search area, and updating the incremental motion information based on SAD candidates of the SAD list.
[0088] In the video encoding / decoding method and apparatus according to the present disclosure, the SAD list may specify a SAD candidate at each search position in a search area.
[0089] In the video encoding / decoding method and apparatus according to the present disclosure, the search area may include an area extending N sample lines from a boundary of a reference block, wherein the reference block may be an area indicated by initial motion information of the current block.
[0090] In the video encoding / decoding method and apparatus according to the present disclosure, the SAD candidate may be determined as a SAD value between an L0 block and an L1 block, and the SAD value may be calculated based on some samples in the L0 block and the L1 block.
[0091] In the video encoding / decoding method and apparatus according to the present disclosure, the position of the L0 block may be determined based on the position of the L0 reference block of the current block and a predetermined offset, wherein the offset may include at least one of a non-directional offset and a directional offset.
[0092] In the video encoding / decoding method and apparatus according to the present disclosure, updating of incremental motion information is performed based on a comparison result between a reference SAD candidate and a predetermined threshold, wherein the reference SAD candidate may mean a SAD candidate corresponding to a non-directional offset.
[0093] In the video encoding / decoding method and apparatus according to the present disclosure, when the reference SAD candidate is greater than or equal to a threshold, a SAD candidate having a minimum value among the SAD candidates in the SAD list is identified, and the incremental motion information can be updated based on an offset corresponding to the identified SAD candidate.
[0094] In the video encoding / decoding method and apparatus according to the present disclosure, when the reference SAD candidate is less than a threshold value, the incremental motion information may be updated based on parameters calculated using all or some of the SAD candidates included in the SAD list.
[0095] In the video encoding / decoding method and apparatus according to the present disclosure, improvement of initial motion information may be performed on a sub-block basis in consideration of the size of a current block.
[0096] In the video encoding / decoding method and apparatus according to the present disclosure, improvement of initial motion information may be selectively performed based on at least one of the size of the current block, the distance between the current picture and the reference picture, the inter prediction mode, the prediction direction, and the resolution of the motion information.
[0097] Disclosed Embodiments The embodiments of the present disclosure will be described in detail with reference to the accompanying drawings attached to the present disclosure, so that a person of ordinary skill in the technical field to which the present disclosure belongs can easily implement the present disclosure. However, the present disclosure can be implemented in various different forms and may not be limited to the exemplary embodiments shown herein. In addition, in the accompanying drawings, parts not related to the description are omitted in order to clearly illustrate the present disclosure. In the accompanying drawings, similar reference numerals are assigned to similar parts throughout the specification.
[0098] It will be understood that when an element is referred to as being “connected to” another element, it can be directly on the other element, directly connected to the other element, or one or more intervening elements may be present.
[0099] Throughout the specification, the meaning that a certain element includes another element does not exclude other components, but means that other components may be further included unless there is a specific description contrary thereto.
[0100] It will be understood that, although the terms “first,” “second,” etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used to distinguish one element from another.
[0101] In addition, in the embodiments of the device and method described in the present disclosure, some components of the device or some steps of the method may be omitted. Some components of the device or some steps of the method may be changed in arrangement order. In addition, some components of the device or some steps of the method may be replaced with other components or other steps.
[0102] Furthermore, some components or some steps of the first embodiment of the present disclosure may be added to the second embodiment of the present disclosure, or may replace some components or some steps of the second embodiment.
[0103] In addition, the components shown in the embodiments of the present disclosure are shown independently to represent the functions of different features. This does not mean that each component constitutes independent hardware or software. That is, for ease of description, each component is described separately. At least two components can be combined into a single component, or a single component can be divided into multiple components. The embodiment in which at least two components are combined into a single component and the embodiment in which a single component is divided into multiple components can be included in the scope of the present disclosure, as long as they do not deviate from the essence of the present disclosure.
[0104] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. In describing the present disclosure, repeated descriptions about the same components are omitted.
[0105] Figure 1 1 is a block flow diagram schematically illustrating a configuration of a video encoding device 100. The video encoding device is configured to encode video data. The video encoding device may include a block partition unit, a prediction unit, a transform unit, a quantization unit, an entropy encoding unit, an inverse quantization unit, an inverse transform unit, an adder, an in-loop filter, a memory, and a subtractor.
[0106] The block partitioning unit 101 partitions from a coding block of a maximum size (hereinafter referred to as a maximum coding block) to a coding block of a minimum size (hereinafter referred to as a minimum coding block). There are various block partitioning methods. Quadtree partitioning (hereinafter referred to as QT partitioning) partitions the current coding block into four blocks of the same size. Binary tree partitioning (hereinafter referred to as BT partitioning) partitions the coding block into two blocks of the same size in the horizontal direction or the vertical direction. In addition to the QT and BT partitioning methods, there are various partitioning methods. In addition, various partitioning methods can be applied simultaneously.
[0107] Figure 2 The operation flow of the block partitioning unit in the video encoding device is shown. As a partitioning method of the higher depth block, one of QT partitioning 201 and BT partitioning 203 may be selected. When QT partitioning is performed on the higher depth block, the current block may be determined by generating a lower depth block by dividing the higher depth block into four (202). When BT partitioning is performed on the higher depth block, the current block may be determined by generating a lower depth block by dividing the higher depth block into two (204).
[0108] The prediction unit 102 generates a prediction block of the current original block using pixels adjacent to the current block to be predicted (hereinafter referred to as the prediction block) or pixels in a reference picture that has been encoded / decoded. One or more prediction blocks can be generated within the coding block. When there is a prediction block in the coding block, the prediction block has the same shape as the coding block. Video signal prediction mainly consists of intra-frame prediction and inter-frame prediction. In intra-frame prediction, a prediction block is generated using adjacent pixels of the current block. In inter-frame prediction, a prediction block is generated using a block that is most similar to the current block in a reference picture that has been encoded / decoded. Subsequently, the residual block can be determined as the difference between the original block and the prediction block. The residual block can undergo various techniques (such as rate-distortion optimization (RDO)) to determine the optimal prediction mode of the prediction block. The RDO cost calculation formula is the same as Formula 1.
[0109] [Formula 1]
[0110] D, R, and J represent the degradation due to quantization, the rate of compressed stream, and the RD cost, respectively. Φ is the coding mode. λ is the Lagrange multiplier and is used as a proportional correction coefficient to match the unit of error amount with the unit of bit amount. In order to select the optimal coding mode during the encoding process, J (i.e., RD cost value) when the optimal coding mode is applied must be smaller than J (i.e., RD cost value) when other coding modes are applied. The RD cost value is calculated while considering both bit rate and error.
[0111] Figure 3 300 is a flowchart showing the operation flow in the prediction unit of a video encoding device. When intra-frame prediction (301) is performed using original information and reconstruction information, the optimal intra-frame prediction mode (302) of each prediction mode is determined using the RD cost value, and a prediction is generated. When inter-frame prediction (303) is performed using the reconstruction information, the RD cost value is calculated for the skip mode, the merge mode, and the AMVP mode. The merge candidate detector 304 configures a set of candidate motion information for the skip mode and the merge mode. Among the sets of candidate motion information, the optimal motion information is determined using the RD cost value (305). The AMVP candidate detector 306 configures a set of candidate motion information for the AMVP mode. Motion estimation is performed using the set of candidate motion information to determine the optimal motion information (307). A prediction block (308) is generated by performing motion compensation using the optimal motion information determined in each mode.
[0112] Figure 4The operation flow 400 of the motion estimator of the prediction unit in the video encoding device is shown. The motion estimation starting point in the reconstructed picture is determined using the candidate motion information of the AMVP mode (401). In this case, the AMVP candidate motion information (i.e., the information means the motion vector predictor) may not be used to determine the motion estimation starting point. However, the motion estimation may start at any pre-set starting point. The maximum / minimum precision (N, M) for searching the motion vector for motion estimation is determined (402), and the motion estimation pattern under the current estimation precision N is determined, and motion estimation is performed (403). Here, the motion estimation pattern refers to Fig.10 . Fig.10 An example 1000 of four motion estimation styles is shown. The pixel point marked with R in each example refers to the pixel point currently being estimated. The pixel point marked with S1 refers to the pixel point that is under the first step of motion estimation for each style among the pixel points adjacent to the R pixel. The pixel point marked with S2 refers to the pixel point from the adjacent pixel point of the same style among the S1 pixel point to the pixel point with the smallest motion estimation cost value. In this way, motion estimation can be repeatedly performed until point SL (L=1, 2, ...). After performing motion estimation, the current motion estimation accuracy N is refined (404). When the refined accuracy N becomes less than the minimum estimation accuracy M, the optimal motion information corresponding to the current accuracy is determined as the optimal motion information of the current block, and the flowchart terminates. Otherwise, the flowchart returns to step 403 and the above process is repeated.
[0113] The above inter prediction can be composed of three modes (skip mode, merge mode, AMVP mode). Each prediction mode uses motion information (prediction direction information, reference picture information, motion vector) to obtain a prediction block for the current block, and there may be additional prediction modes using motion information.
[0114] The skip mode uses the motion information of the pre-reconstructed area to determine the optimal prediction information. A motion information candidate group is configured within the reconstructed area, and a candidate with the lowest RD cost value in the candidate group is used as prediction information to generate a prediction block. Here, the method of configuring the motion information candidate group is the same as the method of configuring the motion information candidate group in the merge mode to be described below, so its description will be omitted.
[0115] The merge mode is the same as the skip mode in that the motion information of the pre-reconstructed area is used to determine the optimal prediction information. However, in the skip mode, the motion information that allows the prediction error to be 0 is searched in the motion information candidate group, while in the merge mode, the motion information that allows the prediction error to be non-0 is searched in the motion information candidate group. In a similar manner to the skip mode, in the merge mode, the motion information candidate group is configured within the reconstructed area, and the candidate with the lowest RD cost value in the candidate group is used as the prediction information to generate the prediction block.
[0116] Figure 5 A method 500 for generating a motion information candidate group in skip mode and merge mode is shown. The maximum number of motion information candidates can be determined equally in a video encoding device and a video decoding device. The number information can be sent in advance in a higher header of the video encoding device (the higher header means a parameter sent from a higher layer of the block (such as a video parameter layer, a sequence parameter layer, and a picture parameter layer)). In the description of step S501 and step S502, the motion information derived using the corresponding motion information is included in the motion information candidate group only when the spatial candidate block and the temporal candidate block are encoded in the inter-frame prediction mode. In step S501, 4 candidates are selected from 5 spatial candidate blocks around the current block within the same picture. For the position of the spatial candidate (700), refer to Figure 7 The position of each candidate can be changed to any block in the reconstruction area. The spatial candidates are considered in the order of A1, A2, A3, A4 and A5, and the motion information of the first available spatial candidate block is determined as the spatial candidate. When there is duplicate motion information, only the motion information of the candidate with a high priority is considered. In step S502, one candidate of the two temporal candidate blocks is selected. For the position of the temporal candidate, refer to Figure 7 . The position of each candidate is determined based on a block in a collocated picture having the same position as the position of the current block in the current picture. Here, the video encoding device and the video decoding device may set the collocated picture in the reconstructed picture having the same condition. The B1 block and the B2 block are considered in order, and the motion information of the first available candidate block is determined as the temporal candidate. Reference Figure 8 Method 800 for determining motion information of a temporal candidate. The motion information of each of the candidate blocks B1 and B2 in the co-located picture indicates a prediction block in the reference picture B. (However, the reference picture of each candidate block may be different from each other. In this description, for convenience, all these reference pictures are referred to as reference pictures B.) The ratio of the distance between the current picture and the reference picture A and the distance between the co-located picture and the reference picture B is calculated. The motion vector of the candidate block is scaled by the calculated ratio to determine the motion vector of the temporal candidate motion information. Formula 2 refers to a scaling formula.
[0117] [Formula 2]
[0118] MV represents the motion vector of the temporal candidate block motion information, MV scale Represents a scaled motion vector, TB represents the temporal distance between the co-located picture and the reference picture B, and TD represents the temporal distance between the current picture and the reference picture A. In addition, the reference picture A and the reference picture B may be the same reference picture. The scaled motion vector is determined as the motion vector of the temporal candidate, and the reference picture information of the temporal candidate motion information is determined as the reference picture of the current picture, thereby deriving the motion information of the temporal candidate. Step S503 is performed only when the maximum number of motion information candidates is not met in steps S501 and S502. This step S503 is to add a new bidirectional motion information candidate as a combination of the motion information candidates derived from the previous step. The bidirectional motion information candidate refers to a new candidate combination of motion information in the past or future direction previously derived. Fig. 9 Table 900 in FIG. 5 shows the priorities of the bidirectional motion information candidate combinations. In addition to the combinations in the table, there may be other combinations. The table shows an example. When the maximum number of motion information candidates is not satisfied even with the use of bidirectional motion information candidates, step S504 is performed. In step S504, the maximum number of motion information candidates is satisfied by fixing the motion vector of the motion information candidate to a zero motion vector and changing the reference picture based on the prediction direction.
[0119] In the AMVP mode, optimal motion information is determined via motion estimation for each reference picture based on a prediction direction. Here, the prediction direction may be unidirectional including only one of the past / future directions, or may be bidirectional including both the past and future directions. A prediction block is generated by performing motion compensation using the optimal motion information determined via motion estimation. Here, a motion information candidate group for motion estimation is derived for each reference picture according to the prediction direction. The motion information candidate group is used as a starting point for motion estimation. A method 600 for deriving a motion information candidate group for motion estimation in the AMVP mode may be referred to. Figure 6 .
[0120] The maximum number of motion information candidates may be determined equally in the video encoding device and the video decoding device. The number information may be sent in advance in a higher header of the video encoding device. Only when the spatial candidate block and the temporal candidate block are encoded in the inter-frame prediction mode in the description of step S601 and step S602, the motion information derived using the corresponding motion information is included in the motion information candidate group. In step S601, Figure 5Different from the description of step S501 in , the number of derived spatial candidates (two) may vary, and the priority used to select the spatial candidates may also vary. The remaining description is the same as the description of step S501. The description of step S602 is the same as the description of step S502. In step S603, when there is duplicate motion information in the candidates derived so far, the duplicate motion information is removed. Step S604 is the same as the description of step S504. The motion information candidate with the smallest RD cost value among the motion information candidates derived in this manner is selected as the optimal motion information candidate, and the optimal motion information of the AMVP mode is obtained through motion estimation processing based on the corresponding motion information.
[0121] The transform unit 103 generates a transform block by transforming the residual block of the difference between the original block and the prediction block. The transform block is the smallest unit for the transform and quantization process. The transform unit transforms the residual signal to the frequency region to generate a transform block with transform coefficients. Various transform methods such as discrete cosine transform (DCT), discrete sine transform (DS) and KLT (Karhunen Loeve transform) can be used as methods for transforming the residual signal to the frequency region. Based on these methods, transform coefficients are generated by transforming the residual signal into the frequency region. In order to conveniently use the transform method, matrix operations are performed using base vectors. Depending on which prediction mode is used to encode the prediction block, the transform method can be used together during the matrix operation. For example, when performing intra-frame prediction, based on the prediction mode, discrete cosine transform can be used in the horizontal direction, and discrete sine transform can be used in the vertical direction.
[0122] The quantization unit 104 quantizes the transform block to generate a quantized transform block. That is, the quantization unit quantizes the transform coefficients of the transform block generated by the transform unit 103 to generate a quantized transform block having quantized transform coefficients. Dead zone uniform threshold quantization (DZUTQ) or a quantization weighting matrix may be used as a quantization method. However, various quantization methods, such as improved quantization thereof, may be used.
[0123] In the above description, it is shown and described that the video encoding device includes a transform unit and a quantization unit. However, the transform unit and the quantization unit may be optionally included therein. That is, the video encoding device may transform the residual block to generate a transform block, and may not perform a quantization process. Alternatively, the video encoding device may not transform the residual block into a frequency coefficient but may only perform a quantization process. Alternatively, the video encoding device may not perform both the transform process and the quantization process. Although the video encoding device does not perform a transform process and / or a quantization process, the block input to the input end of the entropy encoding unit is generally referred to as a "quantized transform block".
[0124] The entropy coding unit 105 encodes the quantized transform block and outputs a bitstream. That is, the entropy coding unit encodes the coefficients of the quantized transform block output from the quantization unit using various encoding methods such as an entropy encoding method. The entropy coding unit 105 generates and outputs a bitstream including additional information necessary for decoding the corresponding block in a video decoding device to be described later (for example, information about a prediction mode (the information about a prediction mode may include motion information or intra-frame prediction mode information determined by a prediction unit), quantization coefficients, etc.).
[0125] The inverse quantization unit 106 reconstructs the inverse quantized transform block by inversely performing a quantization method used for quantization on the quantized transform block.
[0126] The inverse transform unit 107 reconstructs the residual block by inversely transforming the inversely quantized transform block using the same method as that used for transforming. That is, the transform method used in the transform unit is performed inversely.
[0127] In the above description, the inverse quantization unit and the inverse transform unit may perform inverse quantization and inverse transform by inversely using the quantization method and the transform method used in the quantization unit and the transform unit, respectively. Optionally, when only quantization is performed but no transform is performed in the transform unit and the quantization unit, only inverse quantization may be performed and inverse transform may not be performed. Optionally, when neither transform nor quantization is performed in the quantization unit and the transform unit, the inverse quantization unit and the inverse transform unit may not perform both inverse transform and inverse quantization, or may be omitted and not included in the video encoding device.
[0128] The adder 108 reconstructs the current block by adding the residual signal generated by the inverse transform unit and the prediction block generated through the prediction.
[0129] The filter 109 may also filter the entire picture after reconstructing all blocks in the current picture. The filter may perform filtering such as deblocking filtering and SAO (Sample Adaptive Offset). Deblocking filtering refers to a process of reducing block distortion that occurs when encoding video data on a block basis. SAO (Sample Adaptive Offset) refers to a process of minimizing the difference between a reconstructed video and an original video by subtracting a specified value from a reconstructed pixel or adding a specified value to a reconstructed pixel.
[0130] The memory 110 stores a reconstructed current block obtained by adding a residual signal generated by the inverse transform unit to a prediction block generated by prediction and then performing additional filtering using an in-loop filter. The reconstructed current block may be used to predict a next block or a next picture.
[0131] The subtractor 111 generates a residual block by subtracting the prediction block from the current original block.
[0132] Fig.111100 is a flowchart showing a coding flow of coding information in a video coding device. In step S1101, information indicating whether a current coding block is partitioned is encoded. In step S1102, it is determined whether the current coding block is partitioned. When it is determined in step S1102 that the current coding block is partitioned, in step S1103, information about which partitioning method is used in QT partitioning, BT horizontal partitioning, and BT vertical partitioning is encoded. In step S1104, the current coding block is partitioned based on the partitioning method. In step S1105, after transferring to the first sub-coding block partitioned within the current coding block, the process returns to step S1101. When it is determined in step S1102 that the current coding block is not partitioned, information about activation of a skip mode is encoded in step S1106. In step S1107, it is determined whether the skip mode is activated. When the skip mode is activated in step S1107, merge candidate index information of the skip mode is encoded in step S1113. Thereafter, the processing transfers to step S1123. The description thereof will be described in detail below. When the skip mode is not activated in step S1107, the prediction mode is encoded in step S1108. In step S1109, it is determined whether the prediction mode is an inter-frame prediction mode or an intra-frame prediction mode. When the prediction mode is an intra-frame prediction mode in step S1109, the intra-frame prediction mode information is encoded in step S1110, and the processing transfers to step S1123. The description thereof will be described in detail below. When the prediction mode is an inter-frame prediction mode in step S1109, information about the activation of the merge mode is encoded in step S1111. In step S1112, it is determined whether the merge mode is activated. When the merge mode is activated in step S1112, the processing transfers to step S1113 to encode the merge candidate index information of the merge mode. When the merge mode is not activated in step S1112, the prediction direction is encoded in step S1114. Here, the prediction direction may be one of the past direction, the future direction, and the bidirectional direction. In step S1115, it is determined whether the prediction direction is the future direction. When the prediction direction is not the future direction in step S1115, the past reference picture index information is encoded in step S1116. In step S1117, the MVD (motion vector difference) information of the past direction is encoded. In step S1118, the motion vector predictor (MVP) information of the past direction is encoded. When the prediction direction is the future direction or bidirectional in step S1115, or when step S1118 ends, it is determined in step S1119 whether the prediction direction is the past direction. When the prediction direction is not the past direction in step S1119, the reference picture index information of the future direction is encoded in step S1120. In step S1121, the MVD information of the future direction is encoded.In step S1122, the MVP information of the future direction is encoded. When the prediction direction is the past direction or bidirectional in step S1119, or when step S1122 ends, it is determined in step S1123 whether the encoding of all sub-coding blocks is finished. Here, the processing of this step is performed even after step S1113 is completed. When completed, the flowchart is terminated. When not completed, in step S1124, the processing is transferred from the current sub-coding block to the next sub-coding block, and the above processing is repeated from step S1101.
[0133] Fig.12 12 is a block flow diagram briefly showing the configuration of a video decoding device. The video decoding device is a device for decoding video data, and may generally include a block entropy decoding unit, an inverse quantization unit, an inverse transform unit, a prediction unit, an adder, an in-loop filter, and a memory. The encoding block in the video encoding device is referred to as a decoding block in the video decoding device.
[0134] The entropy decoding unit 1201 interprets a bit stream transmitted from the video encoding device to read out various information and quantized transform coefficients required to decode a block.
[0135] The inverse quantization unit 1202 reconstructs an inverse quantized block having inverse quantized coefficients by inversely performing a quantization method used for quantization on the quantized coefficients decoded by the entropy decoding unit.
[0136] The inverse transform unit 1203 reconstructs a residual block having a difference signal by inversely transforming the inversely quantized transform block using the same method as that used for transforming. The transform method used in the transform unit is performed inversely.
[0137] The prediction unit 1204 generates a prediction block using the prediction mode information decoded by the entropy decoding unit. Here, the prediction unit 1204 uses the same prediction method as that performed by the prediction unit of the video encoding device.
[0138] The adder 1205 reconstructs the current block by adding the residual signal reconstructed in the inverse transform unit and the prediction block generated through the prediction.
[0139] The filter 1206 performs additional filtering on the entire picture after reconstructing all blocks in the current picture, and the filtering includes deblocking filtering and SAO (Sample Adaptive Offset). The details are the same as those described with reference to the in-loop filter of the video encoding device described above.
[0140] The memory 1207 stores the reconstructed current block obtained by adding the residual signal generated by the inverse transform unit to the prediction block generated by the prediction and then filtering by an in-loop filter. The reconstructed current block can be used to predict the next block or the next picture.
[0141] Fig.13 An operation flow 1300 of a block partitioning unit in a video decoding device is shown. Partition information is extracted from a higher depth block layer (1301), and a partitioning method is selected from QT partitioning (1302) or BT partitioning (1304). When QT partitioning is performed, a current block is determined by generating a lower depth block by quartering a higher depth block (1303). When BT partitioning is performed, a current block is determined by generating a lower depth block by bisecting a higher depth block (1305).
[0142] Fig.14 1400 is a flowchart showing the operation flow in the prediction unit of the video decoding device. When the prediction mode is the intra prediction mode, a prediction block is generated by determining the optimal intra prediction mode information (1401) and performing intra prediction (1402). When the prediction mode is the inter prediction mode, the optimal prediction mode is determined in the skip mode, the merge mode and the AMVP mode (1403). When decoding is performed in the skip mode or the merge mode, the merge candidate detector configures a candidate motion information set for the skip mode and the merge mode (1404). Among the candidate motion information sets, the optimal motion information is determined (1405). When it is decoded in the AMVP mode, the AMVP candidate detector configures a candidate motion information set for the AMVP mode (1406). Among the candidate motion information candidates, the optimal motion information is determined based on the transmitted MVP information (1407). Thereafter, motion compensation is performed using the optimal motion information determined in each mode to generate a prediction block (1408).
[0143] Fig.151500 is a flowchart showing a decoding flow of encoding information in a video decoding device. In step S1501, information indicating whether a current decoding block is partitioned is decoded. In step S1502, it is determined whether the current decoding block is partitioned. When the current decoding block is partitioned in step S1502, information about which partitioning method is used in QT partitioning, BT horizontal partitioning, and BT vertical partitioning is decoded in step S1503. In step S1504, the current decoding block is partitioned according to the partitioning method. In step S1505, after transferring to the first sub-decoded block partitioned within the current decoding block, the process returns to step S1501. When the current decoding block is not partitioned in step S1502, information about activation of a skip mode is decoded in step S1506. In step S1507, it is determined whether the skip mode is activated. When the skip mode is activated in step S1507, merge candidate index information for the skip mode is decoded in step S1513. Then, the process transfers to step S1523. The description thereof will be described in detail below. When the skip mode is not activated in step S1507, the prediction mode is decoded in step S1508. In step S1509, it is determined whether the prediction mode is an inter-frame prediction mode or an intra-frame prediction mode. When the prediction mode is an intra-frame prediction mode in step S1509, the intra-frame prediction mode information is decoded in step S1510. The processing is transferred to step S1523. The description thereof will be described in detail below. When the prediction mode is an inter-frame prediction mode in step S1509, information about the activation of the merge mode is decoded in step S1511. In step S1512, it is determined whether the merge mode is activated. When the merge mode is activated in step S1512, the processing is transferred to step S1513, and the merge candidate index information for the merge mode is decoded in step S1513. When the merge mode is not activated in step S1512, the prediction direction is decoded in step S1514. Here, the prediction direction can be one of the past direction, the future direction, and the bidirectional direction. In step S1515, it is determined whether the prediction direction is the future direction. When the prediction direction is not the future direction in step S1515, the reference picture index information of the past direction is decoded in step S1516. In step S1517, the MVD (motion vector difference) information of the past direction is decoded. In step S1518, the motion vector predictor (MVP) information of the past direction is decoded. When the prediction direction is the future direction or bidirectional in step S1515, or when step S1518 is completed, it is determined whether the prediction direction is the past direction in step S1519. When the prediction direction is not the past direction in step S1519, the reference picture index information of the future direction is decoded in step S1520. In step S1521, the MVD information of the future direction is decoded.In step S1522, the MVP information of the future direction is decoded. When the prediction direction is the past direction or the bidirectional direction in step S1519, or when step S1522 is terminated, it is determined in step S1523 whether the decoding of all sub-decoding blocks is completed. Here, the processing of this step is performed even after step S1513 is completed. When completed, the flowchart is terminated. When not completed, in step S1524, the processing is transferred from the current sub-decoding block to the next sub-decoding block, and the above processing is repeated from step S1501.
[0144] In this embodiment, a method of partitioning the current coding block by searching for motion boundary points at the boundary of the current coding block using motion information around the current coding block is shown. In addition, a method of predicting each sub-coding block using a decoder-side motion vector derivation (DMVD) mode is shown.
[0145] Fig.16 FIG. 1 shows a block partitioning unit 1600 in a video encoding device according to an embodiment. A higher depth block undergoes processing similar to Figure 2 The same processes 1601 to 1604 as processes 201 to 204 in the above are performed to determine the current block. When the MT (motion tree) partitions the higher depth block (1605), the motion boundary point is searched (1606). Here, MT partitioning refers to a method of partitioning a block at a motion boundary point by searching for a motion boundary point in the higher depth block. The method of searching for a motion boundary point will be described in detail below. The current block is determined by partitioning the higher depth block (1607) based on the searched motion boundary point. The method for partitioning the higher depth block is described in detail below.
[0146] Fig.17 FIG. 1 shows a block partitioning unit 1700 in a video decoding device according to an embodiment. Extract partition information for a higher depth block ( 1701 ). Determine one of a QT partition, a BT partition, or a MT partition based on the partition information. Execute Fig.13 The current block is determined by performing steps 1702 to 1705 which are the same as steps 1302 to 1305 in step 1302 to 1305. In the MT partition (1706), a motion boundary point is searched (1707). The method for searching the motion boundary point will be described in detail below. The current block is determined by partitioning the higher depth block (1708) based on the searched motion boundary point. The method for partitioning the higher depth block is described in detail below.
[0147] exist Fig.16 and Fig.17 In the , QT, BT, and MT partitions can be applied to the range from the largest coding block to the smallest coding block. However, depending on the size and depth of the coding block, only some partitioning methods can be used to perform partitioning. For example, assume that the size of the largest coding block is 64×64 and the size of the smallest coding block is 4×4. Fig.32 It can be referred to the case where the size of the coding block applicable to the QT partition is 64×64 to 16×16, the size of the maximum coding block applicable to the BT partition and the MT partition is 16×16, and the partitionable depth is 3. Fig.32 In the diagram 3200 of FIG. 3200 , the solid line represents the QT partition, the dotted line represents the BT partition, and the combined dotted line represents the MT partition. Fig.32 As shown, the partitioning of the coding block can be performed under the above conditions. When MT partitioning occurs, the current coding block can be partitioned into coding blocks of odd value sizes. When a coding block with a size of 17×8 occurs and BT partitioning is performed on the coding block, the coding block can be partitioned into 9×4 and 8×4 coding blocks. In addition, the maximum size, minimum size, and partitionable depth of the coding block for each partitioning method can be sent in advance in a higher header.
[0148] Fig.18 FIG. 1 shows a prediction unit 1800 in a video encoding device according to an embodiment. Figure 3 Processes 1801 to 1807 are the same as processes 301 to 307 in . The optimal intra prediction mode determined via processes 1801 to 1802 may be used to generate a prediction block. Alternatively, motion compensation 1810 may be performed using the optimal motion information determined via processes 1803 to 1807 to generate a prediction block. In addition, when inter prediction is performed, there is a motion information determination method using a DMVD mode. In order to perform motion estimation in the same manner in a video encoding / decoding device, initial motion information (1808) is determined using motion information of a reconstructed region. Motion estimation (1809) is performed using the determined initial motion information to determine the optimal motion information. Then, a prediction block (1810) is generated by performing motion compensation.
[0149] Fig.19 FIG. 1 shows a prediction unit 1900 in a video decoding device according to an embodiment. Fig.14 The same processes 1401 to 1407 as in the processes 1901 to 1907. The optimal intra prediction mode determined via the processes 1901 to 1902 may be used to generate a prediction block. Alternatively, motion compensation (1910) may be performed using the optimal motion information determined via the processes 1903 to 1907 to generate a prediction block. In addition, when inter prediction is performed, there is a motion information determination method using a DMVD mode. In order to perform motion estimation in the same manner in a video encoding / decoding device, initial motion information (1908) is determined using motion information of a reconstructed region. Motion estimation (1909) is performed using the determined initial motion information to determine the optimal motion information. A prediction block (1910) is generated by performing motion compensation.
[0150] By using Fig.18 and Fig.19 The optimal motion information obtained by the motion estimation of the template block in the DMVD motion estimator 1809 and 1909 is applied to the current prediction block. Optionally, the reconstructed motion information in the template block can be determined as the optimal motion information of the current prediction block. In addition, when the current coding block is generated by MT partitioning of a higher depth block, only the DMVD mode can be used as the prediction mode of the current coding block. In addition, the optimal motion information derived in the DMVD mode can be used as a candidate for the AMVP mode.
[0151] Fig. 20 A DMVD motion estimator 2000 in a prediction unit in a video encoding / decoding device is shown. After performing a DMVD mode using initial motion information determined by a DMVD initial motion information detector in the video encoding / decoding device, optimal motion information is determined. The DMVD mode includes a mode using a template (hereinafter referred to as a "template matching mode") and a mode not using a template (hereinafter referred to as a "bidirectional matching mode"). When the bidirectional matching mode is used (2001), a unidirectional motion vector of each initial motion information is linearly scaled to a reference picture in the opposite prediction direction. Here, the scaling of the motion vector is performed in proportion to the distance between the current picture and the reference picture in each direction. After determining the bidirectional motion vector like this (2002), the motion vector in each direction that minimizes the difference between the prediction block in the past direction and the prediction block in the future direction is determined as the optimal motion information (2003). Fig.21 Diagram 2102 in (diagram 2100) shows a method in which motion vectors in the past direction and the future direction of the current block are linearly generated in the bidirectional matching mode, and then a prediction block of the current block is generated as the average of the two prediction blocks in the bidirectional direction. When the template matching mode is used (2004), a template block is determined in the reconstruction area around the current block (2005). Fig. 22The template blocks are configured as shown in Examples 2201 to 2203 in (illustrated diagram 2200). In illustrated diagram 2201, the template blocks are determined in the lower left (template A), upper left (template B), upper left (template C), and upper right (template D) of the current block, respectively. The size and shape of each template block may be determined differently. In illustrated diagram 2202, the template blocks are determined in the lower left (template A), upper left (template B), upper left (template C), and upper right (template D) of the current block, respectively. The difference between them is that both the reconstruction areas on the left and above adjacent to the current block are used in illustrated diagram 2202. Schematic diagram 2203 refers to a method of generating a template block by simultaneously considering the template block generation methods in illustrated diagrams 2201 and 2202. In addition, the template block may be generated in the reconstruction area around the current block via various methods, including a method of determining the reconstruction areas on the left and above adjacent to the current block as a single template block. However, information indicating the shape and size of the template block may be sent from the video encoding device. After performing motion estimation (2006) of detecting a prediction block most similar to each determined template block from a reference picture, the optimal motion information of each template block is estimated, and then the motion information most suitable for the current block is detected from the motion information. Therefore, the detected motion information is determined as the optimal motion information. Here, when estimating the motion information optimal for the template block, it can be based on Fig.10 Motion estimation is performed by any style selected from the four motion estimation styles shown. The cost value of motion estimation means the sum of the amount of prediction error and the amount of virtual bits of motion information. The prediction error can be obtained through various calculation methods such as SAD (sum of absolute differences), SATD (sum of absolute Hadamard transform differences) and SSD (sum of squared differences). Equations (3), (4) and (5) show the calculation methods of SAD, SATD and SSD respectively.
[0152] [Formula 3]
[0153] [Formula 4]
[0154] [Formula 5]
[0155] Wherein i, j represent pixel positions, template (i, j) represents pixels of the template block, and PredBlk (i, j) represents pixels of the prediction block. Here, the HT() function of Formula 4 refers to the function value obtained by performing Hadamard transform on the differential block between the template block and the prediction block. The virtual bit amount of motion information is not the information actually sent, but is obtained by calculating the virtual bit amount of motion information that is expected to be the same in the video encoding device and the video decoding device. For example, the difference vector size between the motion vector of the initial motion information and the motion vector in the motion information under the current motion estimation can be calculated and determined as the virtual bit amount. In addition, the virtual bit amount of motion information can also be calculated based on the bit amount of reference picture information. Fig.21 Figure 2101 in the figure shows a method in which, when each of the reconstructed areas adjacent to the left and above the current block is used as a single template block, a prediction block of the template block that is most similar to the template block is detected, and then the blocks adjacent to the corresponding template block are determined as the prediction blocks of the current block.
[0156] Fig.23 2300 is a flowchart showing a method for searching for a motion boundary point according to an embodiment. The flowchart is executed in a process of performing motion boundary point detection (1606, 1707) by a block partitioning unit in a video encoding / decoding device according to an embodiment. Index information of an initial motion boundary point (hereinafter referred to as "MB Idx") is initialized to -1, and the number of template blocks is initialized to 2. NumOfMB NumOfMB represents the number of motion boundary points. In this embodiment, it is assumed that the number of motion boundary points is limited to at most two. However, the number may be greater than 2. The method initializes all parameters in the CostBuf[N] buffer to infinity and all parameters in the IdxBuf[N] buffer to -1. Thereafter, an initial motion information list is determined (S2301). The method for determining the initial motion information list is described in detail in Fig.25 . Fig.25 The index in Table 2500 represents the priority of the initial motion information. Candidate motion information in AMVP mode, candidate motion information in merge mode, motion information of sub-blocks in the reconstructed area in the upper, left, upper left, upper right and lower left directions around the current block, zero motion information, etc. can be regarded as initial motion information. In addition, various motion information candidates based on the reconstruction information can be used. Thereafter, MB Idx is updated by increasing MB Idx by 1 (S2302). The method creates a template block in the reconstructed area based on the current motion boundary point (S2303), and searches for the optimal motion information of each template block (S2304). For a detailed description of steps S2303 and S2304, refer to Fig.28 and Fig.29 . Fig.28 and Fig.29 A method for determining template blocks in a reconstruction region based on current motion boundary points and then deriving optimal motion information for each template block is shown. Fig.28 In FIG. 2800 , only some reconstructed regions are generated as template blocks based on motion boundary points. The prediction block most similar to the template A block is detected from the reference picture, and the prediction block most similar to the template B block is detected from the reference picture. The template C block refers to a block that is a combination of the template A and B blocks. The block most similar to the template C block is detected from the reference picture. Fig.29 In the diagram 2900 of Fig.28 Differently, the template A, B, and C blocks are determined based on the motion boundary points in the reconstruction area which is the left area and the top area adjacent to the current block. The prediction block most similar to each template block is detected from the reference picture. Thereafter, the partition cost value (PartCost) at the current motion boundary point is calculated using the cost value (SAD, SATD, SSD cost value, etc.) of the optimal motion information of each template block. MB Idx ) (S2305). Here, various information can be used to determine the partition cost value. The first method for determining the partition cost value is: the sum of the cost value of template A block and the cost value of template B block is determined as the partition cost value. Here, the cost value of template C block is not used. For the second method for determining the partition cost value, refer to Formula 6.
[0157] [Formula 6] Partition cost value = [1-(Cost value of template A block + Cost value of template B block) / Cost value of template C block] × 100.0 (Wherein, the cost value of template C block ≥ (the cost value of template A block + the cost value of template B block)) Formula 6 refers to calculating the percentage change of the sum of the cost values of template A block and template B block relative to the cost value of template C block.
[0158] The method determines whether the partition cost value at the current motion boundary point is less than one or more element values in the CostBuf buffer (S2306). When the current partition cost value is less than one or more element values in the Costbuf buffer, the current partition cost value is stored in the Costbuf[N-1] buffer, and the current MB Idx is stored in the IdxBuf[N-1] buffer (S2307). Thereafter, the element values in the CostBuf buffer are sorted in ascending order (S2308). The method sorts the element values in the IdxBuf buffer in the same manner as the CostBuf buffer sorting order (S2309). For example, when N is 4, the 4 elements of the CostBuf buffer are stored as {100, 200, 150, 180}, and the 4 elements of the IdxBuf buffer are stored as {0, 3, 2, 8}, CostBuf is sorted in the order of {100, 150, 180, 200}, and IdxBuf is sorted in the order of {0, 2, 8, 3} in the same order as the order in the CostBuf buffer. When the current partition cost value is greater than any element in the Costbuf buffer, or when step S2309 terminates, it is determined whether the current MB Idx is the last search candidate partition boundary point (S2310). When the current motion boundary point is the last search candidate partition boundary point, the flowchart is terminated. Otherwise, the method returns to step S2302, where the MB Idx is updated and the above process is repeated.
[0159] Fig.24 A method 2400 for detecting motion boundary points according to one embodiment is shown, and for Fig.23 The accelerated algorithm is shown in the flowchart. Fig.23 The flowchart may be executed in the process of motion boundary point detection (1606, 1707) performed by the block partition unit in the video encoding / decoding device according to one embodiment, or may be executed in the DMVD motion estimator of the prediction unit. Fig.23 The initial information in the table is the same as that in the table. The method for setting the initial motion information list (S2401) is the same as that in the table. Fig.23 The description of step S2301 in is the same. The left and top areas adjacent to the current block in the reconstructed area and the upper left, upper right and lower left partial areas not adjacent to the current block, if necessary, are respectively determined as template blocks (S2402). The cost values of all motion information within the motion estimation range using the initial motion information of the template block are calculated and stored based on lines (S2403). Here, the method of storing the cost value based on lines refers to Fig.26 . Fig.26A method 2600 is shown for storing cost values on a line basis in an area determined to be a template block in a reconstruction area. H-1 This means that the cost value corresponding to the motion information in the left area adjacent to the current block is stored on a line basis. W-1 This means that the cost value corresponding to the motion information in the top area adjacent to the current block is stored on a line basis. Thereafter, the process of updating MBIdx (S2404) is the same as Fig.23 The description of step S2302 in FIG. Thereafter, the template block is recreated in the reconstruction area based on the current motion boundary point (S2405). The method of recreating the template block is the same as Fig.23 The same as the description of step S2303 in . In step S2403, the cost value of each line of the optimal motion information of the recreated template block is pre-calculated. Therefore, the partition cost value at the current motion boundary point is calculated using the corresponding cost value of each line (S2406). Here, the method for calculating the partition cost value is the same as Figure 3 The method described in step S2305 of FIG. 2 is the same as that described in step S2305 of FIG. Fig.23 When the current MB Idx is not the last searched candidate partition boundary point in step S2411, the method returns to step S2404, and the above process is repeated in step S2404.
[0160] Can be based on Fig.23 and Fig.24 The higher depth block is partitioned at one or two motion boundary points. Here, when the difference between the two higher partition cost values in the CostBuf buffer is greater than a certain threshold, the corresponding block may be partitioned at one motion boundary point. Otherwise, when the two higher partition cost values are similar to each other, the block may be partitioned at two motion boundary points. In addition, when the corresponding motion boundary point information is sent to the video decoding device, the video decoding device may omit the partition. Fig.23 and Fig.24 The content for transmitting the motion boundary point information is not included in the encoding / decoding information in the video encoding / decoding device to be described later. The optimal MB Idx among all MB Idxs may be directly transmitted. However, after determining M candidate MB Idxs, the motion boundary point having the optimal RD cost value among the M MB Idxs may be transmitted.
[0161] Fig. 27 A method 2700 for partitioning a higher depth block (current block) based on the number of motion boundary points is shown. Fig.23 and 24The process of partition cost can be used to obtain MB Idx in the order of smaller partition cost values. Therefore, when only one motion boundary point is used, the motion boundary point corresponding to the MB Idx with the smallest partition cost value is determined as the partition boundary point as shown in FIG2701. Therefore, the current block is partitioned into lower depth blocks (prediction blocks) A and B. The reconstruction of the lower depth block A is completed, and the optimal motion information of the lower depth block B can be derived by performing motion estimation again using the template A block and the area under the lower depth block A. When two motion boundary points are used and Fig.23 and Fig.24 When there is at least one MB Idx in each of the top and left directions in the IdxBuf array of the image, as shown in diagram 2702, the motion boundary point corresponding to the MB Idx with the smallest partition cost value in the top direction and the motion boundary point corresponding to the MB Idx with the smallest partition cost value in the left direction are determined as partition boundary points. Then the current block is partitioned into lower depth blocks A, B, C, and D. Reconstruction of the lower depth block A is completed, and then the optimal motion information for the lower depth block B can be derived by performing motion estimation again using the template D block and the right area adjacent to the lower depth block A. Reconstruction of the lower depth block A is completed, and the optimal motion information for the lower depth block C can be derived by performing motion estimation again using the template A block and the area under the lower depth block A. Reconstruction of the lower depth blocks B and C is completed, and the optimal motion information for the lower depth block D can be derived by performing motion estimation again using the right area adjacent to the lower depth block C and the area under the lower depth block B.
[0162] The prediction block A may be obtained through the above-mentioned DMVD motion estimation using each of the template B and C blocks as one template block. The prediction block B may be obtained through the above-mentioned DMVD motion estimation using the template A block as one template block. However, when the prediction block A is first reconstructed by performing transform / quantization and inverse quantization / inverse transform processing not based on the current block but based on the prediction block, the prediction block B may be obtained through DMVD motion estimation using the template C' block and the template A block in the reconstructed area of the prediction block A as one template block. When two motion boundary points are used and Fig.23 and Fig.24 When there is at least one MB Idx in each of the top direction and the left direction in the IdxBuf array of the current block, two motion boundary points corresponding to the MB Idx having the minimum partition cost value in the top direction and the MB Idx having the minimum partition cost value in the left direction are determined as partition boundary points as shown in diagram 2702. The current block is partitioned into lower depth blocks A, B, C, and D.
[0163] Block A may be obtained via the above-mentioned DMVD motion estimation using each of the template B and C blocks as one template block. Prediction block B may be obtained via the above-mentioned DMVD motion estimation using the template block D as one template block. Prediction block C may be obtained via the above-mentioned DMVD motion estimation using the template block A as one template block. Prediction block D may be obtained via the above-mentioned DMVD motion estimation using the prediction blocks D and A as one template block. However, when transform / quantization and inverse quantization / inverse transform processing is performed not based on the current block but based on the prediction block, prediction block A may be reconstructed, and then prediction block B may be generated and reconstructed using the template B' block and the template D block as one template block, and then prediction block C may be generated and reconstructed using the template A block and the template C' block as one template block, and then prediction block D may be generated and reconstructed using the template D' block and the template A' block as one template block.
[0164] Fig.303000 is a flowchart showing a process of encoding some encoding information in an entropy encoding unit in a video encoding device. In this embodiment, an exemplary method for encoding partition information and prediction information of a block will be described. First, information indicating whether a coding block is partitioned is encoded (S3001). It is determined whether the coding block is partitioned (S3002). When it is determined that the coding block is partitioned, information indicating which of QT partitioning, BT partitioning, and MT partitioning is performed is encoded (S3003). Here, when MT partitioning is performed by fixedly using one or two motion boundary points, there is no additional encoding information. However, when one or two motion boundary points are adaptively used according to the characteristics of the current block, information about how many motion boundary points are used is additionally encoded. In addition, when the partition range exceeds the preset partition range for each partition method, one of the remaining partition methods other than the corresponding partition method whose partition range exceeds the preset partition range is encoded. Thereafter, the current coding block is partitioned according to the partition method at the partition boundary point (S3004), and then the process is transferred to the first encoding object sub-block after partitioning (S3005). Thereafter, the process returns to step S3001, in which the above process is repeated. When it is determined whether the coding block is not partitioned (S3002), it is determined whether the current coding block is a block partitioned based on the MT partitioning method (S3006). When the current coding block is not a block partitioned based on the MT partitioning method, information about the activation of the skip mode is encoded (S3007). It is determined whether the skip mode is activated (S3008). When the skip mode is not activated, the prediction mode information is encoded (S3009). The type of the prediction mode is determined (S3010). When the corresponding prediction mode is not the inter-frame prediction mode, the intra-frame prediction information is encoded (S3011), and then the process is transferred to step S3030. The subsequent steps will be described in detail below. When it is determined in step S3008 that the skip mode is activated, the DMVD mode operation information is encoded (S3012). It is determined whether the DMVD mode is activated (S3013). When the DMVD mode is activated, the template information is encoded (S3014). Here, the template information refers to motion information indicating which template block among the template blocks in the reconstruction area is used to predict the current block. When there is one template block, the corresponding information may not be encoded. Thereafter, the processing is transferred to step S3030. The subsequent steps will be described in detail below. When the current coding block is a block partitioned based on the MT partitioning method in step S3006, the template information is encoded (S3019). Then, the processing is transferred to step S3030. The subsequent steps will also be described in detail below. When the DMVD mode is not activated in step S3013, the merge candidate index information for the skip mode is encoded (S3020).When the prediction mode is the inter-frame prediction mode in step S3010, information about the activation of the merge mode is encoded (S3015). It is determined whether the merge mode is activated (S3016). When the merge mode is activated, information about the activation of the DMVD mode is encoded (S3017). It is determined whether the DMVD mode is activated (S3018). When the DMVD mode is activated, the template information is encoded (S3019). When the DMVD mode is not activated, the merge candidate index information for the merge mode is encoded (S3020). After steps S3019 and S3020 are terminated, the processing is transferred to step S3030. The subsequent steps will also be described in detail below. When the merge mode is not activated in step S3016, the prediction direction is encoded (S3021). The prediction direction may mean one of the past direction, the future direction, and the bidirectional direction. It is determined whether the prediction direction is the future direction (S3022). When the prediction direction is not the future direction, the reference picture index information of the past direction, the MVD information of the past direction, and the MVP information of the past direction are encoded (S3023, S3024, S3025). When the prediction direction is the future direction or after step S3025 is completed, it is determined whether the prediction direction is the past direction (S3026). When the prediction direction is not the past direction, the reference picture index information of the future direction, the MVD information of the future direction, and the MVP information of the future direction are encoded (S3027, S3028, S3029). When the prediction direction is the past direction or after step S3029 is completed, the process is transferred to step S3030. In the corresponding step S3030, it is determined whether the encoding of all sub-coding blocks is completed (S3030). When the encoding of all sub-coding blocks is completed, the flowchart is terminated. Otherwise, the process is transferred to the next sub-coding block (S3031), and then the process returns to step S3001, and the above process is repeated in step S3001.
[0165] Fig.31 3100 is a flowchart showing the process of decoding some coded information in an entropy decoding unit in a video decoding device. Fig.30A method for decoding encoding information of some encoding information in an entropy encoding unit in a video encoding device. First, information indicating whether a decoding block is partitioned is decoded (S3101). It is determined whether the decoding block is partitioned (S3102). When the decoding block is partitioned, information indicating which partitioning method among QT partitioning, BT partitioning and MT partitioning is performed is decoded (S3103). Here, when MT partitioning is performed by fixedly using one or two motion boundary points, there is no additional decoding information. However, when one or two motion boundary points are adaptively used according to the characteristics of the current block, information about how many motion boundary points are used is additionally decoded. Thereafter, the current decoding block is partitioned according to the partitioning method at the partition boundary point (S3104), and then the processing is transferred to the first sub-decoding block of the partition (S3105). Thereafter, the processing returns to step S3101, and the above-mentioned processing is repeated in step S3101. When it is determined in S3102 that the decoding block is not partitioned, it is determined whether the current decoding block is a block partitioned based on the MT partitioning method (S3106). When the current decoding block is not a block partitioned based on the MT partitioning method, information about the activation of the skip mode is decoded (S3107). It is determined whether the skip mode is activated (S3108). When the skip mode is not activated, the prediction mode information is decoded (S3109). The type of the prediction mode is determined (S3110). When the corresponding prediction mode is not the inter-frame prediction mode, the intra-frame prediction information is decoded (S3111), and then the processing is transferred to step S3130. The subsequent steps will be described in detail below. When it is determined in step S3108 that the skip mode is activated, information about the activation of the DMVD mode is decoded (S3112). It is determined whether the DMVD mode is activated (S3113). When the DMVD mode is activated, the template information is decoded (S3114). Thereafter, the processing is transferred to step S3130. The subsequent steps will be described in detail below. When the current decoding block is a block partitioned based on the MT partitioning method in step S3106, the template information is decoded (S3119). The processing then transfers to step S3130. The subsequent steps will also be described in detail below. When the DMVD mode is not activated in step S3113, the merge candidate index information of the skip mode is decoded (S3120). When the prediction mode is the inter-frame prediction mode in step S3110, the information about the activation of the merge mode is decoded (S3115). It is determined whether the merge mode is activated (S3116). When the merge mode is activated, the information about the activation of the DMVD mode is decoded (S3117). It is determined whether the DMVD mode is activated (S3118). When the DMVD mode is activated, the template information is decoded (S3119). When the DMVD mode is not activated, the merge candidate index information for the merge mode is decoded (S3120).After steps S3119 and S3120 are terminated, the process transfers to step S3130. The subsequent steps will also be described in detail below. When the merge mode is not activated in step S3116, the prediction direction is decoded (S3121). The prediction direction may mean one of the past direction, the future direction, and the bidirectional direction. It is determined whether the prediction direction is the future direction (S3122). When the prediction direction is not the future direction, the reference picture index information of the past direction, the MVD information of the past direction, and the MVP information of the past direction are decoded (S3123, S3124, S3125). When the prediction direction is the future direction or after step S3125 is completed, it is determined whether the prediction direction is the past direction (S3126). When the prediction direction is not the past direction, the reference picture index information of the future direction, the MVD information of the future direction, and the MVP information of the future direction are decoded (S3127, S3128, S3129). When the prediction direction is the past direction or after step S3129 is completed, the process transfers to step S3130. In the corresponding step S3130, it is determined whether the decoding of all sub-decoding blocks is completed (S3130). When the decoding of all sub-decoding blocks is completed, the flowchart is terminated. Otherwise, the process is transferred to the next sub-decoding block (S3131), and then the process returns to step S3101, and the above process is repeated in step S3101.
[0166] Fig.33 A DMVD motion estimator 3300 in a prediction unit in a video encoding / decoding device is shown. After performing a DMVD mode using initial motion information determined by a DMVD initial motion information detector in the video encoding / decoding device, optimal motion information is determined. The DMVD mode includes a mode using a template (hereinafter referred to as a "template matching mode") and a mode not using a template (hereinafter referred to as a "bidirectional matching mode"). When the bidirectional matching mode is used (3301), a unidirectional motion vector of each initial motion information is linearly scaled to a reference picture in the opposite prediction direction. Here, the scaling of the motion vector is performed in proportion to the distance between the current picture and the reference picture in each direction. After determining the bidirectional motion vector (3302), the motion vector in each direction that minimizes the difference between the prediction block in the past direction and the prediction block in the future direction is determined as the optimal motion information (3303). Fig.21Diagram 2102 in shows a method of linearly generating motion vectors in the past and future directions of the current block in the bidirectional matching mode, and then generating a prediction block of the current block as an average of two bidirectional prediction blocks. When the template matching mode is used (3304), the number of template blocks in the reconstruction area is determined. When a single template block is used (3305) (hereinafter referred to as "single template matching mode"), the left and top reconstruction areas adjacent to the current block are determined as template blocks, and the corresponding template blocks are used to determine the optimal motion information via motion estimation 3306. Fig.21 2101 in FIG. 2102 shows a method for searching for a prediction block of a template block most similar to a template block in a single template matching mode, and then determining a block adjacent to the corresponding template block as a prediction block of a current block. Here, when estimating the optimal motion information for the template block, the prediction block may be determined based on the prediction information from Fig.10 The motion estimation is performed by selecting any of the four motion estimation styles shown. The cost value of the motion estimation means the sum of the amount of prediction error and the amount of virtual bits of motion information. The prediction error can be obtained through various calculation methods such as SAD (Sum of Absolute Difference), SATD (Sum of Absolute Hadamard Transform Difference), and SSD (Sum of Squared Difference). This is consistent with the reference Fig. 20 Therefore, the detailed description thereof will be omitted.
[0167] When multiple template blocks are used (3307) (hereinafter referred to as "multi-template matching mode"), it is determined whether to use multiple template blocks in the reconstruction area in a fixed manner or to adaptively generate template blocks using motion boundary points. When the former is adopted (3308), it can be as follows Fig. 22The template blocks are configured as shown in the example diagrams 2201 to 2203 in . In diagram 2201, the template blocks are respectively determined as the lower left (template A), upper left (template B), upper left (template C) and upper right (template D) adjacent to the current block. The size and shape of each template block can be determined differently. In a similar manner to diagram 2201, in diagram 2202, the template blocks are respectively determined as the lower left (template A), upper left (template B), upper left (template C) and upper right (template D) adjacent to the current block. The difference between them is that both the left and upper reconstruction areas adjacent to the current block are used in diagram 2202. Diagram 2203 refers to a method for generating template blocks by considering the template block generation methods in diagrams 2201 and 2202 at the same time. In addition, template blocks can be generated in the reconstruction area around the current block via various methods, including a method of determining the left and upper reconstruction areas adjacent to the current block as a single template block. However, information indicating the shape and size of the template block can be sent from the video encoding device. After performing motion estimation (3306) of detecting the prediction block most similar to each determined template block from the reference picture, the optimal motion information of each template block is estimated, and then the motion information most suitable for the current block is detected from the estimated motion information. Therefore, the detected motion information is determined as the optimal motion information. When a template block is generated using motion boundary points instead of fixed template blocks, motion boundary points are first detected (3309). The motion boundary point detection method will be described in detail below. The optimal motion information for each template block is estimated by motion estimation for each template block divided based on the determined motion boundary points (3310). The optimal motion information of the adjacent template blocks is applied to each prediction block divided based on the motion boundary points. In addition, in the multi-template matching mode, the optimal motion information obtained by using the motion estimation of the template block can be applied to the current prediction block. Optionally, the reconstructed motion information in the template block can be directly determined as the optimal motion information of the current prediction block.
[0168] Fig.34 3400 is a flowchart showing a method for determining the slope of a motion boundary line of a current block partition at a determined motion boundary point. The first initial information initializes the motion boundary line slope index information (hereinafter referred to as "MB Angle") to -1, and the optimal partition cost value (hereinafter referred to as "PartCost Best”) is initialized to infinity. The MB angle is updated by increasing the current MB angle by 1 (S3401). Here, various methods can be used to pre-set the MB angle in the same manner in the video encoding / decoding device. Thereafter, the motion boundary search area is partitioned into two template blocks by partitioning the motion boundary search area along the slope direction of the motion boundary line corresponding to the current MB angle. Then, the method calculates the cost value of the optimal motion information for each template block via DMVD motion estimation (S3402). The partition cost value of the current MB angle (hereinafter referred to as “PartCost”) is calculated using the cost value of the optimal motion information for each template block. MB Angle ”) (S3403). Calculate PartCost MB Angle The method is similar to the above Fig.23 The method is the same as in. The detailed description of steps S3402 and S3403 refers to Fig.35 .
[0169] Fig.35 3500 is a diagram showing a method of determining a template block in a motion boundary search area based on a current motion boundary point along an inclined direction of a motion boundary line and deriving optimal motion information for each template block. The method detects a prediction block that is most similar to a template A block from a reference picture, and detects a prediction block that is most similar to a template B block from a reference picture. A template C block refers to a block that is a combination of template A and B blocks. The method detects a block that is most similar to a template C block from a reference picture. A line that extends across the current block in an inclined direction of a motion boundary line is called a virtual motion boundary line. In this case, an actual motion boundary line regarding pixels on and along which the virtual motion boundary line extends can be determined based on which of the two prediction blocks around the virtual line contains more pixel areas. The method determines PartCost MB Angle Is it less than the PartCost in the current MB angle? Best (S3404). When PartCost MB Angle Less than PartCost in the current MB angle Best When PartCost MB Angle Stored in PartCost Best The method stores the current MB angle in the optimal motion boundary slope index information (S3405). MB Angle Greater than or equal to PartCost in the current MB angle Best When or after step S3405 is completed, determine whether the current MB angle is the last search candidate motion boundary line slope index (S3406). When the current MB angle is the last search candidate motion boundary line slope index, terminate the flowchart. When this is not the case, the method returns to step S3401, and repeats the above process in step S3401.
[0170] When a reconstruction process is not first performed on the basis of each prediction block partitioned based on a motion boundary point, the method may create a prediction block for a current block and then may perform a filtering process at a prediction block boundary.
[0171] exist Fig.36 In FIG. 3600 of FIG. 3601 , the filtering method is described using an example in which there is a motion boundary point and partitioning is performed along a motion boundary line having a straight horizontal direction rather than a diagonal direction. The filtering method at the prediction block boundary may vary. For example, filtering may be performed using a weighted sum of adjacent pixels of each prediction block in the filtering area. Formula 7 is a filtering formula.
[0172] [Formula 7]
[0173] In Equation 7, a' and b' are filter values of pixels predicted by a and b. W1 to W4 are weight coefficients, W1+W2=1, and W3+W4=1. In this case, the pixels predicted by a and b can be filtered by substituting 0.8 into W1 and W4 and substituting 0.2 into W2 and W3. In addition, various filtering methods can be used.
[0174] exist Fig.37 In FIG. 3700 of FIG. 3701 , the filtering method is described using an example in which there is one motion boundary point and partitioning is performed along a motion boundary line having an oblique direction. Fig.36 Filtering is performed in a manner similar to the description of 3702 using the weighted sum of equation 7 based on the pixels of the prediction block adjacent to the motion boundary line. In the filtering method of diagram 3702, when the prediction block is divided based on the virtual motion boundary line, all pixels passing through the virtual motion boundary line are included in each partition block (partition blocks A and B). Thereafter, a prediction block is created for each partition block. Filtering may be performed based on the virtual motion boundary line instead of the actual motion boundary line. For each pixel passing through the corresponding boundary line, calculate which prediction block contains a larger number of pixel areas. In this diagram, the pixel area areas p and q are calculated based on the line through which the virtual motion boundary line passes in the pixel S to be filtered. Pixel S is filtered using equation 8.
[0175] [Formula 8]
[0176] S' is the filtered S pixel, pixel a is the predicted pixel of partition block A, and pixel b is the predicted pixel of partition block b. Therefore, filtering may be performed as in the examples of Figures 3701 and 3702. In addition, various filtering methods may be used.
[0177] Fig.383800 is a flowchart showing a process of encoding prediction information in an entropy encoding unit in a video encoding device. First, information about the activation of the skip mode is encoded (S3801). It is determined whether the skip mode is activated (S3802). When the skip mode is activated, the prediction mode is encoded (S3803). It is determined whether the prediction mode is an inter-frame prediction mode (S3804). When the prediction mode is not an inter-frame prediction mode, the intra-frame prediction information is encoded and the flowchart is terminated. When the skip mode is activated in step S3802, information about the activation of the DMVD mode is encoded (S3806). Here, the information about the activation of the DMVD mode indicates whether the DMVD mode is activated (bidirectional matching mode or template matching mode). For example, when the DMVD mode is not activated, the information is encoded as 0. When the DMVD mode is a bidirectional matching mode, the information is encoded as 10. When the DMVD mode is a template matching mode, the information is encoded as 11. It is determined whether the DMVD mode is activated (S3807). When the DMVD mode is activated (i.e., bidirectional matching mode or template matching mode), the template information is encoded (S3808). Here, the template information is encoded only when the template matching mode is a multi-template matching mode. In the multi-template matching mode, when a fixed template block is used, the template information indicates which template block's motion information is used to create a prediction block. When a motion boundary point is used, the template information indicates which template block's motion information in the template block partitioned based on the motion boundary point is used to create the current prediction block. The template information is encoded, and then the flowchart is terminated. When the DMVD mode is not activated, the merge candidate index information of the skip mode is encoded (S3809), and then the flowchart is terminated. When the prediction mode is an inter-frame prediction mode in step S3804, information about the activation of the merge mode is encoded (S3810). It is determined whether the merge mode is activated (S3811). When the merge mode is not activated, the prediction direction is encoded (S3812). The prediction direction may mean one of a past direction, a future direction, and a bidirectional direction. Determine whether the prediction direction is a future direction (S3813). When the prediction direction is not a future direction, encode past reference picture index information, past MVD information, and past MVP information (S3814, S3815, S3816). When the prediction direction is a future direction or after step S3816 is completed, determine whether the prediction direction is a past direction (S3817). When the prediction direction is not a past direction, encode future reference picture index information, future MVD information, and future MVP information (S3818, S3819, S3820). Then, terminate the flowchart. When the prediction direction is a past direction in step S3817, terminate the flowchart. When the merge mode is activated in step S3811, encode information about the activation of the DMVD mode (S3821).The description is the same as that of step S3806. It is determined whether the DMVD mode is activated (S3822). When the DMVD mode is activated, the merge candidate index information for the merge mode is encoded, and then the flowchart is terminated. When the DMVD mode is not activated in step S3822, the template information is encoded (S3823), and then the flowchart ends.
[0178] Fig.393900 is a flowchart showing a process of decoding prediction information in an entropy decoding unit in a video decoding device. First, information about activation of a skip mode is decoded (S3901). It is determined whether the skip mode is activated (S3902). When the skip mode is activated, the prediction mode is decoded (S3903). It is determined whether the prediction mode is an inter-frame prediction mode (S3904). When the prediction mode is not an inter-frame prediction mode, the intra-frame prediction information is decoded and the flowchart is terminated. When the skip mode is activated in step S3902, information about activation of the DMVD mode is decoded (S3906). Here, the information about activation of the DMVD mode indicates whether the DMVD mode is activated (bidirectional matching mode or template matching mode). For example, when the DMVD mode is not activated, the information is decoded as 0. When the DMVD mode is a bidirectional matching mode, the information is decoded as 10. When the DMVD mode is a template matching mode, the information is decoded as 11. It is determined whether the DMVD mode is activated (S3907). When the DMVD mode is activated (i.e., bidirectional matching mode or template matching mode), the template information is decoded (S3908) and the flowchart is terminated. When the DMVD mode is not activated, the merge candidate index information for the skip mode is decoded (S3909), and then the flowchart is terminated. When the prediction mode is the inter-frame prediction mode in step S3904, the information about the activation of the merge mode is decoded (S3910). It is determined whether the merge mode is activated (S3911). When the merge mode is not activated, the prediction direction is decoded (S3912). The prediction direction may mean one of the past direction, the future direction, and the bidirectional direction. It is determined whether the prediction direction is the future direction (S3913). When the prediction direction is not the future direction, the reference picture index information of the past direction, the MVD information of the past direction, and the MVP information of the past direction are decoded (S3914, S3915, S3916). When the prediction direction is the future direction or after step S3916 is completed, determine whether the prediction direction is the past direction (S3917). When the prediction direction is not the past direction, decode the reference picture index information of the future direction, the MVD information of the future direction, and the MVP information of the future direction (S3918, S3919, S3920). Then, terminate the flowchart. When the prediction direction is the past direction in step S3917, terminate the flowchart. When the merge mode is activated in step S3911, decode the information about the activation of the DMVD mode (S3921). The description is the same as the description of step S3906. Determine whether the DMVD mode is activated (S3922). When the DMVD mode is activated, decode the merge candidate index information for the merge mode, and then terminate the flowchart.When the DMVD mode is not activated in step S3922 , the template information is decoded ( S3923 ), and then the flowchart ends.
[0179] The initial motion information may be created based on information signaled from the encoding device. In the template matching mode, one item of initial motion information may be generated, and in the bidirectional matching mode, two items of initial motion information may be generated. The improvement process may determine the incremental motion information that minimizes the difference between the pixel values of the L0 reference block and the L1 reference block of the current block, and may improve the initial motion information based on the determined incremental motion information. Fig.43 A method for determining incremental motion information is described in detail.
[0180] The improved motion information (MV ref ) to perform motion compensation.
[0181] Improved motion information of the current block (MV ref ) may be stored in a buffer of a decoder and may be used as a motion information predictor of a block to be decoded after the current block (hereinafter referred to as a first block). The current block may be a neighboring block that is spatially / temporally adjacent to the first block. The current block may not be spatially / temporally adjacent to the first block, but may belong to the same CTU, tile, or tile group as the first block.
[0182] Optionally, the improved motion information (MV ref ) can be used only for motion compensation of the current block, but may not be stored in the decoder's buffer. That is, the improved motion information (MV ref ) may not be used as a motion information predictor for the first block.
[0183] Optionally, initial motion information (MV rec ) can be used for motion compensation of the current block. Improved motion information (MV ref ) can be stored in a buffer of the decoder and used as a motion information predictor for the first block.
[0184] A predetermined in-loop filter may be applied to the motion compensated current block. The in-loop filter may include at least one of a deblocking filter, a sample adaptive offset (SAO), or an adaptive loop filter (ALF). Improved motion information (MV ref ) can be used to determine the properties of the in-loop filter. Here, properties can mean boundary strength (bs), filtering strength, filter coefficients, number of filter taps, filter type, etc.
[0185] Fig.40 A method of performing intra prediction in a prediction unit of a video encoding / decoding device in an embodiment to which the present disclosure is applied is shown.
[0186] Reference Fig.40 , the method may determine an intra prediction mode of a current block (S4000).
[0187] The current block may be divided into a luminance block and a chrominance block. The intra prediction mode predefined in the decoding device may include a non-directional mode (planar mode, DC mode) and N directional modes. Here, the value of N may be an integer of 33, 65 or more. The intra prediction mode of the current block is determined for each of the luminance block and the chrominance block. This will be described in detail below.
[0188] The intra prediction mode of the luminance block can be derived based on a candidate list (MPM list) and a candidate index (MPM Idx). The candidate list includes multiple candidate modes. The candidate mode can be determined based on the intra prediction mode of the neighboring blocks of the current block.
[0189] The number of candidate modes is k. k may be an integer of 3, 4, 5, 6 or more. The number of candidate modes included in the candidate list may be a fixed number predefined in the encoding / decoding device, or may be changed based on the properties of the block. Here, the properties of the block may include size, width (w) or height (h), ratio between width and height, area, shape, partition type, partition depth, scanning order, encoding / decoding order, availability of intra-frame prediction mode, etc. The block may represent at least one of the current block or adjacent blocks. Optionally, the encoding device determines the number of optimal candidate modes, and then may encode the determined number into information and signal the information to the decoding device. The decoding device may determine the number of candidate modes based on the information signaled. In this case, the information may indicate the maximum / minimum number of candidate modes included in the candidate list.
[0190] Specifically, the candidate list may include at least one of an intra prediction mode (modeN), modeN-n, or modeN+n of a neighboring block, or a default mode.
[0191] For example, the neighboring block may include at least one of left (L), top (T), bottom left (BL), top left (TL), and top right (TR) neighboring blocks adjacent to the current block. Here, the value of n may be 1, 2, 3, 4, or a greater integer.
[0192] The default mode may include at least one of a planar mode, a DC mode, and a predetermined direction mode. The predetermined direction mode may include at least one of a vertical mode (modeV) and a horizontal mode (modeH), modeV-k, modeV+k, modeH-k, and modeH+k. Here, the value of k may be greater than or equal to 1, or may be any value of a natural number less than or equal to 15.
[0193] The candidate index may specify a candidate mode that is the same as the intra prediction mode of the luma block among the candidate modes of the candidate list. That is, the candidate mode specified by the candidate index may be set as the intra prediction mode of the luma block.
[0194] Optionally, the intra-frame prediction mode of the luminance block can be derived by applying a predetermined offset to the candidate mode specified by the candidate index. Here, the offset can be determined based on at least one of the size or shape of the luminance block or the value of the candidate mode specified by the candidate index. The offset can be determined as an integer of 0, 1, 2, 3 or more, or can be determined as the total number of intra-frame prediction modes or directional modes predefined in the decoding device. The intra-frame prediction mode of the luminance block can be derived by adding the offset to the candidate mode specified by the candidate index or subtracting the offset from the candidate mode specified by the candidate index. The application of the offset can be selectively performed in consideration of the properties of the block as described above.
[0195] Based on information (intra_chroma_pred_mode) signaled from the encoding device, the intra prediction mode of the chroma block may be derived as shown in Table 1 or 2 below.
[0196] [Table 1]
[0197] According to Table 1, the intra prediction mode of the chroma block may be determined based on the information transmitted by the signal, the intra prediction mode of the luminance block, and a table predefined in the decoding device. In Table 1, mode 66 means a diagonal mode in the upper right direction. Mode 50 means a vertical mode. Mode 18 means a horizontal mode. Mode 1 may mean a DC mode. For example, when the value of the information intra_chroma_pred_mode transmitted by the signal is 4, the intra prediction mode of the chroma block may be set to be the same as the intra prediction mode of the luminance block.
[0198] [Table 2]
[0199] When prediction based on inter-component references is allowed for chroma blocks, Table 2 may be applied. In this case, the intra prediction mode is derived in the same manner as in Table 1. A repeated description thereof will be omitted. However, Table 2 supports mode 81, mode 82, and mode 83 as intra prediction modes for chroma blocks, and they are prediction modes based on inter-component references.
[0200] Reference Fig.40 , the current block may be reconstructed based on the intra prediction mode derived in S4000 and at least one of the neighboring samples ( S4010 ).
[0201] Reconstruction can be performed in units of sub-blocks of the current block. To this end, the current block can be partitioned into multiple sub-blocks. For details on the partitioning method of the current block into sub-blocks, refer to Fig.41 .
[0202] For example, subblocks belonging to the current block may share a single derived intra prediction mode. In this case, different neighboring samples may be used for each of the subblocks. Alternatively, subblocks belonging to the current block may share a single candidate list as derived. Candidate indices may be determined independently for each subblock.
[0203] In one example, when the derived intra prediction mode corresponds to an inter-component reference-based prediction mode, the chrominance block may be predicted from a previously reconstructed luma block. Fig.42 This is described in detail.
[0204] Fig.41 The sub-block based intra prediction method in the embodiment to which the present disclosure is applied is shown.
[0205] As described above, the current block may be partitioned into a plurality of sub-blocks. Here, partitioning may be performed based on at least one of a quadtree (QT), a binary tree (BT), and a ternary tree (TT) as a partition type predefined in the decoding device. Optionally, the current block may correspond to a leaf node. A leaf node may mean a coding block that is no longer partitioned into smaller coding blocks. That is, partitioning the current block into sub-blocks may mean partitioning that is additionally performed after partitioning is finally performed based on a partition type predefined in the decoding device.
[0206] Reference Fig.41 , partitioning can be performed based on the size of the current block (Example 1).
[0207] Specifically, when the size of the current block 4100 is less than a predetermined threshold size, the current block may be partitioned into p sub-blocks in the vertical or horizontal direction. On the contrary, when the size of the current block 4110 is greater than or equal to the threshold size, the current block may be partitioned into q sub-blocks in the vertical or horizontal direction or may be partitioned into q sub-blocks in the vertical and horizontal directions. Here, p may be an integer of 1, 2, 3 or more. q may be an integer of 2, 3, 4 or more. However, p may be set to be less than q.
[0208] The threshold size may be signaled from the encoding device or may be a fixed value predefined in the decoding device. For example, the threshold size may be expressed as N×M, where each of N and M may be 4, 8, 16, or greater. N and M may be equal to or different from each other.
[0209] Optionally, when the size of the current block is smaller than a predefined threshold size, the current block may not be partitioned (non-split). Otherwise, the current block may be partitioned into 2, 4 or 8 sub-blocks.
[0210] In another example, partitioning may be performed based on the shape of the current block (Embodiment 2).
[0211] Specifically, when the shape of the current block is a square, the current block may be partitioned into 4 sub-blocks. Otherwise, the current block may be partitioned into 2 sub-blocks. Conversely, when the shape of the current block is a square, the current block may be partitioned into 2 sub-blocks. Otherwise, the current block may be partitioned into 4 sub-blocks.
[0212] Optionally, when the shape of the current block is a square, the current block may be partitioned into 2, 4, or 8 sub-blocks. Otherwise, the current block may not be partitioned. Conversely, when the shape of the current block is a square, the current block may not be partitioned. Otherwise, the current block may be partitioned into 2, 4, or 8 sub-blocks.
[0213] One of the above-mentioned embodiments 1 and 2 may be selectively applied. Alternatively, partitioning may be performed based on a combination of embodiments 1 and 2.
[0214] When partitioned into 2 sub-blocks, the current block may be partitioned into 2 sub-blocks in the vertical or horizontal direction. When partitioned into 4 sub-blocks, the current block may be partitioned into 4 sub-blocks in the vertical or horizontal direction, or may be partitioned into 4 sub-blocks in the vertical and horizontal directions. When partitioned into 8 sub-blocks, the current block may be partitioned into 8 sub-blocks in the vertical or horizontal direction, or may be partitioned into 8 sub-blocks in the vertical and horizontal directions.
[0215] The above embodiments illustrate the case where the current block is partitioned into 2, 4 or 8 sub-blocks. However, the present disclosure is not limited thereto. The current block may be partitioned into 3 sub-blocks in the vertical or horizontal direction. In this case, the ratio between the width or height of the 3 sub-blocks may be 1:1:2, 1:2:1 or 2:1:1.
[0216] Partition information about whether the current block is partitioned into sub-blocks, whether the current block is partitioned into 4 sub-blocks, the partition direction, and the number of partitions can be notified by a signal from the encoding device, and can be variably determined by the decoding device based on predetermined encoding parameters. Here, the encoding parameters may include block size / shape, partition type (4 partitions, 2 partitions, 3 partitions), intra-frame prediction mode, range / position of adjacent pixels for intra-frame prediction, component type (e.g., luminance, chrominance), maximum / minimum size of transform block, transform type (e.g., transform skip, DCT2, DST7, DCT8), etc. Based on the partition information, the current block may be partitioned into multiple sub-blocks or not partitioned.
[0217] Fig.42 The prediction method based on inter-component reference in the embodiment to which the present disclosure is applied is shown.
[0218] The current block can be divided into a luminance block and a chrominance block according to the component type. The pixels of the reconstructed luminance block can be used to predict the chrominance block. This is called inter-component reference. In this embodiment, it is assumed that the size of the chrominance block is (nTbW×nTbH) and the size of the luminance block corresponding to the chrominance block is (2 nTbW×2 That is, the input video sequence may be encoded in 4:2:0 color format.
[0219] Reference Fig.42 , a luma region for inter-component reference of a chroma block may be specified (S4200).
[0220] The luminance region may include at least one of a luminance block and a template block adjacent to the luminance block (hereinafter referred to as an adjacent region). Here, the luminance block may be defined as including pixels pY[x][y] (x=0.nTBW 2-1, y=0.nTBH 2-1). The pixel may mean a reconstructed value before applying the in-loop filter or a reconstructed value after applying the in-loop filter.
[0221] The adjacent region may include at least one of a left adjacent region, a top adjacent region, and an upper left adjacent region. The left adjacent region may be set to include pixels pY[x][y] (x=-1...-3, y=0.2 The top adjacent area may be set to include the pixel pY[x][y] (x=0.2 The setting may be performed only when the value of numSampT is greater than 0. The upper left adjacent area may be set to an area including at least one of the pixels pY[x][y] (x=-1...-3, y=-1, ...-3). The setting may be performed only when the upper left area of the luminance block is available.
[0222] The above variables numSampL and numSampT may be determined based on the intra prediction mode of the current block.
[0223] For example, when the intra prediction mode of the current block is INTRA_LT_CCLM, variables may be derived based on Equation 9. Here, INTRA_LT_CCLM may mean a mode in which inter-component referencing is performed based on a left neighboring region and a top neighboring region of the current block.
[0224] [Formula 9] numSampT = availT? nTbW: 0 numSampL = availL?nTbH: 0 According to Formula 9, when the top neighboring area of the current block is available, numSampT is derived as nTbW. Otherwise, numSampT may be derived as 0. Similarly, when the left neighboring area of the current block is available, numSampL is derived as nTbH. Otherwise, numSampL may be derived as 0.
[0225] In contrast, when the intra prediction mode of the current block is not INTRA_LT_CCLM, variables may be derived based on Equation 10 below.
[0226] [Formula 10] numSampT = (availT&&predModeIntra = = INTRA_T_CCLM)? (nTbW +numTopRight): 0 numSampL = (availL&&predModeIntra = = INTRA_L_CCLM)? (nTbH +numLeftBelow): 0 In Formula 10, INTRA_T_CCLM may refer to a mode for performing inter-component reference based on the top neighboring area of the current block. INTRA_L_CCLM may refer to a mode for performing inter-component reference based on the left neighboring area of the current block. numTopRight may refer to the number of all or some pixels belonging to the upper right neighboring area of the chroma block. Some pixels may refer to available pixels among the pixels of the lowest pixel row belonging to the corresponding area. In the availability determination, it may be determined sequentially in the left-to-right direction whether the pixel is available. This process may be performed until an unavailable pixel is found. numLeftBelow may refer to the number of all or some pixels belonging to the lower left neighboring area of the chroma block. Some pixels may refer to available pixels among the pixels of the rightmost pixel line (column) belonging to the corresponding area. Similarly, in the availability determination, it may be determined sequentially in the top-to-bottom direction whether the pixel is available. This process may be performed until an unavailable pixel is found.
[0227] The step of specifying a brightness region may further include down-sampling the specified brightness region.
[0228] Downsampling may include at least one of the following: 1. Downsampling of the luminance block; 2. Downsampling of the left adjacent area of the luminance block; 3. Downsampling of the top adjacent area of the luminance block. This will be described in detail below.
[0229] 1. Downsampling of Luma Blocks (Example 1) Based on the corresponding pixel pY[2 x][2 The pixel pDsY[x][y] (x=0..nTbW-1, y=0..nTbH-1) of the downsampled luminance block is derived from the adjacent pixels and the adjacent pixels. The adjacent pixels may refer to at least one of the left adjacent pixels, the right adjacent pixels, the top adjacent pixels, and the bottom adjacent pixels of the corresponding pixel. For example, the pixel pDsY[x][y] may be derived based on the following formula 11.
[0230] [Formula 11] pDsY[x][y]= (pY[ 2 x ][ 2 y-1] + pY[ 2 x-1 ][ 2 y] + 4 pY[ 2 x ][ 2 y] + pY[ 2 x + 1 ][ 2 y] + pY[ 2 x ][ 2 y + 1] + 4)>>3 However, there may be a situation where the left / top neighboring area of the current block is not available. When the left neighboring area of the current block is not available, the corresponding pixel pY[0][2 y] and the neighboring pixels of the corresponding pixel derive the pixel pDsY[0][y] (y=1...nTbH-1) of the downsampled luminance block. The neighboring pixel may mean at least one of the top neighboring pixel and the bottom neighboring pixel of the corresponding pixel. For example, the pixel pDsY[0][y] (y=1..nTbH-1) may be derived based on the following formula 12.
[0231] [Formula 12] pDsY[ 0 ][ y] = (pY[ 0 ][ 2 y-1] + 2 pY[ 0 ][ 2 y] + pY[ 0 ][ 2 y +1] + 2)>>2 When the top neighboring area of the current block is not available, the corresponding pixel pY[2 The pixel pDsY[x][0] (x=1..nTBW-1) of the downsampled luminance block is derived from the neighboring pixels of the corresponding pixel. The neighboring pixel may refer to at least one of the left neighboring pixel and the right neighboring pixel of the corresponding pixel. For example, the pixel pDsY[x][0] (x=1..nTbW-1) may be derived based on the following formula 13.
[0232] [Formula 13] pDsY[ x ][ 0] = (pY[ 2 x-1][0]+2 pY[ 2 x ][ 0]+ pY[ 2 x + 1 ][0]+ 2)>>2 The pixel pDsY[0][0] of the downsampled luma block may be derived based on the corresponding pixel pY[0][0] of the luma block and / or the neighboring pixels of the corresponding pixel. The position of the neighboring pixels may vary depending on whether the left / top neighboring area of the current block is available.
[0233] For example, when the left adjacent region is available but the top adjacent region is not available, pDsY[0][0] may be derived based on the following Equation 14.
[0234] [Formula 14] pDsY[ 0 ][ 0]= (pY[ -1 ][ 0]+ 2 pY[ 0 ][ 0]+ pY[ 1 ][ 0]+ 2)>>2 On the contrary, when the left adjacent region is not available but the top adjacent region is available, pDsY[0][0] may be derived based on the following Equation 15.
[0235] [Formula 15] pDsY[ 0 ][ 0]= (pY[ 0 ][ -1]+ 2 pY[ 0 ][ 0]+ pY[ 0 ][ 1]+ 2)>>2 In another example, when both the left neighboring area and the top neighboring area are not available, pDsY[0][0] may be set to the corresponding pixel pY[0][0] of the luma block.
[0236] (Example 2) Based on the corresponding pixel pY[2 x][2 The pixel pDsY[x][y] (x=0..nTbW-1, y=0..nTbH-1) of the downsampled luminance block is derived based on the neighboring pixels of the corresponding pixel and the neighboring pixels of the corresponding pixel. The neighboring pixels may mean at least one of the bottom neighboring pixels, the left neighboring pixels, the right neighboring pixels, the lower left neighboring pixels, and the lower right neighboring pixels of the corresponding pixel. For example, the pixel pDsY[x][y] may be derived based on the following formula 16.
[0237] [Formula 16] pDsY[ x ][ y]= (pY[ 2 x-1 ][ 2 y] + pY[ 2 x-1 ][ 2 y + 1] + 2 pY[ 2 x ][ 2 y] + 2 pY[ 2 x ][ 2 y + 1] + pY[ 2 x + 1 ][ 2 y] + pY[ 2 x+ 1 ][ 2 y + 1] + 4 )>>3 However, when the left neighboring area of the current block is not available, the corresponding pixel pY[0][2 y] and its bottom neighboring pixels derive the pixel pDsY[0][y] (y=0..nTBH-1) of the downsampled luminance block. For example, the pixel pDsY[0][y] (y=0..nTBH-1) can be derived based on the following equation 17.
[0238] [Formula 17] pDsY[ 0 ][ y] = (pY[ 0 ][ 2 y] + pY[ 0 ][ 2 y + 1] + 1)>>1 Downsampling of the luma block may be performed based on one of the embodiments 1 and 2 as described above. Here, one of the embodiments 1 and 2 may be selectively performed based on a predetermined flag. The flag may indicate whether the downsampled luma pixel has the same position as the original luma pixel. For example, when the flag is a first value, the downsampled luma pixel has the same position as the original luma pixel. On the contrary, when the flag is a second value, the downsampled luma pixel has the same position as the original luma pixel in the horizontal direction, but has a position shifted by half a pixel in the vertical direction.
[0239] 2. Downsampling of the left adjacent area of the luminance block (Example 1) Based on the corresponding pixel pY[-2][2 y] and the neighboring pixels of the corresponding pixel derive the pixel pLeftDsY [y] (y=0..numSampL-1) of the downsampled left adjacent area. The neighboring pixel may mean at least one of the left adjacent pixel, the right adjacent pixel, the top adjacent pixel, and the bottom adjacent pixel of the corresponding pixel. For example, the pixel pLeftDsY [y] may be derived based on the following formula 18.
[0240] [Formula 18] pLeftDsY[ y] = (pY[ -2 ][ 2 y-1] + pY[ -3 ][ 2 y] + 4 pY[ -2 ][ 2 y] + pY[ -1 ][ 2 y] + pY[ -2 ][ 2 y + 1] + 4)>>3 However, when the upper left neighboring area of the current block is not available, the pixel pLeftDsY[0] of the downsampled left neighboring area may be derived based on the corresponding pixel pY[-2][0] of the left neighboring area and the neighboring pixel of the corresponding pixel. The neighboring pixel may mean at least one of the left neighboring pixel and the right neighboring pixel of the corresponding pixel. For example, the pixel pLeftDsY[0] may be derived based on the following Equation 19.
[0241] [Formula 19] pLeftDsY[ 0] = (pY[ -3 ][ 0] + 2 pY[ -2 ][ 0]+ pY[ -1 ][ 0]+ 2)>>2 (Example 2) Based on the corresponding pixel pY[-2][2 y] and the neighboring pixels around the corresponding pixel derive the pixel pLeftDsY[y] (y=0..numSampL-1) of the downsampled left neighboring area. The neighboring pixel may mean at least one of the bottom neighboring pixel, the left neighboring pixel, the right neighboring pixel, the lower left neighboring pixel, and the lower right neighboring pixel of the corresponding pixel. For example, the pixel pLeftDsY[y] may be derived based on the following formula 20.
[0242] [Formula 20] pLeftDsY[ y] = (pY[ -1 ][ 2 y] + pY[ -1 ][ 2 y + 1] + 2 pY[ -2 ][2 y] + 2 pY[ -2] [2 y + 1] + pY[ -3 ][ 2 y] + pY[ -3 ][ 2 y + 1] + 4)>>3 Similarly, downsampling of the left adjacent area may be performed based on one of the embodiments 1 and 2 described above. Here, one of the embodiments 1 and 2 may be selected based on a predetermined flag. The flag indicates whether the downsampled luminance pixel has the same position as the original luminance pixel. This is the same as described above.
[0243] Downsampling of the left neighboring area may be performed only when the numSampL value is greater than 0. When the numSampL value is greater than 0, it may mean that the left neighboring area of the current block is available, and the intra prediction mode of the current block is INTRA_LT_CCLM or INTRA_L_CCLM.
[0244] 3. Downsampling of the top adjacent area of the luminance block (Example 1) The pixels pTopDsY[x] (x=0..numSampT-1) of the downsampled top neighboring region may be derived considering whether the top neighboring region belongs to a CTU different from the CTU to which the luma block belongs.
[0245] When the top neighboring region belongs to the same CTU as the CTU of the luma block, the corresponding pixel pY[2 The pixel pTopDsY[x] of the top adjacent area of the downsampled image is derived from the adjacent pixels of the corresponding pixel. The adjacent pixel may refer to at least one of the left adjacent pixel, the right adjacent pixel, the top adjacent pixel, and the bottom adjacent pixel of the corresponding pixel. For example, the pixel pTopDsY[x] may be derived based on the following formula 21.
[0246] [Formula 21] pTopDsY[ x] = (pY[ 2 x ][ -3]+ pY[ 2 x-1][-2]+ 4 pY[ 2 x ][ -2]+pY[ 2 x + 1 ][ -2]+ pY[ 2 x ][ -1]+ 4)>>3 On the contrary, when the top neighboring region belongs to a CTU different from the luma block, the corresponding pixel pY[2 The pixel pTopDsY[x] of the top adjacent area of the downsampled image is derived from the adjacent pixels of the corresponding pixel. The adjacent pixel may refer to at least one of the left adjacent pixel and the right adjacent pixel of the corresponding pixel. For example, the pixel pTopDsY[x] may be derived based on the following formula 22.
[0247] [Formula 22] pTopDsY[ x] = (pY[ 2 x-1][-1]+2 pY[ 2 x ][ -1]+ pY[ 2 x + 1 ][ -1] + 2)>>2 Alternatively, when the upper left neighboring area of the current block is not available, the neighboring pixel may mean at least one of the top neighboring pixel and the bottom neighboring pixel of the corresponding pixel. For example, the pixel pTopDsY[0] may be derived based on the following formula 23.
[0248] [Formula 23] pTopDsY[ 0] = (pY[ 0 ][ -3] + 2 pY[ 0 ][ -2]+ pY[ 0 ][ -1]+ 2)>>2 Optionally, when the top left neighboring region of the current block is not available and the top neighboring region belongs to a CTU different from the CTU of the luma block, the pixel pTopDsY[0] may be set to the pixel pY[0][-1] of the top neighboring region.
[0249] (Example 2) The pixels pTopDsY[x] (x=0..numSampT-1) of the downsampled top neighboring region may be derived taking into account whether the top neighboring region belongs to a CTU different from the CTU of the luma block.
[0250] When the top neighboring region belongs to the same CTU as the CTU of the luma block, the corresponding pixel pY[2 The pixel pTopDsY[x] of the top adjacent area of the downsampled image is derived from the adjacent pixels of the corresponding pixel. The adjacent pixel may refer to at least one of the bottom adjacent pixel, the left adjacent pixel, the right adjacent pixel, the lower left adjacent pixel, and the lower right adjacent pixel of the corresponding pixel. For example, the pixel pTopDsY[x] may be derived based on the following formula 24.
[0251] [Equation 24] pTopDsY[ x] = (pY[ 2 x-1 ][ -2]+ pY[ 2 x-1][-1]+2 pY[ 2 x ][ -2]+ 2 pY[ 2 x ][ -1]+ pY[ 2 x + 1 ][ -2]+ pY[ 2 x + 1 ][ -1] + 4) >> 3 On the contrary, when the top neighboring region belongs to a CTU different from the CTU of the luma block, the corresponding pixel pY[2 The pixel pTopDsY[x] of the top adjacent area of the downsampled image is derived from the adjacent pixels of the corresponding pixel. The adjacent pixel may refer to at least one of the left adjacent pixel and the right adjacent pixel of the corresponding pixel. For example, the pixel pTopDsY[x] may be derived based on the following formula 25.
[0252] [Formula 25] pTopDsY[ x] = (pY[ 2 x-1][-1]+2 pY[ 2 x ][ -1]+ pY[ 2 x + 1 ][ -1] + 2)>>2 Alternatively, when the upper left neighboring area of the current block is not available, the neighboring pixel may mean at least one of the top neighboring pixel and the bottom neighboring pixel of the corresponding pixel. For example, the pixel pTopDsY[0] may be derived based on the following formula 26.
[0253] [Equation 26] pTopDsY[ 0] = (pY[ 0 ][ -2] + pY[ 0 ][ -1] + 1)>>1 Optionally, when the top left neighboring region of the current block is not available and the top neighboring region belongs to a CTU different from the CTU of the luma block, the pixel pTopDsY[0] may be set to the pixel pY[0][-1] of the top neighboring region.
[0254] In a similar manner, downsampling of the top adjacent region may be performed based on one of the embodiments 1 and 2 described above. Here, one of the embodiments 1 and 2 may be selected based on a predetermined flag. The flag indicates whether the downsampled luminance pixel has the same position as the original luminance pixel. This is the same as described above.
[0255] In one example, downsampling of the top neighboring region may be performed only when the numSampT value is greater than 0. When the numSampT value is greater than 0, it may mean that the top neighboring region of the current block is available, and the intra prediction mode of the current block is INTRA_LT_CCLM or INTRA_T_CCLM.
[0256] Parameters for inter-component reference of a chroma block may be derived ( S4210 ).
[0257] The parameters may be determined in consideration of the intra prediction mode of the current block. That is, the relevant parameters may be derived in a variably manner according to whether the intra prediction mode of the current block is INTRA_LT_CCLM, INTRA_L_CCLM, or INTRA_T_CCLM.
[0258] The parameter may be calculated based on a linear change between an adjacent region of a luma block and an adjacent region of a chroma block. The parameter may be derived using at least one of pixels of a luma region and pixels of a top / left adjacent region of a chroma block. Here, the luma region may include at least one of a luma block and a top / left adjacent region of the luma block. The luma region may mean an area to which the aforementioned downsampling is applied.
[0259] The chroma block may be predicted based on the luma region and the parameter (S4220).
[0260] For example, the parameters may be applied to pixels of a downsampled luma block to generate a predicted block for a chroma block.
[0261] Fig.43 A method for determining incremental motion information in an embodiment to which the present disclosure is applied is shown.
[0262] For the convenience of description, it is assumed that the initial value of the incremental motion information is set to 0 in the decoder.
[0263] The decoder may determine the optimal reference block by performing a search operation based on the position of the reference block of the current block. Here, the reference block for the current block may mean an area indicated by the initial motion information. The decoder may update the incremental motion information based on the position of the optimal reference block. Here, the reference block may include an L0 / L1 reference block. The L0 reference block may refer to a reference block of a reference picture belonging to reference picture list 0 (L0), and the L1 reference block may mean a reference block of a reference picture belonging to reference picture list 1 (L1).
[0264] Reference Fig.43 , the decoder may determine a search area for improving motion information (S4300).
[0265] The search area may be determined as an area including at least one of a reference block and an adjacent area of the reference block. In this case, the position of the upper left sample of the reference block may serve as a reference position for the search. The search area may be determined in each of the L0 direction and the L1 direction. The adjacent area may mean an area extending N sample lines from the boundary of the reference block. Here, N may be an integer of 1, 2, 3 or more.
[0266] The neighboring area may be located in at least one direction of the left, top, right, bottom, top left, bottom left, top right, and bottom right of the reference block. Here, when the current block is W×H, the search area may be expressed as (W+2N)×(H+2N). However, in order to reduce the complexity of the improvement process, the neighboring area may be limited to only some of the above directions. For example, the neighboring area may be limited to the left and top neighboring areas of the reference block, or may be limited to the right and bottom neighboring areas of the reference block.
[0267] The number of sample lines (N) may be a fixed value predefined in the decoder, or may be variably determined in consideration of the properties of the block. Here, the properties of the block may refer to block size / shape, block position, inter-frame prediction mode, component type, etc. The block position may refer to whether the reference block is located on the boundary of the picture or on a predetermined fragment area. The fragment area may refer to a tile, a coding tree block column / row, or a coding tree block. For example, based on the properties of the block, the number of sample lines may be selectively determined to be 0, 1, or 2.
[0268] Reference Fig.43 , the decoder may determine a sum of absolute differences (SAD) at each search position within the search area ( S4310 ).
[0269] Specifically, a list of SADs at each search position in a specified search area (hereinafter referred to as a SAD list) may be determined. Here, the SAD list may include multiple SAD candidates. The number of SAD candidates constituting the SAD list may be M. M may be an integer greater than or equal to 2. The maximum value of M may be limited to 9.
[0270] The SAD candidate may be determined based on the SAD value between the L0 block and the L1 block. Here, the SAD value may be calculated based on all samples belonging to the L0 block and the L1 block, or may be calculated based on some of the samples belonging to the L0 block and the L1 block. Here, some samples refer to subblocks of each of the L0 block and the L1 block. In this case, at least one of the width or height of the subblock may be 1 / 2 of the width or height of the L0 / L1 block. That is, each of the L0 block and the L1 block may have a size of W×H, and the some samples may refer to subblocks having a size of W×H / 2, W / 2×H, or W / 2×H / 2. Here, when the some samples have a size of W×H / 2, the some samples may refer to the top subblock (or bottom subblock) in each of the L0 block and the L1 block. When the size of the some samples is W / 2×H, the some samples may refer to the left subblock (or right subblock) in each of the L0 block and the L1 block. When the size of the some samples is W / 2×H / 2, the some samples may refer to the upper left sub-block in each of the L0 block and the L1 block. However, the present disclosure is not limited thereto. Optionally, the some samples may be defined as a group of even or odd sample lines (vertical or horizontal direction) of each of the L0 block and the L1 block.
[0271] The position of the L0 block may be determined based on the position of the L0 reference block of the current block and a predetermined offset. The offset may refer to a disparity vector between the position of the L0 reference block and the search position. That is, the search position may be a position shifted by p in the x-axis direction and by q in the y-axis direction from the position (x0, y0) of the L0 reference block. Here, each of p and q may be at least one of -1, 0, and 1. Here, the disparity vector created based on the combination of p and q may refer to an offset. The position of the L0 block may be determined as a position shifted by (p, q) from the position (x0, y0) of the L0 reference block. The size (or absolute value) of each of p and q is 0 or 1. However, the present disclosure is not limited thereto. For example, the size of each of p and q may be 2, 3, or more.
[0272] The offset may include at least one of a non-directional offset (0, 0) and a directional offset. The directional offset may include an offset in at least one of left, right, top, bottom, top left, top right, bottom left, and bottom right. For example, the directional offset may include at least one of (-1, 0), (0, 1), (0, -1), (0, 1), (-1, -1), (-1, 1), (1, -1), and (1, 1).
[0273] In the same manner, the position of the L1 block may be determined based on the position of the L1 reference block of the current block and a predetermined offset. Here, the offset of the L1 block may be determined based on the offset of the L0 block. For example, when the offset of the L0 block is (p, q), the offset of the L1 block may be determined as (-p, -q).
[0274] Information about the size and / or direction of the offset as described above may be predefined in the decoder, or may be encoded by the encoder and may be signaled to the decoder. The information may be variably determined in consideration of the properties of the block as described above.
[0275] In an example, the offset may be defined as shown in Table 3 below.
[0276] [Table 3]
[0277] Table 3 defines the offset for determining the search position at each index i. However, Table 3 does not limit the position of the offset corresponding to index i. Instead, the position of the offset at each index may be different from those in Table 3. The offset according to Table 3 may include a non-directional offset (0, 0) and an offset in eight directions as described above.
[0278] In this case, the 0th SAD candidate may be determined based on the position (x, y) and offset (-1, -1) of the reference block. Specifically, a position offset by the offset (-1, -1) from the position (x0, y0) of the L0 reference block may be set as the search position. A W×H block including the search position as the upper left sample may be determined as the L0 block.
[0279] Similarly, a position offset by an offset (1, 1) from a position (x1, y1) of the L1 reference block may be set as a search position. A W×H block including the search position as an upper left sample may be determined as an L1 block. The 0th SAD candidate may be determined by calculating the SAD between the L0 block and the L1 block.
[0280] Through the above process, the first to eighth SAD candidates may be determined, and a SAD list including 9 SAD candidates may be determined.
[0281] Table 3 does not limit the number of offsets used to improve motion information. Only k offsets out of the 9 offsets may be used. Here, k may be any value from 2 to 8. For example, in Table 3, three offsets such as [0, 4, 8], [1, 4, 7], [2, 4, 6], [3, 4, 5], etc. may be used, or four offsets such as [0, 1, 3, 4], [4, 5, 7, 8], etc. may be used, or six offsets such as [0, 1, 3, 4, 6, 7,], [0, 1, 2, 3, 4, 5,], etc. may be used.
[0282] In an example, the offset may be defined as shown in the following Table 4. That is, the offset may consist of only a non-directional offset (0, 0), offsets (-1, 0) and (1, 0) in the horizontal direction, and offsets (0, -1) and (0, 1) in the vertical direction.
[0283] [Table 4]
[0284] In this manner, the size and / or shape of the search area as described above may be variably determined based on the size and / or number of offsets.The search area may have a geometric shape other than a square shape.
[0285] Reference Fig.43 , the incremental motion information may be updated based on the SAD at each search position (S4320).
[0286] Specifically, one of the first method and the second method for improving motion information about the current block may be selected based on the SAD candidates of the SAD list as described above.
[0287] The selection may be performed based on a comparison result between the reference SAD candidate and a predetermined threshold value (Embodiment 1).
[0288] The reference SAD candidate may mean a SAD candidate corresponding to an offset (0, 0). Alternatively, the reference SAD candidate may mean a SAD candidate corresponding to a position of a reference block or a reference position changed using a first method to be described later.
[0289] The threshold value may be determined based on at least one of the width (W) and the height (H) of the current block or the reference block. For example, the threshold value may be determined as W H.W. (H / 2), (W / 2) H, 2 W H, 4 W H et al.
[0290] When the reference SAD candidate is greater than or equal to the threshold, the incremental motion information may be updated based on the first method. Otherwise, the incremental motion information may be updated based on the second method.
[0291] (1) Regarding the first method The SAD candidate having the minimum value among the SAD candidates in the SAD list may be identified. The identification method will be described later. The incremental motion information may be updated based on the offset corresponding to the identified SAD candidate.
[0292] Then, the search position corresponding to the identified SAD candidate can be changed to a reference position for searching. The SAD list determination and / or identification of the SAD candidate with the minimum value as described above can be re-executed based on the changed reference position. The repeated description thereof will be omitted. Then, the incremental motion information can be updated based on the result of the re-execution.
[0293] (2) Regarding the second method The incremental motion information may be updated based on parameters calculated using all or some of the SAD candidates in the SAD list.
[0294] The parameter may include at least one of a parameter dX related to the x-axis direction and a parameter dY related to the y-axis direction. For example, the parameter may be calculated as follows. For ease of description, the calculation will be described based on Table 3 as described above.
[0295] When the sum of sadList[3] and sadList[5] is equal to (sadList[4]<<1), dX may be set to 0. Here, sadList[3] may mean a SAD candidate corresponding to an offset (-1, 0), sadList[5] may mean a SAD candidate corresponding to an offset (1, 0), and sadList[4] may mean a SAD candidate corresponding to an offset (0, 0).
[0296] When the sum of sadList[3] and sadList[5] is greater than (sadList[4]<<1), dX may be determined based on the following Equation 27.
[0297] [Formula 27] dX=((sadList[3]sadList[5])<<3) / (sadList[3]+sadList[5](sadList[4]<<1)) When the sum of sadList[1] and sadList[7] is equal to (sadList[4]<<1), dY may be set to 0. Here, sadList[1] may refer to a SAD candidate corresponding to an offset (0, -1), sadList[7] may refer to a SAD candidate corresponding to an offset (0, 1), and sadList[4] may refer to a SAD candidate corresponding to an offset (0, 0).
[0298] When the sum of sadList[1] and sadList[7] is greater than (sadList[4]<<1), dY may be determined based on the following Equation 28.
[0299] [Equation 28] dY=((sadList[1]sadList[7])<<3) / (sadList[1]+sadList[7](sadList[4]<<1)) Based on the calculated parameters, the incremental motion information may be updated.
[0300] Alternatively, the above selection may be performed based on whether the reference SAD candidate is a SAD candidate having a minimum value among the SAD candidates included in the SAD list (Embodiment 2).
[0301] For example, when the SAD candidate with the minimum value among the SAD candidates in the SAD list is not the reference SAD candidate, the incremental motion information may be updated based on the first method. Otherwise, the incremental motion information may be updated based on the second method. The first method and the second method are the same as described above, so their detailed descriptions will be omitted.
[0302] Alternatively, the above selection may be performed based on a combination of the first method and the second method (Embodiment 3).
[0303] For example, when the reference SAD candidate is less than a threshold, the incremental motion information may be updated using the parameters according to the second method as described above.
[0304] However, when the reference SAD candidate is greater than or equal to the threshold, the SAD candidate with the minimum value among the SAD candidates in the SAD list may be identified. When the identified SAD candidate is the reference SAD candidate, the incremental motion information may be updated using the parameters according to the second method as described above. Conversely, when the identified SAD candidate is not the reference SAD candidate, the incremental motion information may be updated according to the first method as described above.
[0305] A method for identifying a SAD candidate having a minimum value will be described. The SAD candidates in the SAD list may be grouped into 2, 3 or more groups. Hereinafter, for ease of description, an example in which the SAD candidates are grouped into two groups will be described.
[0306] A plurality of SAD candidates may be grouped into a first group and a second group. Each group may include at least two SAD candidates among the SAD candidates in the SAD list. However, the group may be configured only so that the reference SAD candidate is not included therein. A minimum operation may be applied to each group to extract the SAD candidate with the minimum value in each group.
[0307] A SAD candidate having a minimum value among the SAD candidates extracted from the first group and the SAD candidates extracted from the second group (hereinafter referred to as a temporary SAD candidate) may be extracted again.
[0308] Then, the SAD candidate with the minimum value may be identified based on the comparison result between the temporary SAD candidate and the reference SAD candidate. For example, when the temporary SAD candidate is smaller than the reference SAD candidate, the temporary SAD candidate may be identified as the minimum SAD candidate in the SAD list. Conversely, when the temporary SAD candidate is greater than or equal to the reference SAD candidate, the reference SAD candidate may be identified as the minimum SAD candidate in the SAD list.
[0309] The motion derivation method (in particular the motion information improvement method) in the decoder as described above may be performed on a sub-block basis and on a block size basis.
[0310] For example, when the size of the current block is greater than or equal to a predetermined threshold size, the motion information improvement method may be performed based on the threshold size. Conversely, when the size of the current block is less than the predetermined threshold size, the motion information improvement method may be performed based on the current block. Here, the threshold size may be determined such that at least one of the width and height of the block is a T value. The T value may be an integer of 16, 32, 64, or more.
[0311] The motion derivation methods (particularly the motion information improvement methods) in the decoder as described above may be performed in a limited manner based on block size, distance between current and reference pictures, inter prediction mode, prediction direction, or units or resolution of motion information.
[0312] For example, the motion information improvement method may be performed only when any one of the width or height of the current block is greater than or equal to 8, 16, or 32. Alternatively, the motion information improvement method may be applied only when the width and height of the current block are greater than or equal to 8, 16, or 32. Alternatively, the motion information improvement method may be applied only when the area of the current block or the number of samples thereof is greater than or equal to 64, 128, or 256.
[0313] The motion information improving method may be applied only when a difference between the POC of the current picture and the POC of the L0 reference picture and a difference between the POC of the current picture and the POC of the L1 reference picture are equal to each other.
[0314] This improvement method may be applied only when the inter prediction mode of the current block is the merge mode. When the inter prediction mode of the current block is other modes (eg, affine mode, AMVP mode, etc.), the motion information improvement method may not be applied.
[0315] The motion information improvement method can be applied only when the current block undergoes bidirectional prediction. In addition, the motion information improvement method can be applied only when the unit of the motion information is an integer pixel. The motion information improvement method can be applied only when the unit of the motion information is equal to or less than a quarter pixel or a half pixel.
[0316] The various embodiments of the present disclosure are not intended to list all possible combinations of the embodiments, but to illustrate representative aspects of the present disclosure. The features in the various embodiments may be applied independently or in combination of two or more thereof.
[0317] In addition, various embodiments of the present disclosure may be implemented in hardware, firmware, software, or a combination thereof. In a hardware-based implementation, at least one of an ASIC (application specific integrated circuit), a DSP (digital signal processor), a DSPD (digital signal processing device), a PLD (programmable logic device), an FPGA (field programmable gate array), a general purpose processor, a controller, a microcontroller, a microprocessor, etc. may be used.
[0318] The scope of the present disclosure may include software or machine-executable instructions (e.g., operating systems, applications, firmware, programs, etc.) that enable operations according to the methods of the present disclosure to be performed on a device or computer, as well as non-transitory computer-readable media in which such software or instructions executable on a device or computer are stored.
[0319] Industrial availability The present disclosure may be used to encode / decode video data.
Claims
1. A method for decoding video data using a decoding device, the method comprising: Receiving first partition information indicating whether to partition a current block into a plurality of sub-blocks from a bitstream, the first partition information being related to sub-block-based intra prediction; In response to the first division information indicating that the current block is divided into the plurality of sub-blocks, dividing the current block into the plurality of sub-blocks; obtaining the intra prediction mode of the current block based on the intra prediction mode of the current block signaled from the bitstream; as well as Based on the intra prediction mode of the current block, a prediction block of the current block is obtained by performing the sub-block-based intra prediction in units of sub-blocks, wherein, based on the second division information notified by the signal sent from the bitstream, the current block is divided into the plurality of sub-blocks in only one direction, vertically or horizontally; wherein the number of the plurality of sub-blocks obtained by dividing the current block is adaptively determined based on whether the size of the current block is smaller than a threshold size predefined in the decoding device, wherein, in response to the size of the current block being smaller than the threshold size, the number of the plurality of sub-blocks is determined to be equal to 2, wherein, in response to the size of the current block being greater than or equal to the threshold size, the number of the plurality of sub-blocks is determined to be equal to 4, The threshold size is defined as 8×8.
2. A method for encoding video data using an encoding device, the method comprising: Divide the current block into multiple sub-blocks; Based on the intra prediction mode of the current block, obtaining a prediction block of the current block by performing sub-block-based intra prediction in units of sub-blocks; as well as encoding first partition information indicating whether to partition the current block into the plurality of sub-blocks into a bitstream, the first partition information being related to the sub-block-based intra prediction, wherein prediction information about the intra prediction mode of the current block is encoded as a bitstream, The current block is divided into the plurality of sub-blocks only in one direction, vertically or horizontally. wherein, based on the partition direction of the current block, the second partition information is encoded into a bit stream, wherein the number of the plurality of sub-blocks obtained by dividing the current block is adaptively determined based on whether the size of the current block is smaller than a threshold size predefined in the decoding device, wherein, in response to the size of the current block being smaller than the threshold size, the number of the plurality of sub-blocks is determined to be equal to 2, wherein, in response to the size of the current block being greater than or equal to the threshold size, the number of the plurality of sub-blocks is determined to be equal to 4, The threshold size is defined as 8×8.
3. A method for transmitting a bit stream generated by a coding method, the coding method comprising: Divide the current block into multiple sub-blocks; Based on the intra prediction mode of the current block, obtaining a prediction block of the current block by performing sub-block-based intra prediction in units of sub-blocks; as well as encoding first partition information indicating whether to partition the current block into the plurality of sub-blocks into a bitstream, the first partition information being related to the sub-block-based intra prediction, wherein prediction information about the intra prediction mode of the current block is encoded as a bitstream, The current block is divided into the plurality of sub-blocks only in one direction, vertically or horizontally. wherein, based on the partition direction of the current block, the second partition information is encoded into a bit stream, wherein the number of the plurality of sub-blocks obtained by dividing the current block is adaptively determined based on whether the size of the current block is smaller than a threshold size predefined in the decoding device, wherein, in response to the size of the current block being smaller than the threshold size, the number of the plurality of sub-blocks is determined to be equal to 2, wherein, in response to the size of the current block being greater than or equal to the threshold size, the number of the plurality of sub-blocks is determined to be equal to 4, The threshold size is defined as 8×8.
4. A method for decoding video data using a decoding device, the method comprising: Determine initial motion information of the current block; Determining incremental motion information of the current block; improving the initial motion information of the current block by using the incremental motion information; as well as performing motion compensation for the current block by using the improved motion information, Wherein, determining the incremental motion information includes: determining a search area for improving the initial motion information; generating a sum of absolute differences (SAD) list from the search area, the SAD list comprising SAD candidates, each of the SAD candidates corresponding to each of the search positions in the search area; and updating the incremental motion information based on the SAD candidates in the SAD list, The SAD candidate is determined as the SAD value between the L0 block and the L1 block. wherein the SAD value is calculated based on samples of only the even-numbered sample lines in the L0 block and the L1 block, The current block is obtained by dividing the current picture based on division information notified by a signal sent from a bit stream.
5. A method for encoding video data using an encoding device, the method comprising: Determine initial motion information of the current block; Determining incremental motion information of the current block; improving the initial motion information of the current block by using the incremental motion information; as well as performing motion compensation for the current block by using the improved motion information, Wherein, determining the incremental motion information includes: determining a search area for improving the initial motion information; generating a sum of absolute differences (SAD) list from the search area, the SAD list comprising SAD candidates, each of the SAD candidates corresponding to each of the search positions in the search area; and updating the incremental motion information based on the SAD candidates in the SAD list, The SAD candidate is determined as the SAD value between the L0 block and the L1 block. wherein the SAD value is calculated based on samples of only the even-numbered sample lines in the L0 block and the L1 block, Wherein, based on the partition of the current picture used to obtain the current block, the partition information is encoded into a bit stream.
6. A method for transmitting a bit stream generated by a coding method, the coding method comprising: Determine initial motion information of the current block; Determining incremental motion information of the current block; improving the initial motion information of the current block by using the incremental motion information; as well as performing motion compensation for the current block by using the improved motion information, Wherein, determining the incremental motion information includes: determining a search area for improving the initial motion information; generating a sum of absolute differences (SAD) list from the search area, the SAD list comprising SAD candidates, each of the SAD candidates corresponding to each of the search positions in the search area; and updating the incremental motion information based on the SAD candidates in the SAD list, The SAD candidate is determined as the SAD value between the L0 block and the L1 block. wherein the SAD value is calculated based on samples of only the even-numbered sample lines in the L0 block and the L1 block, Wherein, based on the partition of the current picture used to obtain the current block, the partition information is encoded into a bit stream.