Decoding device and method for predicting block divided into arbitrary shape
By dividing video blocks into non-rectangular blocks and combining intra- and inter-frame prediction technologies, the shortcomings of existing video compression technologies in high-resolution and high-frame-rate video processing are solved, and more efficient encoding and decoding effects are achieved.
Patent Information
- Application Number
- CN202510268644.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-09
- Filing Date
- 2020-12-15
- Publication Date
- 2025-06-06
AI Technical Summary
Existing video compression technologies work hard to meet the growing demands when processing video data at high resolution and high frame rate.
By dividing the current block into non-rectangular blocks and performing intra prediction and inter prediction in these blocks, combined with adaptive aptive filtering processing, the efficiency of encoding and decoding is improved.
It achieves higher encoding efficiency and image quality, can effectively process high resolution and high frame rate video data, reducing the size of the bitstream.
Smart Images

Figure CN120111218A_ABST
Abstract
Description
[0001] This application is a divisional application of the PCT patent application entered into China with Chinese patent application number 202080086347.6, invention name “Decoding device and method for predicting blocks divided into arbitrary shapes”, and filing date December 15, 2020.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to Korean Patent Application No. 10-2019-0168016 filed on December 16, 2019 and Korean Patent Application No. 10-2020-0003143 filed on January 9, 2020, which are hereby incorporated by reference in their entirety. Technical Field
[0004] The present invention relates to encoding and decoding of a video, and more particularly, to a method and apparatus for further improving encoding and decoding efficiency by performing inter-frame prediction and intra-frame prediction on blocks divided in a given shape. Background Art
[0005] Since video data has a larger data volume than audio data or still image data, a large amount of hardware resources (including memory) are required to store or transmit the data in its original form before compression processing.
[0006] Accordingly, storing or transmitting video data is usually accompanied by compressing it by using an encoder before a decoder can receive, decompress and reproduce the compressed video data. Existing video compression technologies include H.264 / AVC and High Efficiency Video Coding (HEVC), which has an approximately 40% improvement in coding efficiency over H.264 / AVC.
[0007] However, the continuous increase in size, resolution and frame rate of video images and the resulting increase in the amount of data to be encoded require a new and superior compression technology with greater improvement in coding efficiency and higher improvement in image quality compared to existing compression technologies. Summary of the invention
[0008] To meet these needs, an object of the present invention is to provide an improved video encoding and decoding technology. Specifically, one aspect of the present invention relates to a technology for improving the efficiency of encoding and decoding by classifying non-rectangular blocks divided from a block into blocks for inter-frame prediction and blocks for intra-frame prediction.
[0009] Furthermore, another aspect of the present invention relates to a technique for improving the efficiency of encoding and decoding by simplifying an adaptive filtering process.
[0010] According to one aspect, the present invention provides a method for predicting a current block based on a first mode. The method includes: dividing the current block into non-rectangular blocks based on a partition mode syntax element; determining an intra-frame block for intra-frame prediction and an inter-frame block for inter-frame prediction in the non-rectangular block; and deriving a prediction sample of a first region including the inter-frame block based on motion information, and deriving a prediction sample of a second region including the intra-frame block based on the intra-frame prediction mode.
[0011] According to another aspect, the present invention provides a decoding device for predicting a current block based on a first mode. The device includes an entropy decoder and a predictor. The entropy decoder is configured to divide the current block into non-rectangular blocks based on a partition mode syntax element. The predictor is configured to determine an intra-frame block to be intra-predicted and an inter-frame block to be inter-predicted in the non-rectangular block, derive a prediction sample of a first region including the inter-frame block based on motion information, and derive a prediction sample of a second region including the intra-frame block based on the intra-frame prediction mode.
[0012] Compared with the conventional method that performs only inter-frame prediction, the present invention can further expand its applicability because intra-frame prediction can be performed on a non-rectangular block, not just inter-frame prediction.
[0013] Furthermore, the present invention can improve the performance of intra prediction because intra prediction of any non-rectangular block can be performed with reference to the inter prediction value of another non-rectangular block.
[0014] Furthermore, the present invention can effectively remove discontinuity occurring in a block edge by applying weights to an inter prediction value and an intra prediction value based on prediction types of neighboring blocks.
[0015] Furthermore, the present invention can improve bit efficiency because determination of an inter-block and an intra-block, whether to perform a mixing process, and whether to apply a deblocking filter can be determined based on a 1-bit flag.
[0016] In addition, the present invention can improve the efficiency of encoding and decoding because adaptive filtering can be simplified by integrating the feature extraction process of sample adaptive offset and the feature extraction process of adaptive loop filtering. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a block diagram showing a video encoding device capable of implementing the technology of the present invention.
[0018] Figure 2 is a schematic diagram for explaining a method of dividing a block by using a QTBTTT structure.
[0019] Figure 3a is a schematic diagram showing multiple intra prediction modes.
[0020] Figure 3b is a schematic diagram showing a plurality of intra prediction modes including a wide-angle intra prediction mode.
[0021] Figure 4 is a block diagram showing a video encoding device that can implement the technology of the present invention.
[0022] Figure 5 is an exemplary diagram for describing a triangle partitioning pattern.
[0023] Figure 6 is a flowchart for describing an example of a method of predicting a current block based on the first mode proposed in the present invention.
[0024] Figure 7 is an exemplary diagram for describing partition edges.
[0025] Figure 8 is an exemplary diagram of neighboring blocks used to determine whether to apply the first mode.
[0026] Fig. 9 is an exemplary diagram for describing an example of a method of determining an inter block and an intra block.
[0027] Fig.10 is an exemplary diagram for describing an example of a method of determining an intra prediction mode.
[0028] Fig.11 is a flowchart for describing an example of a method of using prediction samples of an inter block as reference samples for prediction of an intra block.
[0029] Fig.12 and Fig.13 Is used to describe Fig.11 An exemplary diagram of an example of a method.
[0030] Fig.14 is a flow chart for describing an example of a method of removing discontinuities in partition edges.
[0031] Fig.15 is used to describe the derivation Fig.14 An exemplary diagram of an example of a method of using weights in a method.
[0032] Fig.16 is an exemplary block diagram of a filtering unit.
[0033] Fig.17 is an exemplary block diagram of the filtering unit proposed in the present invention.
[0034] Fig.18 is a flowchart for describing an example of an adaptive filtering method.
[0035] Fig.19a and Fig.19b is an exemplary diagram for describing a primary difference filter and a secondary difference filter.
[0036] Fig.20a and Fig.20b is an exemplary diagram of partitioning the reconstructed samples in directions.
[0037] Fig.21 is an exemplary diagram of a reconstructed sample for describing a method of determining a category of the reconstructed sample. DETAILED DESCRIPTION
[0038] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying drawings. In the following description, the same reference numerals preferably represent the same elements, although the elements are shown in different drawings. In addition, in the following description of some embodiments, when it is considered to be obscure the subject of the present invention, for the sake of clarity and brevity, the specific description of the related known components and functions will be omitted.
[0039] Figure 1 1 is a block diagram showing a video encoding device capable of implementing the technology of the present invention. Figure 1 A video encoding apparatus and elements of the apparatus are described below.
[0040] The video encoding device includes: an image divider 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a filtering unit 180 and a memory 190.
[0041] Each element of the video encoding device may be implemented in hardware or software or a combination of hardware and software. The functions of each element may be implemented as software, and a microprocessor may be implemented to execute the software functions corresponding to each element.
[0042] A video includes multiple images. Each image is divided into multiple regions, and encoding is performed on each region. For example, an image is divided into one or more tiles (tile) or / and slices (slice). Here, one or more tiles can be defined as tile groups. Each tile or slice is divided into one or more coding tree units (coding treeunit, CTU). Each CTU is divided into one or more coding units (coding unit, CU) through a tree structure. The information applied to each CU is encoded as the syntax of the CU, and the information commonly applied to the CU included in a CTU is encoded as the syntax of the CTU. In addition, the information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and the information applied to all blocks constituting an image is encoded in the picture parameter set (Picture Parameter Set, PPS) or the picture header. In addition, the information commonly referenced by multiple images is encoded in the sequence parameter set (Sequence Parameter Set, SPS). In addition, the information commonly referenced by one or more SPSs is encoded in the video parameter set (VideoParameter Set, VPS). Information commonly applied to one tile or tile group may be encoded as the syntax of a tile header or tile group header.
[0043] The image divider 110 determines the size of a coding tree unit (CTU). Information on the size of the CTU (CTU size) is encoded as a syntax of an SPS or a PPS and transmitted to a video decoding apparatus.
[0044] The image divider 110 divides each image constituting a video into a plurality of CTUs having a predetermined size, and then recursively divides the CTUs using a tree structure. In the tree structure, a leaf node is used as a coding unit (CU), which is a basic unit of coding.
[0045] The tree structure can be a quadtree (QuadTree, QT), a binary tree (BinaryTree, BT), a ternary tree (TernaryTree, TT), or a structure formed by a combination of two or more QT structures, BT structures and TT structures, wherein the quadtree (QT) is a node (or parent node) that is divided into four slave nodes (or child nodes) of the same size, the binary tree (BT) is a node that is divided into two slave nodes, and the ternary tree (TT) is a node that is divided into three slave nodes at a ratio of 1:2:1. For example, a quadtree plus binary tree (QuadTree plus BinaryTree, QTBT) structure can be used, or a quadtree plus binary tree ternary tree (QuadTree plus BinaryTree TernaryTree, QTBTTT) structure can be used. Here, BTTT can be collectively referred to as a multiple-type tree (MTT).
[0046] Figure 2 The QTBTTT segmentation tree structure is shown as an example. Figure 2 As shown, the CTU can be first split into a QT structure. The QT splitting can be repeated until the size of the split block reaches the minimum block size MinQTSize of the leaf node allowed in the QT. The first flag (QT_split_flag) indicating whether each node of the QT structure is split into four nodes of the lower layer is encoded by the entropy encoder 155 and notified to the video decoding device by a signal. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, it can be further split into one or more BT structures or TT structures. The BT structure and / or the TT structure can have multiple splitting directions. For example, there can be two directions, namely, a direction of splitting the blocks of nodes horizontally and a direction of splitting the blocks vertically. As Figure 2 As shown, when MTT splitting starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is split, a flag indicating the splitting direction (vertical or horizontal) in the case of splitting, and / or a flag indicating the splitting type (binary or trifurcated), and notifies the video decoding device with a signal.
[0047] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into 4 nodes of the lower layer, the CU split flag (split_cu_flag) indicating whether the node is split may be encoded. When the value of the CU split flag (split_cu_flag) indicates that the split is not performed, the block of the node becomes a leaf node in the split tree structure and is used as a coding unit (CU), which is a basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that the split is performed, the video encoding device starts encoding the flags from the first flag in the above-described manner.
[0048] When QTBT is used as another example of a tree structure, there may be two types of segmentation, namely, a type of horizontally segmenting a block into two blocks of the same size (i.e., symmetrical horizontal segmentation) and a type of vertically segmenting a block into two blocks of the same size (i.e., symmetrical vertical segmentation). A segmentation flag (split_flag) indicating whether each node of the BT structure is segmented into blocks of the lower layer and segmentation type information indicating the segmentation type are encoded by the entropy encoder 155 and transmitted to the video decoding device. There may be an additional type of segmenting a block of a node into two asymmetric blocks. An asymmetric segmentation type may include a type of segmenting a block into two rectangular blocks at a size ratio of 1:3, or a type of segmenting a block of a node diagonally.
[0049] A CU may have various sizes according to the QTBT or QTBTTT partitioning of a CTU. Hereinafter, a block corresponding to a CU to be encoded or decoded (i.e., a leaf node of a QTBTTT) is referred to as a "current block". When QTBTTT partitioning is adopted, the shape of the current block may be a square or a rectangle.
[0050] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0051] Typically, each current block in an image can be predictively encoded. Typically, intra-frame prediction techniques (which utilize data from the image that includes the current block) or inter-frame prediction techniques (which utilize data from an image that was encoded before the image that includes the current block) are used. Inter-frame prediction includes both unidirectional prediction and bidirectional prediction.
[0052] The intra-frame predictor 122 predicts pixels in the current block using pixels (reference pixels) located around the current block in the current image including the current block. There are multiple intra-frame prediction modes according to the prediction direction. For example, Figure 3aAs shown, multiple intra prediction modes may include 2 non-directional modes and 65 directional modes, and the 2 non-directional modes include planar mode and DC mode. The neighboring pixels and equations to be used are defined differently for each prediction mode. The following table lists the intra prediction mode numbers and their names.
[0053] In order to perform effective directional prediction for the current block of rectangular shape, the additional use of Figure 3b The dotted arrows in the figure indicate the directional modes (intra-prediction modes 67 to 80 and -1 to -14). These modes may be referred to as "wide-angle intra-prediction modes". Figure 3b In the figure, the arrows indicate the corresponding reference samples used for prediction, rather than indicating the prediction direction. The prediction direction is opposite to the direction indicated by the arrow. The wide-angle intra prediction mode is a mode in which prediction is performed in a direction opposite to a specific direction mode without additional bit transmission when the current block is a rectangular shape. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes that can be used for the current block can be determined based on the ratio of the width to the height of the rectangular current block. For example, when the current block has a rectangular shape with a height less than the width, a wide-angle intra prediction mode (intra prediction modes 67 to 80) with an angle less than 45 degrees can be utilized. When the current block has a rectangular shape with a width greater than the height, a wide-angle intra prediction mode (intra prediction modes -1 to -14) with an angle greater than -135 degrees can be utilized.
[0054] The intra-frame predictor 122 may determine an intra-frame prediction mode to be used when encoding the current block. In some examples, the intra-frame predictor 122 may encode the current block using several intra-frame prediction modes and select an appropriate intra-frame prediction mode to be used from the tested modes. For example, the intra-frame predictor 122 may calculate a rate-distortion value using a rate-distortion analysis of several tested intra-frame prediction modes and may select an intra-frame prediction mode having the best rate-distortion characteristic from the tested modes.
[0055] The intra predictor 122 selects an intra prediction mode from a plurality of intra prediction modes and predicts the current block using adjacent pixels (reference pixels) and equations determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0056] The inter-frame predictor 124 generates a prediction block for the current block by motion compensation. The inter-frame predictor 124 searches for a block that is most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and uses the searched block to generate a prediction block for the current block. Then, the inter-frame predictor generates a motion vector corresponding to the displacement between the current block in the current image and the prediction block in the reference image. Typically, motion estimation is performed on the luminance (luma) component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. The motion information including information about the reference image and information about the motion vector used to predict the current block is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0057] The subtractor 130 subtracts the prediction block generated by the intra predictor 122 or the inter predictor 124 from the current block to generate a residual block.
[0058] The transformer 140 may divide the residual block into one or more transform blocks, perform a transform on the transform block, and transform the residual value of the transform block from the pixel domain to the frequency domain. In the frequency domain, the transform block is referred to as a coefficient block containing one or more transform coefficient values. A two-dimensional (2D) transform kernel may be used for the transform, and a one-dimensional (1D) transform kernel may be used for each of the horizontal transform and the vertical transform. The transform kernel may be based on a discrete cosine transform (DCT), a discrete sine transform (DST), etc.
[0059] The transformer 140 may transform the residual signal in the residual block by using the entire size of the residual block as a transform unit. In addition, the transformer 140 may divide the residual block into two sub-blocks in the horizontal direction or the vertical direction, and may perform the transform on only one of the two sub-blocks. Accordingly, the size of the transform block may be different from the size of the residual block (and thus the size of the prediction block). Non-zero residual sample values may not exist or be very sparse in the untransformed sub-block. The residual samples of the untransformed sub-block may not be signaled and may be regarded as "0" by the video decoding device. There may be several types of partitions according to the partition direction and the partition ratio. The transformer 140 may provide information about the encoding mode (or transform mode) of the residual block to the entropy encoder 155 (for example, the information about the encoding mode includes: information indicating whether to transform the residual block or the sub-block of the residual block, information indicating the partition type selected to divide the residual block into sub-blocks, and information for identifying the sub-block to be transformed, etc.). The entropy encoder 155 may encode information about the encoding mode (or transform mode) of the residual block.
[0060] The quantizer 145 quantizes the transform coefficient output from the transformer 140 and outputs the quantized transform coefficient to the entropy encoder 155. The quantizer 145 may directly quantize a related residual block for a specific block or frame without transformation.
[0061] The rearrangement unit 150 can perform rearrangement of coefficient values with quantized transform coefficients. The rearrangement unit 150 can use coefficient scanning to change the two-dimensional coefficient array into a one-dimensional coefficient sequence. For example, the rearrangement unit 150 can scan the coefficients from the DC coefficient to the coefficients in the high-frequency region by zig-zag scanning or diagonal scanning to output a one-dimensional coefficient sequence. Depending on the size of the transform unit and the intra-frame prediction mode, the zig-zag scan used can be replaced by a vertical scan for scanning the two-dimensional coefficient array in the column direction and a horizontal scan for scanning the two-dimensional block-shaped coefficients in the row direction. In other words, the scanning method to be used can be determined among zig-zag scanning, diagonal scanning, vertical scanning, and horizontal scanning according to the size of the transform unit and the intra-frame prediction mode.
[0062] The entropy encoder 155 encodes the sequence of one-dimensional quantized transform coefficients output from the rearrangement unit 150 by using various encoding methods such as Context-based Adaptive Binary Arithmetic Code (CABAC), Exponential Golomb, etc. to generate a bit stream.
[0063] The entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CU partition flag, QT partition flag, MTT partition type, and MTT partition direction) so that the video decoding device can partition the block in the same manner as the video encoding device. In addition, the entropy encoder 155 encodes information about a prediction type indicating whether the current block is encoded by intra-frame prediction or inter-frame prediction, and encodes intra-frame prediction information (i.e., information about an intra-frame prediction mode) or inter-frame prediction information (i.e., information about a reference image index and a motion vector) according to the prediction type.
[0064] The inverse quantizer 160 inversely quantizes the quantized transform coefficient output from the quantizer 145 to generate a transform coefficient. The inverse transformer 165 transforms the transform coefficient output from the inverse quantizer 160 from the frequency domain to the spatial domain, and reconstructs a residual block.
[0065] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. Pixels in the reconstructed current block are used as reference pixels when performing intra prediction of a subsequent block.
[0066] The filtering unit 180 filters the reconstructed pixels to reduce blocking artifacts, ringing artifacts, and blurring artifacts caused by block-based prediction and transform / quantization. The filtering unit 180 may include a deblocking filter 182 and a sample adaptive offset (SAO) filter 184.
[0067] The deblocking filter 180 filters the boundary between the reconstructed blocks to remove the block artifacts caused by block-by-block encoding / decoding, and the SAO filter 184 performs additional filtering on the deblocking filtered video. The SAO filter 184 is a filter for compensating for the difference between the reconstructed pixels and the original pixels caused by lossy coding.
[0068] The reconstructed blocks filtered by the deblocking filter 182 and the SAO filter 184 are stored in the memory 190. Once all blocks in one image are reconstructed, the reconstructed image can be used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.
[0069] Figure 4 is an exemplary functional block diagram of a video decoding device capable of implementing the technology of the present invention. Figure 4 To describe a video decoding apparatus and its components.
[0070] The video decoding apparatus may include an entropy decoder 410 , a rearrangement unit 415 , an inverse quantizer 420 , an inverse transformer 430 , a predictor 440 , an adder 450 , a filtering unit 460 , and a memory 470 .
[0071] Similar to Figure 1 In the video encoding device of the present invention, each element of the video decoding device can be implemented by hardware, software or a combination of hardware and software. In addition, the function of each element can be implemented by software, and the microprocessor can be implemented to execute the software function corresponding to each element.
[0072] The entropy decoder 410 determines a current block to be decoded by decoding a bit stream generated by a video encoding device and extracting information related to block partitioning, and extracts prediction information required to reconstruct the current block, information about a residual signal, and the like.
[0073] The entropy decoder 410 extracts information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), determines the size of the CTU, and divides the image into CTUs of the determined size. Then, the decoder determines the CTU as the highest level (i.e., the root node) of the tree structure, and extracts partition information about the CTU to partition the CTU using the tree structure.
[0074] For example, when the CTU is segmented using the QTBTTT structure, the first flag (QT_split_flag) related to the segmentation of the QT is extracted to segment each node into four nodes of the sublayer. For the node corresponding to the leaf node of the QT, the second flag (MTT_split_flag) related to the segmentation of the MTT and the information about the segmentation direction (vertical / horizontal) and / or the segmentation type (binary / tripartite) are extracted to segment the corresponding leaf node with the MTT structure. Thus, each node below the leaf node of the QT is recursively segmented with the BT or TT structure.
[0075] As another example, when a CTU is split using a QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split may be extracted. When the corresponding block is split, a first flag (QT_split_flag) may be extracted. In the split operation, after zero or more recursive QT splits, zero or more recursive MTT splits may occur per node. For example, a CTU may directly undergo MTT splitting without undergoing QT splitting, or may only undergo QT splitting multiple times.
[0076] As another example, when the CTU is segmented using the QTBT structure, the first flag (QT_split_flag) related to the QT segmentation is extracted, and each node is segmented into four nodes of the lower layer. Then, the segmentation flag (split_flag) indicating whether the node corresponding to the leaf node of the QT is further segmented with the BT and the segmentation direction information are extracted.
[0077] Once the current block to be decoded is determined by tree structure partitioning, the entropy decoder 410 extracts information about the prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoder 410 extracts syntax elements of intra-prediction information (intra-prediction mode) of the current block. When the prediction type information indicates inter-prediction, the entropy decoder 410 extracts syntax elements for inter-prediction information (i.e., information indicating a motion vector and a reference image referenced by the motion vector).
[0078] On the other hand, the entropy decoder 410 extracts information about the encoding mode of the residual block from the bitstream (for example, information about whether to encode the residual block or only encode a subblock of the residual block, information indicating a division type selected to divide the residual block into subblocks, information for identifying the encoded residual subblock, a quantization parameter, etc.). In addition, the entropy decoder 410 extracts information about the quantized transform coefficient of the current block as information about the residual signal.
[0079] The rearrangement unit 415 may change a sequence of quantized 1D transform coefficients entropy-decoded by the entropy decoder 410 into a 2D coefficient array (ie, block) in the reverse order of coefficient scanning performed by the video encoding apparatus.
[0080] The inverse quantizer 420 inversely quantizes the quantized transform coefficient, and the inverse transformer 430 generates a reconstructed residual block of the current block by reconstructing a residual signal by inversely transforming the inversely quantized transform coefficient from the frequency domain to the spatial domain based on information about the encoding mode of the residual block.
[0081] When the information about the encoding mode of the residual block indicates that the residual block of the current block is encoded in the video encoding device, the inverse transformer 430 uses the size of the current block (and therefore the size of the residual block to be restored) as a transformation unit to generate a reconstructed residual block of the current block by performing an inverse transformation on the inverse quantized transformation coefficients.
[0082] In addition, when the information about the encoding mode of the residual block indicates that only one sub-block of the residual block is encoded in the video encoding device, the inverse transformer 430 uses the size of the transformed sub-block as the transformation unit, performs inverse transformation on the inverse quantized transformation coefficient, reconstructs the residual signal of the transformed sub-block, and sets the residual signal of the untransformed sub-block to "0", thereby generating a reconstructed residual block of the current block.
[0083] The predictor 440 may include an intra predictor 442 and an inter predictor 444. The intra predictor 442 is activated when the prediction type of the current block is intra prediction, and the inter predictor 444 is activated when the prediction type of the current block is inter prediction.
[0084] The intra predictor 442 determines an intra prediction mode of a current block among a plurality of intra prediction modes based on a syntax element of the intra prediction mode extracted from the entropy decoder 410 , and predicts the current block using reference pixels around the current block according to the intra prediction mode.
[0085] The inter predictor 444 determines a motion vector of a current block and a reference image referred to by the motion vector using a syntax element of the inter prediction mode extracted from the entropy decoder 410 , and predicts the current block based on the motion vector and the reference image.
[0086] The adder 450 reconstructs the current block by adding the residual block output from the inverse transformer 430 and the prediction block output from the inter predictor 444 or the intra predictor 442. When intra-predicting a subsequent block to be decoded, pixels in the reconstructed current block are used as reference pixels.
[0087] The filtering unit 460 may include a deblocking filter 462 and an SAO filter 464. The deblocking filter 462 performs deblocking filtering on the boundaries between the reconstructed blocks to remove block artifacts caused by block-by-block decoding. The SAO filter 464 performs additional filtering on the reconstructed blocks after deblocking filtering the corresponding offsets to compensate for the difference between the reconstructed pixels and the original pixels caused by lossy encoding. The reconstructed blocks filtered by the deblocking filter 462 and the SAO filter 464 are stored in the memory 470. When all blocks in one image are reconstructed, the reconstructed image is used as a reference image for inter-frame prediction of blocks in subsequent images to be encoded.
[0088] The encoding device and the decoding device may divide the current block into blocks having a rectangular shape, and may divide the current block into blocks having a triangular shape, so as to improve prediction performance. To this end, the encoding device and the decoding device may also support a triangle partition mode (TPM).
[0089] In TPM, Figure 5 As shown, the encoding device divides the current block into two triangular blocks, and performs inter-frame prediction on each of the triangular blocks based on the merge mode. The encoding device indicates two division forms (division in the lower right diagonal direction ( Figure 5 (a)) and the division in the upper right diagonal direction ( Figure 5 The encoding device and the decoding device may further perform blending on the partition edge to remove discontinuity of sample values occurring in the partition edge.
[0090] However, as the image resolution gradually increases, the size of the prediction block also increases. Accordingly, there is a need for block partitioning with more different shapes (eg, a given shape or a non-rectangular shape) and applying intra prediction to the partitioned blocks.
[0091] The present invention proposes a method for dividing a current block into non-rectangular blocks and a method for applying intra prediction to non-rectangular blocks in addition to inter prediction. The method of the present invention can be expressed as GIIP mode (Geometric Partitioning for Intra and Inter Prediction) or the first mode.
[0092] Figure 6 A flow chart showing an example of the method proposed in the present invention.
[0093] 1. Block division
[0094] The encoding device may divide the current block into non-rectangular blocks, may encode a partition mode syntax element indicating a structure in which the non-rectangular blocks are divided, and may signal the partition mode syntax element to the decoding device. The entropy decoder 410 may decode the partition mode syntax element from the bitstream, and may divide the current block into non-rectangular blocks based on the decoded partition mode syntax element (S604).
[0095] exist Figure 7 In the example of , the vertical distance (d i ) and the angle formed by the partition edge and the horizontal direction of the current block (a i ) to determine the dividing edge or boundary. i Can have values between 0 and 360 degrees. max (That is, d i The maximum value of a i And the width (w) and height (h) of the current block are determined as shown in Equation 1.
[0096] [Equation 1]
[0097]
[0098] In equation 1, Indicates the angle from the horizontal axis to the edge of the block (a i ). The subtraction of b is used to prevent the corners of the current block and the block edge from becoming too close to each other.
[0099] The partition mode syntax element can directly indicate a i and d i Or it can be in a i and d i The index of the partition edge of the current block is indicated in the preset value of. If the partition mode syntax element is implemented as an index, the encoding device and the decoding device can use the i and d i A lookup table of preset values of the partition mode syntax element is used to determine a corresponding to the index indicated by the partition mode syntax element. i and di .
[0100] 2. Determine whether to execute GIIP mode
[0101] According to an embodiment, before performing process S604, process S602 of determining whether to apply the GIIP mode may be performed first. Whether to apply the GIIP mode may be determined based on syntax elements, prediction types of neighboring blocks, angles of partition edges, and the like.
[0102] As an example, the encoding device may determine whether the GIIP mode has been applied to the current block, and set the value of the application syntax element to a value corresponding to the determination result. The encoding device may encode the application syntax element, and may signal the application syntax element to the decoding device. The entropy decoder 410 may decode the application syntax element from the bitstream, and may determine whether to apply the GIIP mode based on the value of the application syntax element (S602).
[0103] The application syntax element may be a 1-bit flag. When the application syntax element == 0, the GIIP mode is not applied. In this case, as in the conventional TPM, inter-frame prediction using the merge index of each of the non-rectangular blocks may be performed. When the application syntax element == 1, the GIIP mode may be applied, inter-frame prediction using motion information may be performed on inter-frame blocks (blocks predicted based on the inter-frame prediction type in the non-rectangular blocks), and intra-frame prediction using the intra-frame prediction mode may be performed on intra-frame blocks (blocks predicted based on the intra-frame prediction type in the non-rectangular blocks).
[0104] As another example, the encoding apparatus and the decoding apparatus may determine whether to apply the GIIP mode based on the number of blocks predicted based on the intra prediction type in a neighboring block located adjacent to the current block.
[0105] exist Figure 8 In the example of , whether to apply the GIIP mode may be determined by considering the prediction types of the neighboring blocks A, B, C, and D or by considering the prediction types of the neighboring blocks A and B.
[0106] If prediction types A, B, C, and D are considered, when the number of neighboring blocks predicted by intra-frame is 3 or more, it can be determined that the GIIP mode is applied. When the number of neighboring blocks predicted by intra-frame is 2, it can be determined whether the GIIP mode is applied based on the value of the application syntax element. When the number of neighboring blocks predicted by intra-frame is 1 or less, it can be determined that the GIIP mode is not applied. According to an embodiment, when the number of neighboring blocks predicted by intra-frame is 2 or more or 3, it can be determined that the GIIP mode is applied. According to an embodiment, when the number of neighboring blocks is less than 2 or 3, it can be determined that the GIIP mode is not applied.
[0107] If the prediction types of adjacent blocks A and B are considered, when the number of adjacent blocks predicted by intra-frame is 2, it can be determined that the GIIP mode is applied. When the number of adjacent blocks predicted by intra-frame is 1, it can be determined whether the GIIP mode is applied based on the value of the application syntax element. When the number of adjacent blocks predicted by intra-frame is 0, it can be determined that the GIIP mode is not applied. According to an embodiment, when the number of adjacent blocks predicted by intra-frame is 1 or more or 2, it can be determined that the GIIP mode is applied. According to an embodiment, when the number of adjacent blocks predicted by intra-frame is less than 1 or 2, it can be determined that the GIIP mode is not applied.
[0108] 3. Determine inter-frame blocks and intra-frame blocks
[0109] Referring back to process S604, when the current block is divided into non-rectangular blocks, process S606 is performed to determine intra blocks to be intra predicted and inter blocks to be inter predicted in the non-rectangular blocks. Process S606 may be a process of determining the prediction type of the non-rectangular blocks.
[0110] The prediction type of the non-rectangular block may be 1) explicitly determined by a type syntax element indicating the prediction type of the non-rectangular block, or 2) may be implicitly determined by one or more of an angle of a partition edge and a vertical distance from the center of the current block to the partition edge (S608).
[0111] In the case of 1), the encoding device may set the prediction type of the non-rectangular block to the value of the type syntax element, encode the type syntax element, and signal the type syntax element to the decoding device. The decoding device may decode the type syntax element from the bitstream. The predictor 440 may determine the prediction type of the non-rectangular block based on the value of the type syntax element.
[0112] In case 2), the predictor 120, the predictor 440 may determine the prediction type of the non-rectangular block based on one or more of the angle of the partition edge (the angle formed by the partition edge and the horizontal direction of the current block) and the vertical distance (the vertical distance between the partition edge and the center).
[0113] like Fig. 9 As shown in (a), when the angle of the dividing edge is 0 to 90 degrees, the non-rectangular block located at the top (top block) is closer to the reference sample (sample marked with a pattern) for intra prediction than the non-rectangular block located at the bottom (bottom block). Accordingly, the top block can be determined as an intra block, and the bottom block can be determined as an inter block.
[0114] Alternatively or additionally, in Fig. 9In the case of (a), as the vertical distance from the partition edge to the center of the current block increases, the number of reference samples of the neighboring top block decreases, and the number of reference samples of the neighboring bottom block increases. If the vertical distance becomes very large and the number of reference samples of the neighboring top block is less than the number of reference samples of the neighboring bottom block, the top block can be determined as an inter-frame block, and the bottom block can be determined as an intra-frame block.
[0115] like Fig. 9 As shown in (b), if the angle of the dividing edge is 90 to 180 degrees, the non-rectangular block (block A) located on the left is closer to the reference sample than the non-rectangular block (block B) located on the right. Accordingly, in this case, block A can be determined as an intra-frame block, and block B can be determined as an inter-frame block.
[0116] Alternatively or additionally, in Fig. 9 In the case of (b), as the vertical distance from the dividing edge to the center increases, the number of reference samples of the neighboring left block (block A) decreases, and the number of reference samples of the neighboring right block (block B) increases. If the vertical distance becomes very large and the number of reference samples of the neighboring block A is less than the number of reference samples of the neighboring block B, block A may be determined as an inter-frame block, and block B may be determined as an intra-frame block.
[0117] like Fig. 9 As shown in (c), if the angle of the partition edge is 90 to 180 degrees and the partition edge is located above the center of the current block, the non-rectangular block (block B) located at the top is closer to the reference sample than the non-rectangular block (block A) located at the bottom. Accordingly, in this case, block B can be determined as an intra-frame block, and block A can be determined as an inter-frame block.
[0118] Alternatively or additionally, in Fig. 9 In the case of (c), as the vertical distance from the partition edge to the center increases, the number of reference samples of the neighboring block B decreases, and the number of reference samples of the neighboring block A increases. If the vertical distance becomes very large and the number of reference samples of the neighboring block B is less than the number of reference samples of the neighboring block A, the block B may be determined as an inter-frame block, and the block A may be determined as an intra-frame block.
[0119] 4. Derive prediction samples for non-rectangular blocks (prediction execution)
[0120] When the inter block and the intra block are determined, a process S610 is performed for deriving or generating a prediction sample for each of the inter block and the intra block.
[0121] The process of deriving the prediction sample of the inter-frame block may be performed based on the motion information. The encoding device may encode the motion information (e.g., merge index) used in the inter-frame prediction of the inter-frame block, and signal the motion information to the decoding device. The decoding device may decode the motion information from the bitstream. The inter-frame predictor 444 may derive the prediction sample of the first region by using the motion information.
[0122] The first region is a region including an inter block and may be included in the current block. Accordingly, the first region may be the same as the inter block, or may be a region including more samples located around a partition edge, or may be the same as the current block. If the first region is the same as the current block, inter prediction samples for the entire current block may be derived.
[0123] Determine intra prediction mode
[0124] The process of deriving the prediction sample of the intra block may be performed based on the intra prediction mode. The encoding device may encode information about the prediction mode (intra prediction mode) used in the intra prediction of the intra block, and may signal the information to the decoding device. The information about the intra prediction mode may be an index (mode index) indicating any one of the preset intra prediction mode candidates. The decoding device may decode the information about the intra prediction mode from the bitstream. The intra predictor 442 may derive the prediction sample of the second region by utilizing the intra prediction mode indicated by the decoded information.
[0125] The second region is a region including an intra block and may be included in the current block. Accordingly, the second region may be the same as the intra block, or may be a region including more samples around a partition edge, or may be the same as the current block. If the second region is the same as the current block, intra prediction samples for the entire current block may be derived.
[0126] According to an embodiment, the intra-frame prediction mode candidate may include one or more directional modes that are the same as or similar to the angle of the partition edge (corresponding to the angle of the partition edge). According to another embodiment, the intra-frame prediction mode candidate may further include a horizontal (horizontal, HOR) mode and a vertical (vertical, VER) mode in addition to the directional mode corresponding to the partition edge, and may further include a planar mode as a non-directional mode.
[0127] like Fig. 9 As shown in (a), if the directional modes corresponding to the angle of the partition edge are mode A and mode B, the intra prediction mode candidates may include directional mode A and directional mode B. In addition to directional mode A and directional mode B, the intra prediction mode candidates may further include one or more of planar mode, HOR mode, and VER mode.
[0128] like Fig. 9 As shown in (b), if the directional mode corresponding to the angle of the partition edge is mode A, the intra prediction mode candidates may include directional mode A. In addition to directional mode A, the intra prediction mode candidates may further include one or more of planar mode and HOR mode.
[0129] like Fig. 9 As shown in (c), if the directional mode corresponding to the angle of the partition edge is mode A, the intra prediction mode candidates may include directional mode A. In addition to directional mode A, the intra prediction mode candidates may further include one or more of planar mode and VER mode.
[0130] According to an embodiment, intra prediction mode candidates may include Figure 3b The 93 directional modes shown in FIG. 1 (except the planar mode and DC mode in the wide-angle intra prediction mode) are independent of the angle of the partition edge, or may include the following: Figure 3a 67 intra prediction modes are shown. According to another embodiment, the intra prediction mode candidates may include 35 intra prediction modes (including HOR directional mode, VER directional mode and planar mode).
[0131] In other embodiments, the intra prediction mode used in the prediction of the intra block is not explicitly indicated by the information about the intra prediction mode, but can be pre-agreed between the encoding device and the decoding device. For example, the intra prediction mode used in the prediction of the intra block can be pre-agreed to be any one of the intra prediction mode candidates, for example, pre-agreed to be a plane mode.
[0132] Derive prediction samples for intra blocks
[0133] The prediction samples of the intra block may be derived by referring only to samples located around the current block (first type), or may be derived by referring to both samples located around the current block and the prediction samples of the inter block (second type).
[0134] Whether to perform the second type can be explicitly indicated by a 1-bit flag, or it can be pre-agreed that the second type is always performed when the GIIP mode is applied (implicit). If it is explicitly indicated whether to perform the second type, the encoding device can determine whether to perform the second type, and can set its result to the value of the 1-bit flag. The encoding device can encode the 1-bit flag, and can signal the 1-bit flag to the decoding device. The entropy decoder 410 can decode the 1-bit flag from the bitstream.
[0135] If the second type is performed, the inter-frame predictor 124, the inter-frame predictor 444 may derive the prediction sample of the inter-frame block by performing a motion compensation process based on the motion information (S1102). The intra-frame predictor 122, the intra-frame predictor 442 may refer to both the samples located around the current block and the prediction samples of the inter-frame block to derive the prediction sample of the intra-frame block (S1104). The intra-frame prediction mode candidates for the second type may include the following: Fig.12 The example covers 360 degrees of directional mode, planar mode and DC mode.
[0136] The process of applying a smoothing filter to the prediction sample of the inter-frame block may be performed between the process S1102 and the process S1104. That is, the intra-frame predictor 122, the intra-frame predictor 442 may refer to the prediction sample of the inter-frame block to which the smoothing filter has been applied to derive the prediction sample of the intra-frame block. Whether the smoothing filter is applied may be explicitly indicated by a 1-bit flag signaled by the encoding device, or may be agreed upon in advance between the encoding device and the decoding device.
[0137] The intra predictor 122 and the intra predictor 442 may refer to some or all of the prediction samples of the inter block to derive the prediction samples of the intra block. For example, the intra predictor 122 and the intra predictor 442 may refer to the prediction samples of the inter block adjacent to the partition edge to derive the prediction samples of the intra block (S1106).
[0138] Fig.13 An example of deriving prediction samples of an intra block based on a planar mode with reference to prediction samples of an inter block adjacent to a partition edge is shown.
[0139] exist Fig.13 In the example of (a), we can refer to the adjacent samples (samples marked with patterns) 0,9 and r 9,0 And the prediction samples of inter-frame blocks (p 1,8 、p 2,7 、p 3,6 、p 4,5 、p 5,4 、p 6,3 、p 7,2 and p 8,1 ) 3,6 and p 7,2 To derive the prediction sample for the sample (p) in the intra block.
[0140] exist Fig.13 In the example of (b), we can refer to the adjacent samples (samples marked with patterns) 0,2 and r 2,0 And the prediction samples of inter-frame blocks (p 1,6 、p 2,5 、p3,4 、p 4,3 、p 5,2 and p 6,1 ) 2,5 and p 5,2 To derive the prediction sample for the sample (p) in the intra block.
[0141] exist Fig.13 In the example of (c), we can refer to the adjacent samples (samples marked with patterns) 0,6 and r 0,9 And the prediction samples of inter-frame blocks (p 1,2 、p 2,3 、p 3,4 、p 4,5 、p 5,6 、p 6,7 and p 7,8 ) 3,4 and p 5,6 To derive the prediction sample for the sample (p) in the intra block.
[0142] exist Fig.13 In the example of (d), we can refer to the adjacent samples (samples marked with patterns) 7,0 and r 9,0 And the prediction samples of inter-frame blocks (p 2,1 、p 3,2 、p 4,3 、p 5,4 、p 6,5 、p 7,6 and p 8,7 ) 4,3 and p 7,6 To derive the prediction sample for the sample (p) in the intra block.
[0143] 5. Prediction sample mixture
[0144] When the derivation of the prediction sample is completed, a hybrid process can be performed to remove or reduce the discontinuity occurring in the partition edge. The hybrid process can be a process of performing a weighted sum of an inter-frame prediction value and an intra-frame prediction value for a target sample to be predicted in the current block based on the distance between the target sample and the partition edge.
[0145] Whether to perform the hybrid process can be explicitly indicated by a 1-bit flag notified from the encoding device to the decoding device by signaling, or can be agreed upon in advance between the encoding device and the decoding device. Whether to perform the hybrid process can be determined based on the intra prediction mode of the intra block, the sample value difference between adjacent samples of the partition edge, etc. For example, when the sample value difference is greater than or less than a preset specific value, the hybrid process can be performed.
[0146] Derivation weights
[0147] If it is determined that hybrid processing is applied, the predictor 120, the predictor 440 may derive an inter-frame weight and an intra-frame weight (S1402). The inter-frame weight and the intra-frame weight may be derived by considering the distance between the target sample in the current block and the partition edge, the prediction type of the adjacent block, the intra-frame prediction mode, etc. The inter-frame weight corresponds to the weight of the inter-frame prediction value to be applied to the target sample. The intra-frame weight corresponds to the weight of the intra-frame prediction value to be applied to the target sample.
[0148] By considering the distance between the target sample and the partition edge, the inter-frame weight and the intra-frame weight can be derived. For example, Equation 2 can be used.
[0149] [Equation 2]
[0150] sampleWeight intra [x][y]=BlendFilter[dist]
[0151] sampleWeight inter [x][y]=Y-Blend Filter[dist]
[0152] In Equation 2, (x, y) indicates the location of the target sample, sampleWeight intra Indicates the intra-frame weight, and sampleWeight inter Indicates the inter-frame weight. Y is determined by the hybrid filter and is set equal to 2 if an n-bit hybrid filter is used. n . Accordingly, if a 3-bit blend filter is used, Y=8. The variable dist is a value obtained by scaling the vertical distance between the target sample and the partition edge, and the variable dist can be obtained by a lookup table based on the distance between the position of the target sample and the partition edge and the angle of the partition edge. When the dist value is 0 to 14, the value of the 3-bit BlendFilter[dist] can be derived by the following Table 1.
[0153] [Table 1]
[0154] dist 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 BlendFilter[dist] 4 4 5 5 5 5 6 6 6 6 7 7 7 7 8
[0155] The inter-frame weight and the intra-frame weight may be derived by further considering the prediction type of the neighboring blocks, for example, through Equation 2, Equation 3, and Equation 4.
[0156] [Equation 3]
[0157] sampleWeight intra [x][y]=BlendFilter[dist]-a
[0158] sampleWeight inter [x][y]=Y-BlendFilter[dist]+a
[0159] [Equation 4]
[0160] sampleWeight intra [x][y]=BlendFilter[dist]+a
[0161] sampleWeight inter [x][y]=Y-BlendFilter[dist]-a
[0162] When the number of neighboring blocks predicted by the intra prediction type (intra neighboring blocks) is less than the number of neighboring blocks predicted by the inter prediction type (inter neighboring blocks), Equation 3 may be applied. That is, when the number of inter neighboring blocks is greater than the number of intra neighboring blocks, the inter weight may be derived by adding the offset (a) to the value "Y-BlendFilter[dist]", and the intra weight may be derived by subtracting the offset (a) from the value BlendFilter[dist]. In this case, the offset (a) may be preset based on the difference between the number of intra neighboring blocks and the number of inter neighboring blocks.
[0163] When the number of intra-frame neighboring blocks is greater than the number of inter-frame neighboring blocks, Equation 4 may be applied. That is, when the number of intra-frame neighboring blocks is greater than the number of inter-frame neighboring blocks, the intra-frame weight may be derived by adding the offset (a) to the value BlendFilter[dist], and the inter-frame weight may be derived by subtracting the offset (a) from the value "Y-BlendFilter[dist]".
[0164] Fig.15 An example of deriving inter-frame weights and intra-frame weights by further considering the prediction types of two neighboring blocks is shown.
[0165] exist Fig.15 In the example, A and B indicate two adjacent blocks, (x, y) indicates the location of the target sample, and d i Indicates the distance between the target sample and the partition edge (represented as a dashed straight line).
[0166] If any one of A and B is predicted based on an intra prediction type and the other is predicted based on an inter prediction type, the inter weight and the intra weight may be derived by Equation 2. If both A and B are predicted based on an inter prediction type, the inter weight and the intra weight may be derived by Equation 3. If both A and B are predicted based on an intra prediction type, the inter weight and the intra weight may be derived by Equation 4.
[0167] Deriving weighted forecasts
[0168] When deriving the inter-frame weight and the intra-frame weight, the predictor 120, the predictor 440 may derive the prediction value for the weighted target sample by a hybrid process in which the intra-frame weight and the inter-frame weight are applied to the intra-frame prediction value and the inter-frame prediction value for the target sample, respectively.
[0169] The process of deriving the weighted prediction value can be performed by Equation 5.
[0170] [Equation 5]
[0171] P blended =(sampleWeight intra *P intra +sampleWeight inter *P inter +2 n-1 )>>n
[0172] In Equation 5, P blended Indicates the weighted prediction value. P intra Indicates the intra prediction value for the target sample. inter Indicates the inter prediction value for the target sample.
[0173] 6. Divide edge deblocking filter
[0174] If the GIIP mode has been implemented, a deblocking filter may be applied to block edges.
[0175] Whether to perform deblocking filtering can be explicitly indicated by a 1-bit flag notified from the encoding device to the decoding device by signaling, or it can be agreed in advance between the encoding device and the decoding device that deblocking filtering is always performed. Deblocking filtering can be performed when the difference between the values of samples of adjacent block edges is greater than or less than a preset specific value (threshold), or when the quantization parameter value of the values of samples of adjacent block edges is greater than or less than a preset parameter value. In addition, whether to perform deblocking filtering can be determined based on the intra-frame prediction mode. The deblocking filter used in deblocking filtering can be determined based on the intra-frame prediction mode.
[0176] When the prediction samples (prediction block) and the residual samples (residual block) are added to reconstruct the image, filtering is performed on the reconstructed image. The filtered image can be used for prediction of another image or stored in the memory 190, memory 470 for display.
[0177] like Fig.16 As shown, the conventional filtering unit may be configured to include a deblocking filter, a SAO, and an adaptive loop filter (ALF). The reconstructed image may be sequentially filtered in the order of the deblocking filter, the SAO, and the ALF to output a filtered image.
[0178] SAO classifies samples within a reconstructed image according to predefined criteria and adaptively applies an offset based on the corresponding classification. That is, SAO adaptively applies an offset to samples. For the application of SAO, mode information (information indicating any one type (mode) of edge offset, band offset, and non-performance of SAO) and offset information are signaled from the encoding device to the decoding device at the CTU level.
[0179] ALF applies adaptive filtering to sample features on a per-block basis in order to minimize the error between the original image and the reconstructed image. For the application of ALF, information about the filter set is signaled from the encoding device to the decoding device at the adaptation parameter set (APS) level, and the filter set is selected at the CTU level. A mode can be derived in block units within a CTU (the mode indicates any one of a total of N categories (modes) to which the corresponding block belongs, based on the direction and activity of the samples within the block). The filter corresponding to the derived mode is applied to the reconstructed image of the block.
[0180] As described above, SAO and ALF apply adaptive offset and filter to samples in sample units. In addition, SAO and ALF perform "processing of deriving sample features" and "processing of applying offset and filter suitable for the features".
[0181] When SAO and ALF are applied, the “process of deriving sample characteristics” is redundantly or individually performed for each of SAO and ALF. The present invention proposes a method of integrating the individually performed processes into one process to efficiently perform a filtering process for a reconstructed image.
[0182] Fig.17 It is an exemplary block diagram of the filtering unit 180 and the filtering unit 460 proposed in the present invention. Fig.18 is a flowchart for describing an example of an adaptive filtering method.
[0183] like Fig.17As shown, the filtering unit 180 and the filtering unit 460 may be configured to include a deblocking filter 182, a deblocking filter 462 and an adaptive filtering unit 1700. The adaptive filtering unit 1700 may be configured to include an extracting unit 1710, a mapping unit 1720, a determining unit 1730, a filtering unit 1740 and a cropping unit 1750.
[0184] A. Feature Extraction
[0185] The extraction unit 1710 may extract one or more features from one or more reconstructed samples within the reconstructed image to which the deblocking filter has been applied ( S1802 ).
[0186] Features may be extracted for each unit in which filtering is performed in the reconstructed sample within the reconstructed image. In this case, the unit in which filtering is performed may be any one of a sample, a block, a CTU, etc. The feature may be a representative value of the reconstructed sample, such as an average value, a most likely value, a central value, a minimum value, a maximum value, etc. of the reconstructed sample. In addition, the feature may be a value calculated using a representative value of the reconstructed sample.
[0187] Features can be extracted by a one-dimensional filter or a two-dimensional filter with three taps. The one-dimensional filter can be applied in several directions, such as the vertical direction, the horizontal direction, and the diagonal direction. For each direction, the value obtained from applying the one-dimensional filter can be extracted as a feature of that direction. For example, the absolute value of the filtered value can be extracted as a feature. The two-dimensional filter can be a first-order differential filter or a second-order differential filter. The two-dimensional filter can be a first-order differential filter or a second-order differential filter defined for several directions, such as the horizontal direction ( Fig.19a (b) Fig.19b (b)), vertical direction ( Fig.19a (a) Fig.19b (a)) and diagonal direction ( Fig.19a (c) and (d), Fig.19b (c) and (d)).
[0188] The features may be extracted using information required for feature extraction. The information used for feature extraction may be filter-related information. The filter-related information may include the shape, size, coefficients, etc. of the filter, may include filter mode information, or may include coefficients of a residual filter. In this case, the extraction unit 1710 of the decoding device may determine the filter by using the information required for feature extraction, and may extract features from the reconstructed samples by using the determined filter.
[0189] The information required for extracting the feature may be pattern information indicating any one of several patterns that can be used for reconstructed samples according to directions. In this case, the extraction unit 1710 of the decoding device may determine the pattern (direction) of the reconstructed samples by using the information required for extracting the feature, and may extract features (currently reconstructed samples and adjacent reconstructed samples) from the reconstructed samples by using the determined pattern.
[0190] Fig.20a and Fig.20b An example of a currently reconstructed sample and adjacent reconstructed samples that can be extracted as features is shown. Features can be such as Fig.20a Some of the currently reconstructed samples (c) and adjacently reconstructed samples (a, b) arranged horizontally in (a), or may be as follows Fig.20a The features may be as follows: Fig.20b Some of the currently reconstructed samples (c) and the adjacent reconstructed samples (a, b) arranged in the 135 degree diagonal direction as shown in (a), or may be as follows Fig.20b (b) shows some of the current reconstructed samples (c) and adjacent reconstructed samples (a, b) arranged in a 45-degree diagonal direction.
[0191] The information required for feature extraction may be encoded by an encoding device and may be notified to a decoding device by one or more signals at an SPS level, a PPS level, a slice header (SH) level, a picture header (PH) level, and a CTU level.
[0192] The extracted features may be input to the mapping unit 1720. According to an embodiment, quantized values for quantizing the extracted features may be input to the mapping unit 1720, or mapped values of the extracted features mapped using a lookup table may be input to the mapping unit 1720. Information about the lookup table may be signaled from the encoding device to the decoding device, or may be agreed upon in advance between the encoding device and the decoding device so as to use the same lookup table.
[0193] B. Category Mapping
[0194] The mapping unit 1720 may determine the category to which the reconstructed sample belongs based on the extracted feature value (S1804). This process may be a process of mapping the extracted feature to the category to which the reconstructed sample belongs in a preset category.
[0195] For example, the class mapping may be performed by using an equation. The equation for class mapping may be signaled from the encoding device to the decoding device, or the parameters of the equation may be signaled from the encoding device to the decoding device.
[0196] As an example, Equation 6 may be used as an equation for category mapping.
[0197] [Equation 6]
[0198] C = A × M + B In Equation 6, C indicates the category of mapping, and A and B indicate two features input to the mapping unit 1720. Either A and B may be a change amplitude centered on the currently reconstructed sample, or may be a direction of change. M may be determined based on A, B, and N (the number of preset categories). For example, if N = 25, C = 0 to 24, A = 0 to 4, and B = 0 to 4, then M = 5 may be determined.
[0199] As another example, the class mapping may be performed using a lookup table. The encoding device may encode the lookup table for class mapping and signal the lookup table to the decoding device, or the encoding device and the decoding device may be preconfigured to use the same lookup table.
[0200] As yet another example, if a currently reconstructed sample and adjacently reconstructed samples are extracted as features and input to the mapping unit 1720 , categories may be mapped based on sample values of the currently reconstructed sample and sample values of the adjacently reconstructed samples.
[0201] refer to Fig.21 , when the value of the adjacent reconstructed sample (a, c) is greater than the value of the current reconstructed sample (b) and the value of the adjacent reconstructed sample (a, c) is equal to Fig.21 When the values of any one of the adjacent reconstructed samples (a, c) are the same as the value of the currently reconstructed sample (b), and the value of the other adjacent reconstructed sample (a, c) is greater than the value of the currently reconstructed sample (b), as shown in (a), the first category can be mapped. Fig.21 (b) and Fig.21 As shown in (c), the second category can be mapped. When the value of any one of the adjacent reconstructed samples (a, c) is the same as the value of the current reconstructed sample (b), and the value of the other adjacent reconstructed sample (a, c) is less than the value of the current reconstructed sample (b), as shown in (c). Fig.21 (d) and Fig.21 As shown in (e), the third category can be mapped. When the value of the currently reconstructed sample (b) is greater than the value of the adjacent reconstructed sample (a, c), and the value of the adjacent reconstructed sample (a, c) is Fig.21When the same as shown in (f) of , the fourth category can be mapped. Fig.21 (a) to Fig.21 The case where (f) does not correspond may be determined as the fifth category. The mapped category may be input to the determination unit 1730.
[0202] C. Determine the filter
[0203] The determination unit 1730 may determine an SAO filter and an ALF filter corresponding to the mapped category.
[0204] The encoding device may generate or determine the optimal filter coefficient (optimal ALF filter) and the offset (optimal SAO filter) for each unit performing filtering based on the type of mapping. For example, the optimal ALF filter and the optimal SAO filter may be determined by applying the least squares method to the values of the original samples and the values of the reconstructed samples to determine the optimal ALF filter, and then determining the offset using the average value of the error between the values of the reconstructed samples and the values of the original samples to which the optimal ALF has been applied.
[0205] Information about the optimal ALF filter (filter coefficient information) and information about the optimal SAO filter (offset information) may be encoded and signaled from the encoding device to the decoding device. The filter coefficient information and the offset information may be signaled at the APS level or a level higher than the APS level (e.g., a slice header level).
[0206] The decoding device can determine the SAO filter and the ALF filter corresponding to the category by using the filter coefficient information and the offset information notified by the signal. For example, the decoding device can determine or update the filter coefficient and offset used in the unit to be filtered currently by using "filter coefficient information and offset information notified by the signal at a higher level", "filter coefficient information and offset information notified by the signal at the APS level", or "filter coefficient and offset used in the previous filtering unit". The determined filter coefficients may be some of all the coefficients of the filter. In addition, some filter coefficients may be shared in subsequent units to be filtered. Some of the shared filter coefficients may be determined by the category or mode information notified by the signal from the encoding device.
[0207] Information about the determined SAO filter and ALF filter may be input to the filtering unit 1740 .
[0208] D. Filtering
[0209] The filtering unit 1740 may filter the reconstructed samples by using the SAO filter and the ALF filter determined by the determination unit 1730 ( S1808 ).
[0210] The process S1808 of filtering the reconstructed samples by the filtering unit 1740 can be divided into a process of applying an SAO filter to the reconstructed samples (SAO filtering) and a process of applying an ALF filter to the reconstructed samples (ALF filtering). The SAO filtering and the ALF filtering can be performed simultaneously. Alternatively, the ALF filtering can be performed after the SAO filtering is performed first.
[0211] If SAO filtering and ALF filtering are performed simultaneously, a process of filtering reconstructed samples can be expressed as Equation 7.
[0212] [Equation 7]
[0213] x=W·X+O
[0214] In Equation 7, W indicates a coefficient vector of an ALF filter, O indicates a SAO filter coefficient (offset), X indicates a reconstructed sample vector corresponding to a filter coefficient position, and x indicates a reconstructed sample obtained from SAO filtering and ALF filtering. W and X are vectors having the same length. As shown in Equation 7, x can be derived by performing a dot product calculation on W and X and then applying O.
[0215] If SAO filtering is performed first and then ALF filtering is performed, the process of filtering the reconstructed samples can be expressed as Equation 8.
[0216] [Equation 8]
[0217] x=W·(X+O)
[0218] As shown in Equation 8, after O is first applied to X, a dot product calculation between (X+O) and W can be performed so that x can be derived.
[0219] E. Cutting
[0220] The clipping unit 1750 may clip (limit) the value of the reconstructed sample based on the value of the filtered reconstructed sample (the value of the current reconstructed sample and the value of the adjacent reconstructed sample) and a threshold value.
[0221] The threshold value may be determined based on the bit depth of the image, etc. The encoding device may encode information about the determined threshold value, and may signal the information to the decoding device. The decoding device may determine the threshold value to be applied to the cropping by using the decoded information about the threshold value.
[0222] Although exemplary embodiments of the present invention have been described for illustrative purposes, it will be appreciated by those skilled in the art that various modifications and changes are possible without departing from the spirit and scope of the present invention. For the sake of brevity and clarity, exemplary embodiments have been described. Accordingly, it will be appreciated by those of ordinary skill that the scope of the embodiments is not limited by the embodiments explicitly described above, but is included in the claims and their equivalents.
Claims
1. A method for decoding a current block based on a first mode, the method include: Divide the current block into non-rectangular blocks; Determining, in the non-rectangular block, an intra-frame block to be subjected to intra-frame prediction and an inter-frame block to be subjected to inter-frame prediction; deriving inter prediction samples of a first region including an inter block based on the motion information; deriving intra prediction samples of a second region including the intra block based on the intra prediction mode; Generate a prediction block of a current block based on the inter-frame prediction samples and the intra-frame prediction samples; reconstructing transform coefficients of the current block from the bitstream; Generate a residual block of the current block by inversely transforming the transform coefficients; Reconstruct the current block based on the prediction block and the residual block, The intra prediction samples of the second region are derived by: constructing an intra prediction mode candidate list including one or more directional modes corresponding to angles formed by partition edges of the non-rectangular block, An intra prediction mode is selected from the intra prediction mode candidate list.
2. The method according to claim 1, in, An intra block and an inter block are determined in the non-rectangular block based on an angle formed by a partition edge of the non-rectangular block and a horizontal direction of the current block.
3. The method according to claim 1, in, An intra block and an inter block are determined in the non-rectangular block based on a vertical distance from a center position of the current block to a partition edge of the non-rectangular block.
4. The method according to claim 1, in, The intra prediction samples of the intra block are derived by using the inter prediction samples of the inter block as reference samples.
5. The method according to claim 4, in, The intra prediction samples of the intra block are derived by using the prediction samples of the partition edge of the adjacent non-rectangular block among the inter prediction samples of the inter block as reference samples.
6. The method according to claim 1, further comprising: include: Whether to apply the first mode is determined based on the number of intra-predicted neighboring blocks among neighboring blocks of the current block, wherein when it is determined to apply the first mode, splitting the current block into non-rectangular blocks is performed.
7. The method according to claim 1, in, Generating the prediction block of the current block includes: deriving intra-weights and inter-weights to be applied to a target sample to be predicted in the current block by considering the distance between the target sample and a partition edge of a non-rectangular block in the current block; and A weighted prediction value of the target sample is derived based on a weighted intra prediction value obtained by applying an intra weight to an intra prediction sample corresponding to the target sample and a weighted inter prediction value obtained by applying an inter weight to an inter prediction sample corresponding to the target sample.
8. The method according to claim 7, in, The intra-weight and inter-weight are derived by further considering whether the neighboring blocks of the current block have been inter-predicted or intra-predicted.
9. A method for encoding a current block based on a first mode, the method include: Divide the current block into non-rectangular blocks; Determining, in the non-rectangular block, an intra-frame block to be subjected to intra-frame prediction and an inter-frame block to be subjected to inter-frame prediction; deriving inter prediction samples of a first region including an inter block based on the motion information; deriving intra prediction samples of a second region including the intra block based on the intra prediction mode; Generate a prediction block of a current block based on the inter-frame prediction samples and the intra-frame prediction samples; Generate a residual block of the current block based on the prediction block; Generate transform coefficients by transforming the residual block; Encode the information of the transform coefficients, The intra prediction samples of the second region are derived by: constructing an intra prediction mode candidate list including one or more directional modes corresponding to angles formed by partition edges of the non-rectangular block, An intra prediction mode is selected from the intra prediction mode candidate list.
10. A method for transmitting a bit stream comprising encoded video data, the method include: Generate a bitstream by encoding the current block; Sending a bitstream to a video decoding device, The current block encoding includes: Divide the current block into non-rectangular blocks; Determining, in the non-rectangular block, an intra-frame block to be subjected to intra-frame prediction and an inter-frame block to be subjected to inter-frame prediction; deriving inter prediction samples of a first region including an inter block based on the motion information; deriving intra prediction samples of a second region including the intra block based on the intra prediction mode; Generate a prediction block of a current block based on the inter-frame prediction samples and the intra-frame prediction samples; Generate a residual block of the current block based on the prediction block; Generate transform coefficients by transforming the residual block; Encode the information of the transform coefficients, The intra prediction samples of the second region are derived by: constructing an intra prediction mode candidate list including one or more directional modes corresponding to angles formed by partition edges of the non-rectangular block, An intra prediction mode is selected from the intra prediction mode candidate list.
Citation Information
Patent Citations
OLED display panel and manufacturing method thereof
KR1020200003143A