Method and apparatus for video coding for adaptive determination of mixing region in geometric partitioning mode
By adopting geometric partition mode and adaptive hybrid area determination method in video encoding and decoding technology, the problem of insufficient encoding and decoding efficiency and video quality in the prior art at high resolution and frame rates is solved, and more efficient video encoding and decoding and improved video quality are achieved.
Patent Information
- Application Number
- CN202380080033.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2023-11-09
- Publication Date
- 2025-06-27
AI Technical Summary
The existing video encoding and decoding technology has insufficient encoding and decoding efficiency and image enhancement effect at high resolution and frame rates, making it difficult to meet the needs of more efficient compression and improving video quality.
Geometric partitioning mode (GPM) is used to generate prediction blocks of sub-blocks partitioned from the current block, and adaptively determine the mixed region for weighted sum of prediction blocks of sub-blocks, and optimize the generation of prediction signals through the mixing matrix.
Improves video encoding and decoding efficiency and enhances video quality, especially under high resolution and high frame rate conditions.
Smart Images

Figure CN120226348A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a video encoding and decoding method and apparatus for adaptively determining a hybrid region in a geometric partitioning mode. Background Art
[0002] The statements in this section merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Since video data has a large amount of data compared to audio data or still image data, video data requires a large amount of hardware resources (including memory) to store or transmit uncompressed video data.
[0004] Accordingly, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression technologies include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), and the Versatile Video Coding (VVC) has a coding and decoding efficiency that is approximately 30% or more higher than that of HEVC.
[0005] However, due to the gradual increase in image size, resolution, and frame rate, the amount of data to be encoded also increases. Accordingly, there is a need to provide a new compression technology with a higher coding and decoding efficiency and improved image enhancement effect than existing compression technologies.
[0006] The VVC technology adopts a geometric partitioning mode (GPM) of the inter-frame prediction technology for prediction based on flexible partitioning rather than square and rectangular partitioning of a Quadtree plus Multiple-type Tree (QT+MTT) partitioning structure. The GPM performs prediction of a current block based on a mode index of a region used for partitioning and motion vector information. Here, the mode index indicates one of predefined geometric partitioning modes for partitioning the current block, and motion vector information can be derived for predicting a sub-region partitioned from the current block.
[0007] The encoder sends the mode index and motion vector information to the decoder. The decoder partitions the current block into two regions according to the parsed geometric partitioning mode. The decoder uses the motion vector information to generate prediction signals for each sub-region, and then performs weighted summation on the generated prediction signals to generate the final prediction block. At this time, the weights for weighted summation can be determined based on the geometric partitioning mode, and the mixed region can be determined as a fixed region based on the size of the current block and the geometric partitioning mode. Due to the foregoing, in order to improve video coding and decoding efficiency and enhance video quality, it is necessary to adaptively determine the mixed region under the geometric partitioning mode. Summary of the Invention
[0008] Technical Problem
[0009] The present invention is committed to providing a video coding and decoding method and apparatus, which generate prediction blocks of sub-blocks partitioned from the current block according to the geometric partitioning mode (GPM), and then adaptively determine the mixed region for weighted summation of the prediction blocks of the sub-blocks.
[0010] Technical Solution
[0011] At least one aspect of the present invention provides a method for reconstructing a current block by a video decoding apparatus. The method includes decoding the geometric partitioning mode of the current block. The method further includes partitioning the current block into a first sub-region and a second sub-region centered on the geometric segmentation boundary according to the geometric partitioning mode. The method further includes generating a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction modes of the first sub-region and the second sub-region. The method further includes determining a final mixed region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and the initial mixed region. The method further includes determining a mixing matrix in the final mixed region. The method further includes generating a final prediction signal of the current block by performing weighted summation on the first prediction signal and the second prediction signal in the final mixed region by using the mixing matrix.
[0012] Another aspect of the present invention provides a method for encoding a current block by a video encoding device. The method includes determining a geometric partitioning pattern of the current block. The method further includes partitioning the current block into a first sub-region and a second sub-region centered on a geometric splitting boundary according to the geometric partitioning pattern. The method further includes generating a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction patterns of the first sub-region and the second sub-region. The method further includes determining a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and an initial mixing region. The method further includes determining a mixing matrix in the final mixing region. The method further includes generating a first final prediction signal of the current block by weighted summing the first prediction signal and the second prediction signal in the final mixing region by using the mixing matrix.
[0013] Another aspect of the present invention provides a computer-readable recording medium storing a bitstream generated by a video encoding method. The video encoding method includes determining a geometric partitioning pattern of the current block. The video encoding method further includes partitioning the current block into a first sub-region and a second sub-region centered on a geometric splitting boundary according to the geometric partitioning pattern. The video encoding method further includes generating a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction patterns of the first sub-region and the second sub-region. The video encoding method further includes determining a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and an initial mixing region. The video encoding method further includes determining a mixing matrix in the final mixing region. The video encoding method further includes generating a final prediction signal of the current block by weighted summing the first prediction signal and the second prediction signal in the final mixing region by using the mixing matrix.
[0014] Beneficial effects
[0015] As described above, the present invention provides a video encoding and decoding method and apparatus, which generate prediction blocks of sub-blocks partitioned from a current block according to a geometric partitioning pattern (GPM), and then adaptively determine a mixing region for weighted summation of the prediction blocks of the sub-blocks. Therefore, the video encoding and decoding method and apparatus improve the video encoding and decoding efficiency and enhance the video quality. Description of the drawings
[0016] Figure 1 is a block diagram of a video encoding device that can implement the technology of the present invention.
[0017] Figure 2 Illustrates a method for partitioning a block by using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.
[0018] Figure 3a andFigure 3b Shows multiple intra prediction modes including wide-angle intra prediction modes.
[0019] Figure 4 Shows adjacent blocks of the current block.
[0020] Figure 5 Is a block diagram of a video decoding device that can implement the technology of the present invention.
[0021] Figure 6 Is a block diagram of a detailed part of a video decoding device according to at least one embodiment of the present invention.
[0022] Figure 7 Is a flowchart of a geometric partitioning mode according to at least one embodiment of the present invention.
[0023] Figure 8 Is a flowchart of a geometric partitioning mode according to another embodiment of the present invention.
[0024] Figure 9 Is a schematic diagram showing an initial mixing region according to at least one embodiment of the present invention.
[0025] Figure 10 Is a schematic diagram showing a mixing matrix according to at least one embodiment of the present invention.
[0026] Figure 11a and Figure 11b Are flowcharts for determining an implicit final mixing region according to some embodiments of the present invention.
[0027] Figures 12a to 12c Is a flowchart showing the determination of the intensity at a geometric segmentation boundary according to some embodiments of the present invention.
[0028] Figure 13a and Figure 13b Are flowcharts showing the determination of the intensity at a geometric segmentation boundary according to other embodiments of the present invention.
[0029] Figures 14a to 14d Is a schematic diagram showing the generation of a final mixing region according to some embodiments of the present invention.
[0030] Figure 15 Is a schematic diagram showing the composition of weights in a final mixing region according to at least one embodiment of the present invention.
[0031] Figure 16 Is a schematic diagram showing the determination of a final mixing region based on template matching according to at least one embodiment of the present invention.
[0032] Figure 17It is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention. Detailed Description
[0033] Hereinafter, some embodiments of the present invention will be described in detail with reference to the accompanying illustrative drawings. In the following description, the same reference numerals denote the same elements, although the elements are shown in different drawings. In addition, in the following description of some embodiments, when the detailed description of related known components and functions is considered to obscure the subject matter of the present invention, the detailed description of the related known components and functions may be omitted for clarity and conciseness.
[0034] Figure 1 It is a block diagram of a video encoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 1 the illustration of, the video encoding device and the components of the device will be described.
[0035] The encoding device may include: an image splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.
[0036] Each component of the encoding device may be implemented as hardware or software, or implemented as a combination of hardware and software. Additionally, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.
[0037] A video consists of one or more sequences including a plurality of images. Each image is segmented into a plurality of regions, and encoding is performed on each region. For example, an image is segmented into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile and / or slice is segmented into one or more coding tree units (CTUs). Additionally, each CTU is segmented into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information applied to the CUs included in one CTU is encoded as the syntax of the CTU. Additionally, information applied to all blocks in one slice is encoded as the syntax of the slice header, while information applied to all blocks constituting one or more images is encoded as the syntax of the Picture Parameter Set (PPS) or the picture header. Furthermore, information commonly referred to by a plurality of images is encoded as the Sequence Parameter Set (SPS). Additionally, information commonly referred to by one or more SPSs is encoded as the Video Parameter Set (VPS). Moreover, information applied to one tile or tile group can also be encoded as the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntax.
[0038] The image splitter 110 determines the size of the coding tree unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and is transmitted to the video decoding device.
[0039] The image splitter 110 segments each image constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively segments the CTUs by using a tree structure. A leaf node in the tree structure becomes a coding unit (CU), and the CU is the basic unit of encoding.
[0040] The tree structure can be a quadtree (QT), where a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), where a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), where a higher node is divided into three lower nodes in a 1:2:1 ratio. The tree structure can also be a structure that mixes two or more of the QT structure, BT structure, and TT structure. For example, a quadtree plus binarytree (QTBT) structure can be used, or a quadtree plus binarytreeternarytree (QTBTTT) structure can be used. Here, the binarytreeternarytree (BTTT) is added to the tree structure to form a multiple-type tree (MTT).
[0041] Figure 2 is a schematic diagram for describing a method of dividing a block by using the QTBTTT structure.
[0042] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. The entropy encoder 155 encodes a first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes and signals it to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root nodes allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or TT structure. For example, there can be two directions, namely, the direction of horizontally dividing the block of the corresponding node and the direction of vertically dividing the block of the corresponding node. As Figure 2 shown, when the MTT division starts, the entropy encoder 155 encodes a second flag (mtt_split_flag) indicating whether the node is divided, and a flag indicating the division direction (vertical or horizontal) and / or a flag indicating the division type (binary or ternary) in the case where the node is divided, and signals it to the video decoding device.
[0043] Alternatively, before encoding a first flag (QT_split_flag) indicating whether each node is split into four lower-layer nodes, a CU split flag (split_cu_flag) indicating whether a node is split may also be encoded. When the value of the CU split flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU, which is the basic unit of encoding. When the value of the CU split flag (split_cu_flag) indicates that each node is split, the video encoding device starts encoding the first flag first according to the above scheme.
[0044] When QTBT is used as another example of a tree structure, there may be two types, that is, a type that horizontally splits the block of the corresponding node into two blocks of the same size (i.e., symmetric horizontal split) and a type that vertically splits the block of the corresponding node into two blocks of the same size (i.e., symmetric vertical split). The entropy encoder 155 encodes a split flag (split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and split type information indicating the split type, and transmits them to the video decoding device. On the other hand, there may additionally be a type in which the block of the corresponding node is split into two asymmetric blocks. The asymmetric form may include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or may also include a form in which the block of the corresponding node is split in the diagonal direction.
[0045] A CU may have various sizes according to the QTBT or QTBTTT split from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of QTBTTT) is referred to as the "current block". When QTBTTT split is adopted, in addition to the square shape, the shape of the current block may also be a rectangular shape.
[0046] The predictor 120 predicts the current block to generate a prediction block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.
[0047] Generally, each of the current blocks in the image can be predictively encoded. Generally, the prediction of the current block can be performed by using an intra prediction technique (which uses data from the image including the current block) or an inter prediction technique (which uses data from the image encoded before the image including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.
[0048] The intra predictor 122 predicts the pixels in the current block by using the pixels (reference pixels) adjacent to the current block in the current image including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3aAs shown, multiple intra-prediction modes may include two non-directional modes including Planar mode and DC mode, and may include 65 directional modes. Adjacent pixels to be used and algorithm equations are defined differently according to each prediction mode.
[0049] For efficient directional prediction of a current block having a rectangular shape, the directional modes shown by the dashed arrows in Figure 3b (#67 to #80, intra-prediction modes #-1 to #-14) may be additionally used. The directional modes may be referred to as "wide angle intra-prediction modes". In Figure 3b , the arrows indicate the corresponding reference samples for prediction, rather than representing the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide angle intra-prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide angle intra-prediction mode, some wide angle intra-prediction modes available for the current block may be determined by the ratio of the width to the height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide angle intra-prediction modes having an angle less than 45 degrees (intra-prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, wide angle intra-prediction modes having an angle greater than -135 degrees are available.
[0050] The intra-predictor 122 may determine the intra-prediction to be used for encoding the current block. In some examples, the intra-predictor 122 may encode the current block by utilizing multiple intra-prediction modes, and may also select an appropriate intra-prediction mode to be used from test modes. For example, the intra-predictor 122 may calculate rate-distortion values by utilizing rate-distortion analysis of multiple tested intra-prediction modes, and may also select the intra-prediction mode having the best rate-distortion characteristics in the test modes.
[0051] The intra-predictor 122 selects one intra-prediction mode from multiple intra-prediction modes, and predicts the current block by utilizing adjacent pixels (reference pixels) and algorithm equations determined according to the selected intra-prediction mode. Information about the selected intra-prediction mode is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0052] The inter - frame predictor 124 generates a predicted block of the current block by using motion - compensation processing. The inter - frame predictor 124 searches for the block most similar to the current block in a reference image that has been encoded and decoded earlier than the current image, and generates a predicted block of the current block by using the searched - for block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current image and the predicted block in the reference image. Generally, motion estimation is performed on the luma component, and the motion vector calculated based on the luma component is used for both the luma component and the chroma component. The entropy encoder 155 encodes the motion information including the information of the reference image and the information about the motion vector used for predicting the current block, and transmits it to the video decoding device.
[0053] The inter - frame predictor 124 may also perform interpolation of the reference image or reference block to increase the prediction accuracy. In other words, sub - samples are interpolated between two consecutive integer samples by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for the block most similar to the current block on the interpolated reference image, the motion vector can represent fractional - unit precision rather than integer - sample - unit precision. For each target region to be encoded, such as units like slices, tiles, CTUs, CUs, etc., the precision or resolution of the motion vector can be set differently. When applying such an adaptive motion vector resolution (AMVR), information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution can be information representing the precision of the motion - vector difference described below.
[0054] On the other hand, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference images and two motion vectors representing the positions of the blocks most similar to the current block in each reference image are used. The inter-frame predictor 124 selects a first reference image and a second reference image from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the blocks most similar to the current block in the corresponding reference images to generate a first reference block and a second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Further, motion information including information about the two reference images used for predicting the current block and information about the two motion vectors is transmitted to the entropy encoder 155. Here, reference picture list 0 may be composed of images in the pre-reconstructed images that are before the current image in the display order, and reference picture list 1 may be composed of images in the pre-reconstructed images that are after the current image in the display order. However, although not particularly limited thereto, pre-reconstructed images after the current image in the display order may be additionally included in reference picture list 0. Conversely, pre-reconstructed images before the current image may also be additionally included in reference picture list 1.
[0055] To minimize the amount of bits consumed for encoding the motion information, various methods can be used.
[0056] For example, when the reference image and motion vector of the current block are the same as those of an adjacent block, information of the adjacent block that can be recognized is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.
[0057] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the adjacent blocks of the current block.
[0058] As the adjacent blocks for deriving the merge candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image can be used, as Figure 4 shown. In addition, in addition to the current image where the current block is located, blocks in the reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as merge candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be additionally used as a merge candidate. If the number of merge candidates selected by the above method is less than the preset number, zero vectors are added to the merge candidates.
[0059] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using adjacent blocks. A merge candidate to be used as the motion information of the current block is selected from among the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and transmitted to the video decoding device.
[0060] The merge skip mode is a special case of the merge mode. After quantization, when all the transform coefficients for entropy coding are close to zero, only the adjacent block selection information is transmitted without transmitting the residual signal. By using the merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.
[0061] Thereafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.
[0062] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.
[0063] In the AMVP mode, the inter-frame predictor 124 derives motion vector prediction candidates for the motion vector of the current block by using adjacent blocks of the current block. As the adjacent blocks for deriving the motion vector prediction candidates, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current image shown Figure 4 can be used. In addition, in addition to the current image where the current block is located, blocks in a reference image (which may be the same as or different from the reference image used for predicting the current block) can also be used as adjacent blocks for deriving the motion vector prediction candidates. For example, the co-located block of the current block in the reference image or a block adjacent to the co-located block can be used. If the number of motion vector candidates selected by the above method is less than the preset number, a zero vector is added to the motion vector candidates.
[0064] The inter-frame predictor 124 derives motion vector prediction candidates by using the motion vectors of adjacent blocks, and determines the motion vector prediction of the motion vector of the current block by using the motion vector prediction candidates. In addition, the motion vector difference is calculated by subtracting the motion vector prediction from the motion vector of the current block.
[0065] Motion vector prediction can be obtained by applying a predefined function (e.g., median and mean calculations, etc.) to motion vector prediction candidates. In this case, the video decoding device also knows the predefined function. In addition, since the neighboring blocks used to derive the motion vector prediction candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction candidates. Accordingly, in this case, the information about the motion vector difference and the information about the reference image used to predict the current block are encoded.
[0066] On the other hand, motion vector prediction can also be determined by a scheme of selecting any one of the motion vector prediction candidates. In this case, the information for identifying the selected motion vector prediction candidate is additionally encoded together with the information about the motion vector difference and the information about the reference image used to predict the current block.
[0067] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra predictor 122 or the inter predictor 124 from the current block.
[0068] The transformer 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can transform the residual signal in the residual block by using the entire size of the residual block as a transform unit, or the residual block can also be divided into multiple sub-blocks, and the transform can be performed by using the sub-blocks as transform units. Alternatively, the residual block is divided into two sub-blocks, namely a transform region and a non-transform region, to transform the residual signal by using only the transform region sub-block as a transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 encodes a flag (cu_sbt_flag) indicating only the transformed sub-block, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag), and signals it to the video decoding device. In addition, the size of the transform region sub-block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, the entropy encoder 155 additionally encodes a flag (cu_sbt_quad_flag) for dividing the corresponding segmentation, and signals it to the video decoding device.
[0069] On the other hand, the transformer 140 may perform the transformation of the residual block separately in the horizontal direction and the vertical direction. For this transformation, various types of transformation functions or transformation matrices may be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation may be defined as a multiple transform set (MTS). The transformer 140 may select a pair of transformation functions in the MTS that has the highest transformation efficiency, and may transform the residual block on each of the horizontal direction and the vertical direction. Information (mts_idx) about the pair of transformation functions in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.
[0070] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 may also quantize the relevant residual block immediately without transforming any block or frame. The quantizer 145 may also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. A quantization matrix applied to the quantized transform coefficients arranged in two dimensions may be encoded and signaled to the video decoding device.
[0071] The rearrangement unit 150 may perform rearrangement of the coefficient values on the quantized residual values.
[0072] The rearrangement unit 150 may change the 2D coefficient array to a 1D coefficient sequence by coefficient scanning. For example, the rearrangement unit 150 may scan the coefficients from the DC coefficient to the high-frequency region using a zig-zag scan or a diagonal scan to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, a vertical scan that scans the 2D coefficient array in the column direction and a horizontal scan that scans the 2D block type coefficients in the row direction may also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scan method to be used may be determined among the zig-zag scan, the diagonal scan, the vertical scan, and the horizontal scan.
[0073] The entropy encoder 155 encodes the sequence of the 1D quantized transform coefficients output from the rearrangement unit 150 using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc. to generate a bitstream.
[0074] In addition, the entropy encoder 155 encodes information related to block partitioning (e.g., CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, and MTT partitioning direction, etc.) so that the video decoding device can partition blocks in the same way as the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (merge index in the case of the merge mode, and information regarding the reference image index and motion vector difference in the case of the AMVP mode) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding the quantization parameter and information regarding the quantization matrix).
[0075] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.
[0076] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next block, the pixels in the reconstructed current block are used as reference pixels.
[0077] The loop filter unit 180 performs filtering on the reconstructed pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180, as an in-loop filter, may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.
[0078] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove blocking artifacts that occur due to block-based coding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked video. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels that occur due to lossy coding. The SAO filter 184 applies an offset in units of CTUs to enhance the subjective image quality and the coding efficiency. On the other hand, the ALF 186 performs block-based filtering and applies different filters by dividing the degree of the boundary and the variation amount of the corresponding block to compensate for distortion. Information about the filter coefficients to be used for the ALF can be encoded and signaled to the video decoding device.
[0079] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter prediction of the blocks within the image to be encoded subsequently.
[0080] The video encoding device can store the bitstream of the encoded video data in a non-volatile storage medium or send the bitstream to the video decoding device through a communication network.
[0081] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present invention. Hereinafter, with reference to Figure 5 ,the video decoding device and the components of the device are described.
[0082] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transformer 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.
[0083] Similar to Figure 1 the video encoding device, each component of the video decoding device can be implemented as hardware or software, or implemented as a combination of hardware and software. In addition, the functions of each component can be implemented as software, and the microprocessor can also be implemented to execute the functions of the software corresponding to each component.
[0084] The entropy decoder 510 extracts information related to block partitioning by decoding the bitstream generated by the video encoding device to determine the current block to be decoded, and extracts the prediction information and the information about the residual signal required for reconstructing the current block.
[0085] The entropy decoder 510 determines the size of a coding tree unit (CTU) by extracting information about the CTU size from a sequence parameter set (SPS) or a picture parameter set (PPS), and divides an image into CTUs with the determined size. In addition, a CTU is determined as the top layer (i.e., the root node) of a tree structure, and the splitting information of the CTU can be extracted to split the CTU by using the tree structure.
[0086] For example, when splitting a CTU by using a QTBTTT structure, a first flag (QT_split_flag) related to the splitting of a quad tree (QT) is first extracted to divide each node into four lower-layer nodes. In addition, a second flag (mtt_split_flag) related to the splitting of a multi-type tree (MTT), a splitting direction (vertical / horizontal), and / or a splitting type (binary / trinary) are extracted for a node corresponding to a leaf node of the QT to divide the corresponding leaf node into an MTT structure. As a result, each node below the leaf node of the QT is recursively divided into a binary tree (BT) or a ternary tree (TT) structure.
[0087] As another example, when splitting a CTU by using a QTBTTT structure, a CU splitting flag (split_cu_flag) indicating whether to split a coding unit (CU) is extracted. When splitting the corresponding block, the first flag (QT_split_flag) may also be extracted. During the splitting process, for each node, zero or more recursive MTT splittings may occur after zero or more recursive QT splittings. For example, for a CTU, the MTT splitting may occur immediately, or conversely, only multiple QT splittings may occur.
[0088] As another example, when splitting a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the splitting of a QT is extracted to divide each node into four lower-layer nodes. In addition, a splitting flag (split_flag) indicating whether to further split the node corresponding to the leaf node of the QT into a BT and splitting direction information are extracted.
[0089] On the other hand, when the entropy decoder 510 determines a current block to be decoded by using the splitting of a tree structure, the entropy decoder 510 extracts information about a prediction type indicating whether the current block is intra-frame predicted or inter-frame predicted. When the prediction type information indicates intra-frame prediction, the entropy decoder 510 extracts a syntax element for intra-frame prediction information (intra-frame prediction mode) of the current block. When the prediction type information indicates inter-frame prediction, the entropy decoder 510 extracts information about syntax elements representing inter-frame prediction information, that is, a motion vector and a reference image to which the motion vector refers.
[0090] In addition, the entropy decoder 510 extracts quantization-related information and extracts information on the quantized transform coefficients of the current block as information on the residual signal.
[0091] The rearrangement unit 515 can change the sequence of 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in the reverse order of the coefficient scan order performed by the video coding device.
[0092] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.
[0093] The inverse transformer 530 reconstructs the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain to generate a residual block of the current block.
[0094] In addition, when the inverse transformer 530 inverse-transforms a partial region (sub-block) of the transform block, the inverse transformer 530 extracts a flag (cu_sbt_flag) for inverse-transforming only the sub-block of the transform block, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the non-inverse-transformed region with the value "0" as the residual signal to generate the final residual block of the current block.
[0095] In addition, when applying MTS, the inverse transformer 530 determines the transform index or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.
[0096] The predictor 540 can include an intra predictor 542 and an inter predictor 544. When the prediction type of the current block is intra prediction, the intra predictor 542 is activated, and when the prediction type of the current block is inter prediction, the inter predictor 544 is activated.
[0097] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block according to the intra prediction mode by using the adjacent reference pixels of the current block.
[0098] The inter - frame predictor 544 determines the motion vector of the current block and the reference image for motion - vector reference by using the syntax element of the inter - frame prediction mode extracted from the entropy decoder 510.
[0099] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 to the prediction block output from the inter - frame predictor 544 or the intra - frame predictor 542. When performing intra - frame prediction on a block to be decoded subsequently, the pixels within the reconstructed current block are used as reference pixels.
[0100] The loop filter unit 560 as an in - loop filter may include a de - blocking filter 562, a SAO filter 564, and an ALF 566. The de - blocking filter 562 performs de - blocking filtering on the boundaries between the reconstructed blocks to remove block artifacts that occur due to block - unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after de - blocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information about the filter coefficients decoded from the bitstream.
[0101] The reconstructed blocks filtered by the de - blocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in an image are reconstructed, the reconstructed image can be used as a reference image for inter - frame prediction of the blocks within the image to be encoded subsequently.
[0102] In some embodiments, the present invention relates to encoding and decoding video images as described above. More specifically, the present invention provides a video encoding and decoding method and apparatus that generate a prediction block of a sub - block partitioned from a current block according to a geometric partitioning mode (GPM), and then adaptively determine a mixing region for weighted summation of the prediction blocks of the sub - blocks.
[0103] The following embodiments can be executed by the predictor 120 in a video encoding device. The following embodiments can also be executed by the predictor 540 in a video decoding device.
[0104] When encoding a current block, the video encoding device can generate signaling information associated with this embodiment from the perspective of optimizing rate - distortion. The video encoding device can encode the signaling information by using the entropy encoder 155 and send the encoded signaling information to the video decoding device. The video decoding device can decode the signaling information associated with the decoding of the current block from the bitstream by using the entropy decoder 510.
[0105] In the following description, the term "target block" may be used interchangeably with the current block or coding unit (CU), or may refer to some regions of the coding unit.
[0106] In addition, a value of true for a flag indicates the case where the flag is set to 1. Additionally, a value of false for a flag indicates the case where the flag is set to 0.
[0107] The following embodiments are described with respect to a video decoding device, but they can also be implemented in a video encoding device in the same or similar manner.
[0108] Figure 6 is a block diagram of a detailed part of a video decoding device according to at least one embodiment of the present invention.
[0109] A video decoding device according to some embodiments may determine a prediction unit and a transform unit, and for a current block corresponding to the determined unit, perform prediction and inverse transformation by using the determined prediction technique and prediction mode to finally generate a reconstructed block of the current block. The operations shown can be performed by an inverse transformer 530, a predictor 540, and an adder 550 in the video decoding device. On the other hand, the same operations as those shown in Figure 6 can be performed by an inverse transformer 165, an image splitter 110, a predictor 120, and an adder 170 in the video encoding device. In this case, the video decoding device utilizes the encoded information parsed from the bitstream, while the video encoding device may utilize the encoded information set from a higher level in terms of minimizing rate distortion. Hereinafter, for ease of description, the embodiments are described centering on the video decoding device. Figure 6 In the following description, for ease of description, the embodiments are described centering on the video decoding device.
[0110] As Figure 5 shown, according to the prediction technique, the predictor 540 includes an intra predictor 542 and an inter predictor 544, but as Figure 6 shown, the predictor 540 may include all or part of a prediction unit determiner 602, a prediction technique determiner 604, a prediction mode determiner 606, and a prediction executor 608.
[0111] When the color format of the input video is the YUV format (YUV420, YUV411, YUV422, YUV444, etc.), the video decoding device may perform prediction and reconstruction of the luminance component, and then may perform prediction and reconstruction of the chrominance component. In other words, the luminance component and the chrominance component may be performed by Figure 6The components shown in are reconstructed sequentially. On the other hand, when the color format of the input video is RGB, the video encoding device may perform a color format conversion from RGB to YUV, and then may encode the converted video. Here, in the case of the YUV format, the color format represents the correspondence between the pixels in the luminance component and the pixels in the chrominance component.
[0112] The prediction unit determiner 602 determines a prediction unit (PU). The prediction technique determiner 604 determines a prediction technique for the prediction unit, such as intra prediction, inter prediction, or intra block copy (IBC) mode, palette mode, etc. The prediction mode determiner 606 determines a detailed prediction mode for the prediction technique. The prediction executor 608 generates a prediction block for the current block according to the determined prediction mode.
[0113] The inverse transformer 530 includes a transform unit determiner 610 and an inverse transform executor 612. The transform unit determiner 610 determines a transform unit (TU) for the inverse quantization signal of the current block, and the inverse transform executor 612 performs an inverse transform on the transform unit represented by the inverse quantization signal to generate a residual signal.
[0114] The adder 550 sums the prediction block and the residual signal to generate a reconstructed block. The reconstructed block is stored in the memory and can be used in the future to predict other blocks.
[0115] The prediction unit determined by the prediction unit determiner 602 may be a sub-block of the current block or a sub-block of a sub-block divided from the current block. In this case, depending on the color format, the prediction unit of the chrominance component may correspond in size to the prediction unit of the luminance component. Alternatively, the prediction units of the luminance component and the chrominance component may be determined separately, and prediction is performed on the prediction unit of the chrominance component.
[0116] The prediction technique determiner 604 determines a prediction technique for the prediction unit. As described above, the prediction technique may be one of inter prediction, intra prediction, IBC mode, and palette mode. In this case, the prediction technique of the chrominance component may be determined to be the same as the prediction technique of the corresponding luminance component without signaling and parsing separate information.
[0117] In one example, when the prediction technique of the current block is not intra prediction, the video decoding device parses 1-bit flag information. For example, when the parsed flag indicates the skip mode, the video decoding device determines that the prediction mode of the current block is the merge mode of inter prediction or the IBC merge mode. The video decoding device may use the prediction signal as the reconstructed signal, thereby omitting the inverse transform process.
[0118] On the other hand, when the parsed flag does not indicate a skip mode for the current block, the prediction technique determiner 604 may parse a series of one-bit flags to determine that the prediction technique for the current block is one of techniques such as inter-frame prediction, intra-frame prediction, IBC mode, palette mode, etc.
[0119] For example, when skip is not applied to the current block and the prediction technique is determined to be inter-frame prediction or IBC mode, the video decoding device parses a 1-bit flag. Depending on the parsed flag, the prediction mode of the current block may be determined to be the merge mode or the advanced motion vector prediction (AMVP) mode.
[0120] The prediction mode determiner 606 determines in detail the prediction mode for the prediction technique.
[0121] As an example, when the prediction technique is inter-frame prediction, the prediction technique determiner 604 may generate a final prediction signal for the current block by parsing a 1-bit flag as follows. For example, according to the geometric partitioning mode, the video decoding device generates a prediction signal by using at least one motion compensation. The video decoding device performs a weighted sum of multiple prediction signals to generate a prediction block, i.e., the prediction signal or predictor for the current block.
[0122] As another example, when the prediction technique is intra-frame prediction, the prediction technique determiner 604 may generate a final prediction signal for the current block by parsing a 1-bit flag as follows. For example, according to the geometric partitioning mode, the video decoding device generates a prediction signal by using multiple intra-frame prediction modes. The video decoding device performs a weighted sum of multiple prediction signals to generate a prediction block for the current block.
[0123] The prediction executor 608 generates a prediction block for the current block according to the determined prediction technique and prediction mode.
[0124] Figure 7 is a flowchart of the geometric partitioning mode according to at least one embodiment of the present invention.
[0125] Figure 7 The exemplary geometric partitioning mode in shows a case where a final prediction signal is generated by performing a weighted sum of two different prediction signals including a prediction signal generated according to at least one motion compensation. When the geometric partitioning mode is applied, the video decoding device parses information for predicting the current block, such as Figure 7 shown.
[0126] The video decoding device decodes a flag indicating the application of the geometric partitioning mode (S700). The video decoding device checks the foregoing flag (S702).
[0127] If the flag indicating the application of the geometric partitioning mode is true (S702 is yes), the video decoding device performs the following steps.
[0128] The video decoding device decodes an index BlendingArea_idx indicating the size of the blending area (S704). In Figure 7 the example of, the step of parsing BlendingArea_idx can be omitted.
[0129] The video decoding device decodes a GPMmode_idx that is an index indicating the geometric partitioning mode (S706).
[0130] For example, the geometric partitioning mode can be defined by using a common look-up table (LUT) between the video encoding device and the video decoding device. The video decoding device can partition the current block into sub-regions according to the parsed index GPMmode_idx.
[0131] The video decoding device decodes indices PredMode_idx_0 and PredMode_idx_1 that are indices representing the prediction modes of each sub-region. The video decoding device generates a prediction block for each sub-region as follows.
[0132] The video decoding device determines whether PredMode_idx_x (x = 0, 1) is greater than maxCandNum (S710). Here, maxCandNum is the maximum number of components of the motion candidate list. If PredMode_idx_x is greater than maxCandNum, PredMode_idx_x indicates an intra prediction mode.
[0133] If PredMode_idx_x is greater than maxCandNum (S710 is yes), the video decoding device uses intra prediction to perform the following steps for predicting the relevant sub-region.
[0134] The video decoding device forms a most probable mode (MPM) list by using the prediction information of the intra prediction block in the adjacent reconstructed region of the current block (S712).
[0135] The video decoding device performs intra prediction on the relevant sub-block (S714).
[0136] The video decoding device decodes the MPM index and derives the intra prediction mode indicated by the MPM index from the MPM list. The video decoding device generates a prediction signal for the relevant sub-block according to the derived intra prediction mode.
[0137] As another example, the video decoding device decodes an index indicating an intra prediction mode without forming an MPM list. According to the intra prediction mode indicated by the index, the video decoding device may generate a prediction signal for a relevant sub-block.
[0138] If PredMode_idx_x is less than or equal to maxCandNum (S710 is NO), the video decoding device performs the following steps for predicting a relevant sub-region using motion compensation.
[0139] The video decoding device forms a motion candidate list (S740). The video decoding device forms a motion candidate list by using motion information in an adjacent reconstructed region of the current block.
[0140] The video decoding device compensates for motion to perform inter prediction for a relevant sub-block (S742). The video decoding device decodes an index indicating a candidate in the motion candidate list. The video decoding device may use the motion information indicated by the decoded index to compensate for motion in the relevant sub-region, and thereby may generate a prediction signal.
[0141] When generating a prediction signal for each sub-region, the video decoding device performs the following steps.
[0142] The video decoding device determines a mixing region (S716).
[0143] The video decoding device performs weighted summation on the mixing region (S718). The video decoding device performs weighted summation to generate a final prediction signal for the current block.
[0144] On the other hand, if a flag indicating the application of the geometric partitioning mode is false (S702 is NO), the video decoding device performs inter prediction for the current block (S730).
[0145] Figure 8 is a flowchart of a geometric partitioning mode according to another embodiment of the present invention.
[0146] Figure 8 The exemplary geometric partitioning mode in shows a case where a final prediction signal is generated by performing weighted summation on two different prediction signals generated according to intra prediction. When the geometric partitioning mode is applied, the video decoding device parses information for predicting the current block, such as Figure 8 shown.
[0147] The video decoding device decodes a flag indicating the application of the geometric partitioning mode (S800). The video decoding device checks the aforementioned flag (S802).
[0148] If the flag indicating the application of the geometric partitioning mode is true (S802 is YES), the video decoding device performs the following steps.
[0149] The video decoding device decodes the BlendingArea_idx indicating the size of the blending area (S804). In Figure 8 the example of
[0150] The video decoding device decodes the GPMmode_idx which is an index indicating the geometric partitioning mode (S806).
[0151] For example, the geometric partitioning mode can be defined by using a common look-up table between the video encoding device and the video decoding device. The video decoding device partitions the current block into sub-regions according to the parsed index GPMmode_idx.
[0152] The video decoding device decodes the PredMode_idx_0 and PredMode_idx_1 which are indices representing the prediction mode of each sub-region (S808). The video decoding device generates an intra prediction block for each sub-region as follows.
[0153] The video decoding device composes an MPM list by using the prediction information of the intra prediction blocks within the adjacent reconstructed regions of the current block (S810).
[0154] The video decoding device performs intra prediction on the relevant sub-blocks (S812).
[0155] The video decoding device decodes the MPM index and derives the intra prediction mode indicated by the MPM index from the MPM list. The video decoding device generates a prediction signal for the relevant sub-blocks according to the derived intra prediction mode.
[0156] As another example, the video decoding device decodes the index indicating the intra prediction mode without composing an MPM list. According to the intra prediction mode indicated by the index, the video decoding device can generate a prediction signal for the relevant sub-blocks.
[0157] The video decoding device determines the blending area (S814).
[0158] The video decoding device performs weighted summation on the blending area (S816). The video decoding device performs weighted summation to generate the final prediction signal of the current block.
[0159] On the other hand, if the flag indicating the application of the geometric partitioning mode is false (NO in S802), the video decoding device performs intra prediction of the current block (S830).
[0160] In Figure 7 and Figure 8In an example, during the weighted summation process for generating the final prediction signal, the blending region can be determined as follows.
[0161] In one example, based on the size of the current block and the geometric partitioning mode determined by GPMmode_idx, the initial blending region is determined as a preset fixed region. In this case, as described above, in Figure 7 or Figure 8 example, the step of parsing BlendingArea_idx can be omitted.
[0162] As another example, the final blending region can be explicitly determined based on BlendingArea_idx.
[0163] As yet another example, the final blending region can be implicitly determined by utilizing the blending region determination process. In this case, as described above, in Figure 7 or Figure 8 example, the step of parsing BlendingArea_idx can be omitted.
[0164] The determination of the blending region and the determination of the blending matrix are described in detail below.
[0165] In one example, the video decoding device generates a prediction signal for each sub-region. For blending, prediction signals for each sub-region are generated for the entire region of the current block. The video decoding device explicitly or implicitly determines the final blending region during the weighted summation process for generating the final prediction signal of the current block. Then, based on the determined final blending region, the video decoding device calculates the blending matrix to be utilized in the weighted summation process.
[0166] For example, the video decoding device calculates the blending matrix W B . In this case, each weight value W B (i, j) in the blending matrix can be an integer value from 0 to 2 n . Here, n is an integer greater than or equal to 0, which can be determined according to the size of the current block and / or the geometric partitioning mode. The geometric partitioning mode can be parsed or deduced.
[0167] Figure 9 is a schematic diagram showing the initial blending region according to at least one embodiment of the present invention.
[0168] As another example, based on the size of the current block, the geometric partitioning mode, and / or the color component of the current block, the video decoding device determines the initial blending region τ, as Figure 9 shown.
[0169] In one example, the video decoding device will Figure 9The initial mixing region τ shown in the figure is determined as the final mixing region. In this case, the process of determining the mixing region is omitted. Additionally, when the initial mixing region is determined as the final mixing region, the decoding process of BlendingArea_idx in the examples of Figure 7 and Figure 8 can be omitted.
[0170] In Figure 9 the example, the current block is partitioned into a left (or upper) sub-region and a right (or lower) sub-region according to the geometric segmentation boundary. Hereinafter, the two sub-regions are respectively referred to as the first sub-region and the second sub-region. The prediction signal in the first sub-region is denoted as P0, and the prediction signal in the second sub-region is denoted as P1. As described above, for mixing, prediction signals for each sub-region are generated for the entire region of the current block. The mixing matrix is applied to the final mixing region. In this case, the mixing matrix W B can be applied to P0 and 1 - W B to P1, or vice versa. Hereinafter, P0 can be used interchangeably with the first prediction signal, and P1 can be used interchangeably with the second prediction signal.
[0171] When the initial mixing region τ is determined as the final mixing region, the video decoding device determines the mixing matrix W B , in the form of a weighting matrix, including gradually increasing and decreasing the weights centered on the geometric segmentation boundary, as shown in Figure 10 . In this case, the value "a" can be an integer value determined based on the size of the current block, the geometric partitioning pattern, the color component, and / or the final mixing region.
[0172] In one example, based on the initial mixing region τ shown in Figure 9 , the video decoding device can determine the final mixing region bτ. In this case, the value "b" can be 2 k (where k is an integer). The video decoding device can determine the final mixing regions on both sides of the geometric segmentation boundary as regions of different sizes. In this case, the range of the value "b" can be determined based on the size of the current block, the geometric partitioning pattern, and / or the color component.
[0173] Hereinafter, the mixing region overlapping with the first sub-region centered on the geometric segmentation boundary is represented by τ0, and the mixing region overlapping with the second sub-region is represented by τ1.
[0174] In one example, the video decoding device can implicitly determine the final mixing region without decoding BlendingArea_idx as shown in Figure 7 and Figure 8 .
[0175] As another example, a video decoding device decodes a flag indicating whether to implicitly determine a final blending area on a per-CU basis. If the flag is true, the video decoding device implicitly determines the final blending area of the current block and determines the corresponding blending matrix. On the other hand, if the aforementioned flag is false, the video decoding device may, for example, determine the initial blending area as the final blending area as described above. Alternatively, the video decoding device may perform decoding of BlendingArea_idx and determine the final blending area based on BlendingArea_idx.
[0176] A method for implicitly determining a final blending area and a blending matrix is described below.
[0177] Figure 11a and Figure 11b is a flowchart for determining an implicit final blending area according to some embodiments of the present invention.
[0178] In Figure 11a 's example, the video decoding device determines the strength of the geometric segmentation boundary based on a predicted value, and determines the final blending area based on the determined strength of the geometric segmentation boundary.
[0179] The video decoding device determines the strength of the geometric segmentation boundary based on a predicted value (S1100).
[0180] The video decoding device determines whether the geometric segmentation boundary is a strong edge (S1102).
[0181] When the geometric segmentation boundary is a strong edge (S1102 is yes), the video decoding device sets the final blending area to 0 (S1104). That is, the video decoding device does not set the final blending areas τ0 and τ1. When the final blending areas τ0 and τ1 are determined to be 0, each weight of the blending matrix W B is composed of 0 and 2 centered on the geometric segmentation boundary k and.
[0182] On the other hand, when the geometric segmentation boundary is not a strong edge (S1102 is no), the video decoding device determines the final blending area based on the predicted value (S1106).
[0183] The following describes the detailed steps for determining the strength of the geometric segmentation boundary based on the predicted value.
[0184] The video decoding device may determine the strength S of the geometric segmentation boundary by using the initial prediction signals P0 and P1 of the sub-regions g . Here, the initial blending areas τ0 i and τ1 iand the prediction modes PM_idx_0 and PM_idx_1 of the sub-regions to generate initial prediction signals P0 and P1.
[0185] Figures 12a to 12c is a flowchart showing the determination of the intensity at the geometric segmentation boundary according to some embodiments of the present invention.
[0186] By using the prediction signals of two sub-regions centered on the geometric segmentation boundary, the video decoding device can calculate the intensity at the geometric segmentation boundary. The video decoding device can determine the intensity of the geometric segmentation boundary by comparing the initial prediction signal values of the regions adjacent to the geometric segmentation boundary or the inner boundary parallel to the geometric segmentation boundary, as Figures 12a to 12c shown.
[0187] On the other hand, the initial mixing regions τ0 i and τ1 i on both sides of the geometric segmentation boundary can be the same, as Figures 12a to 12c shown. In addition, the initial mixing regions τ0 i and τ1 i on both sides of the geometric segmentation boundary can be the same as the final mixing regions τ0 and τ1 on both sides. The final mixing regions τ0 and τ1 on both sides of the geometric segmentation boundary can be different.
[0188] In one example, the video decoding device determines the edge intensity by comparing the sample values of the initial prediction signals P0 and P1 located at the geometric segmentation boundary, as Figure 12a shown. The video decoding device can determine the edge intensity based on Equation 1 or Equation 2.
[0189] [Equation 1]
[0190] |P0(c y ) - P1(c y )| < th c1
[0191] [Equation 2]
[0192]
[0193] In Equation 1 and Equation 2, the thresholds th C1 and th C2 can be pre-determined based on the protocol between the video encoding device and the video decoding device. Alternatively, the thresholds can be signaled / parsed or determined based on the quantization parameter of the current block.
[0194] For each position, when there are "m" or fewer cases that satisfy Equation 1, that is, when there are more than "m" cases that do not satisfy Equation 1, the video decoding device determines the geometric segmentation boundary as a strong edge. In this case, the threshold "m" can be determined based on the size of the current block. Alternatively, if Equation 2 is not satisfied, the video decoding device may determine the geometric segmentation boundary as a strong edge.
[0195] When the geometric segmentation boundary is not determined as a strong edge according to Equation 1 or Equation 2, the video decoding device can determine the edge strength by comparing the initial prediction signal values in the region adjacent to the inner boundary parallel to the geometric segmentation boundary, as Figure 12b and Figure 12c shown.
[0196] For example, by using the inner boundary existing in the first sub-region (i.e., on the left side of the geometric segmentation boundary), as Figure 12b shown, the video decoding device can determine the edge strength according to Equation 3 or Equation 4.
[0197] [Equation 3]
[0198] |P0(l y ) - P1(l y )| < th l1
[0199] [Equation 4]
[0200]
[0201] In Equations 3 and 4, the threshold th l1 and th l2 can be predefined based on the protocol between the video encoding device and the video decoding device. Alternatively, the threshold can be signaled / parsed or determined based on the quantization parameter of the current block.
[0202] For each position, when there are "m" or fewer cases that satisfy Equation 3, that is, when there are more than "m" cases that do not satisfy Equation 3, the video decoding device determines the geometric segmentation boundary as a strong edge. In this case, the threshold "m" can be determined based on the size of the current block. Alternatively, if Equation 4 is not satisfied, the video decoding device may determine the geometric segmentation boundary as a strong edge.
[0203] On the other hand, when there are "m" or fewer cases that satisfy Equation 3, that is, when there are more than "m" cases that do not satisfy Equation 3, the video decoding device can determine the final mixing region τ0 of the first sub-partition as 0. Alternatively, if Equation 4 is not satisfied, the video decoding device can determine the final mixing region τ0 of the first sub-region as 0.
[0204] As another example, by utilizing the inner boundary existing in the second sub-region (i.e., on the right side of the geometric segmentation boundary), as Figure 12c shown, the video decoding device can determine the edge strength according to Equation 5 or Equation 6.
[0205] [Equation 5]
[0206] |P0(r y ) - P1(r y )| < th r1
[0207] [Equation 6]
[0208]
[0209] In Equation 5 and Equation 6, the threshold th can be predefined based on the protocol between the video encoding device and the video decoding device r1 and th r2 . Alternatively, the threshold can be signaled / parsed or determined based on the quantization parameter of the current block.
[0210] For each position, when there are "m" or fewer cases satisfying Equation 5, i.e., when there are more than "m" cases not satisfying Equation 5, the geometric segmentation boundary is determined as a strong edge. In this case, the threshold "m" can be determined based on the size of the current block. Alternatively, if Equation 6 is not satisfied, the video decoding device can determine the geometric segmentation boundary as a strong edge.
[0211] On the other hand, when there are "m" or fewer cases satisfying Equation 5, i.e., when there are more than "m" cases not satisfying Equation 5, the video decoding device can set the final mixing region τ1 of the second sub-partition to 0. Alternatively, if Equation 6 is not satisfied, the video decoding device can set the final mixing region τ1 of the second sub-region to 0.
[0212] Alternatively, as in the example of Figure 12a , when the geometric segmentation boundary is classified as a strong edge based on the predicted value at the geometric segmentation boundary, the video decoding device can further perform a process of determining the strength of the geometric segmentation boundary based on the slope.
[0213] In the example of Figure 11b , the video decoding device determines the strength of the geometric segmentation boundary based on the slope and determines the final mixing region based on the determined strength of the geometric segmentation boundary.
[0214] The video decoding device determines the strength at the geometric segmentation boundary based on the slope (S1120).
[0215] The video decoding device determines whether the geometric segmentation boundary is a strong edge (S1122).
[0216] When the geometric segmentation boundary is a strong edge (S1122 is yes), the video decoding device sets the final mixing region to 0 (S1124). That is, the video decoding device sets the final mixing regions τ0 and τ1 of the first sub-region and the second sub-region to 0. When the final mixing regions τ0 and τ1 are determined to be 0, each weight of the mixing matrix W B consists of 0 and 2 centered on the geometric segmentation boundary k .
[0217] On the other hand, when the geometric segmentation boundary is not a strong edge (S1122 is no), the video decoding device determines the final mixing region based on the predicted value (S1126).
[0218] The following describes the detailed steps for determining the strength of the geometric segmentation boundary based on the slope.
[0219] The video decoding device can determine the slope-based strength S of the geometric segmentation boundary by using the initial prediction signals P0 and P1 of the sub-regions g . Here, the initial prediction signals P0 and P1 can be generated by using the prediction modes PM_idx_0 and PM_idx_1 of the sub-regions as described above. To calculate the slope, the video decoding device uses the initial prediction signals within the initial mixing region τ.
[0220] By using the prediction signals of the two sub-regions centered on the geometric segmentation boundary, the video decoding device can calculate the strength at the geometric segmentation boundary. The video decoding device can determine the strength S of the geometric segmentation boundary by using the sum of the slopes according to Figure 13a , Figure 13b and Equation 7 g .
[0221] [Equation 7]
[0222]
[0223] After calculating the strength of the geometric segmentation boundary by using the sum of the slopes in the sub-region according to Equation 7, the video decoding device compares the strength of the geometric segmentation boundary with a threshold. When the strength of the geometric segmentation boundary is greater than or equal to the threshold, the video decoding device sets the geometric segmentation boundary as a strong edge. In this case, the threshold can be predefined according to the protocol between the video encoding device and the video decoding device. Alternatively, the threshold can be signaled / parsed or determined based on the quantization parameter of the current block.
[0224] When the geometric segmentation boundary is determined as a strong edge, the video decoding device may determine the final blending regions τ0 and τ1 of the sub-regions as 0.
[0225] As an example, in Figure 13a , Figure 13b and Equation 7, the positions of p 2,y and q 2,y can be the intermediate positions between P 1,y and p 3,y respectively, and the intermediate positions between q 1,y and q 3,y .
[0226] As another example, in Figure 13a , Figure 13b and Equation 7, the positions of p 2,y and q 2,y can be those positions that are at a sample distance of τ / 2 away from p 1,y and p 3,y . For example, if τ is odd, the positions of p 2,y and q 2,y can be those positions that are approximately at a sample distance of (τ / 2) away from p 1,y and p 3,y . Alternatively, if τ is odd, the positions of p 2,y and q 2,y can be the average values of the samples at positions that are at a sample distance of approximately (τ / 2) away from p 1,y and p 3,y respectively, and the average values of the samples at positions that are at a sample distance of approximately (τ / 2 + 0.5) away from p 1,y and p 3,y .
[0227] As another example, in Figure 13a , Figure 13b and Equation 7, the samples at positions p 2,y and q 2,y can be those samples that are respectively filtered from the samples at positions that are τ / 2 away from p 1,y and p 3,y . For the filtering, one of the filters such as a Gaussian filter, a smoothing filter, etc. can be used.
[0228] The following describes the detailed steps (S1106 or S1126) for determining the final blending region based on the predicted value.
[0229] When the geometric segmentation boundary is determined as a strong edge, the video decoding device may respectively generate the final blending regions τ0 and τ1 of the two sub-regions centered on the geometric segmentation boundary. As described above, in Figure 12b or Figure 12cIn the example where, when the final mixing region τ0 of the first sub-region or the final mixing region τ1 of the second sub-region is determined to be 0, the video decoding device only generates the final mixing regions of the remaining sub-regions.
[0230] In the following, by using Figures 14a to 14d the example of, the process for generating the final mixing region τ0 of the first sub-region is described. The video decoding device can generate the final mixing region τ1 of the second sub-region in the same manner as in the example of FIG. 14.
[0231] Figures 14a to 14d is a schematic diagram showing the generation of the final mixing region according to some embodiments of the present invention.
[0232] The video decoding device sets the initial mixing candidate region to the maximum mixing region maxτ applicable to the current block. The maximum mixing region maxτ and the minimum mixing region minτ can be determined based on the size of the current block, the geometric partitioning pattern, and / or the color component. Alternatively, the maximum mixing region maxτ can be explicitly determined based on the decoded BlendingArea_idx.
[0233] The video decoding device can determine the final mixing region in the region that is at a mixing candidate region distance from the geometric segmentation boundary according to Equation 8 and Equation 9.
[0234] [Equation 8]
[0235] |P0(t y ) - P1(t y )| < th t1
[0236] [Equation 9]
[0237]
[0238] In Equation 8 and Equation 9, the thresholds th t1 and th t2 can be predefined by the protocol between the video encoding device and the video decoding device. Alternatively, the thresholds can be signaled / parsed or determined based on the quantization parameter of the current block.
[0239] When there are "m" or fewer cases satisfying Equation 8 for each position, the video decoding device determines the current mixing candidate region as the final mixing region τ0. In this case, the threshold "m" can be determined based on the size of the current block. Alternatively, if Equation 9 is satisfied, the video decoding device can determine the current mixing candidate region as the final mixing region τ0.
[0240] If Equation 8 or Equation 9 is not satisfied, the video decoding device may repeat the above process by changing the candidate blending region. For example, the video decoding device may reduce the candidate blending region to half of the previous region, as Figures 14a to 14d shown.
[0241] When the final blending regions τ0 and τ1 are determined as described above, the video decoding device may determine the respective weights of the blending matrix W B in each region centered on the geometric segmentation boundary. For example, the video decoding device may set the weight to 2 k-1 in the region closest to the geometric segmentation boundary, and set the weight to 2 k and 0 in the region farthest from the geometric segmentation boundary. When different blending regions are determined in two sub-regions centered on the geometric segmentation boundary, the video decoding device may organize gradually decreasing weights, as Figure 15 shown.
[0242] On the other hand, Figures 14a to 14d and Figure 15 examples, as well as Equation 8 and Equation 9 as described above, may be applied to both the left and right regions or the top and bottom regions centered on the geometric segmentation boundary of the current block.
[0243] As another example, by utilizing the index BlendingArea_idx decoded according to Figure 7 and Figure 8 examples, the video decoding device may explicitly determine the final blending region.
[0244] For example, the video decoding device decodes the index of each of the blending regions τ0 and τ1, and determines the final blending region based on the decoded index. In this case, each blending region mapped to the index may be an integer multiple of the initial blending region shown in Figure 9 . Each blending region mapped to the index may contribute to the organization of the lookup table. The lookup table may be preset by a protocol between the video encoding device and the video decoding device.
[0245] As another example, the index may combinatorially map the two blending regions τ0 and τ1. The blending regions mapped to the index may be organized as a lookup table. The lookup table may be preset by a protocol between the video encoding device and the video decoding device. The video decoding device may parse an index, and may utilize the index to determine the final blending regions of the two regions from the lookup table.
[0246] A method for implicitly determining the final blending region based on template matching is described below.
[0247] Figure 16It is a schematic diagram showing the determination of the final mixing region based on template matching according to at least one embodiment of the present invention.
[0248] The video decoding device determines the initial mixing region as a predetermined fixed region based on the size of the current block, the geometric partitioning mode, and / or the color component.
[0249] By using the adjacent reconstructed regions of the current block, the adjacent template regions of the initial prediction signals P0 and P1 of the current block, and the geometric partitioning mode of the current block, the video decoding device extends the geometric segmentation boundary to the adjacent templates of each initial prediction signal, as Figure 16 shown. The video decoding device generates candidate mixing regions based on the initial mixing region. After generating the mixing matrix for each candidate mixing region, the video decoding device uses the generated mixing matrix to perform a weighted sum of the adjacent templates, thereby generating a candidate template corresponding to each candidate mixing region. The video decoding device compares the candidate template with the template in the reconstructed region of the current block based on a cost function, and reorders the candidate templates according to the cost function value. Based on the reordered candidate templates, the video decoding device can determine the candidate mixing region corresponding to the candidate template with the minimum cost as the final mixing region, and can use the mixing matrix corresponding to the final mixing region.
[0250] In this case, the size of the template region can be determined based on the size of the current block, or a preset size can be used. In Figure 16 the example, the size of the template region is determined by the size of the current block, "a" is the width of the template on the left side of the current block, and "b" is the height of the template on the top side of the current block. In addition, for the cost function, metrics such as Mean Square Errors (MSE), Sum of Absolute Differences (SAD), Sum of Absolute Transformed Differences (SATD), etc. can be used.
[0251] In one example, the video decoding device generates the final prediction signal P of the current block by performing a weighted sum as shown in Equation 10 based on the mixing matrix W B and the initial prediction signals P0 and P1 calculated by the mixing region determination process. G .
[0252] [Equation 10]
[0253] P G (i,j)=(W B (i,j)×P0+(2 k -W B(i,j)) × P1+(2 k-1 )) >> k
[0254] Then, the video decoding device decodes the residual signal and sums the residual signal and the final prediction signal to generate a reconstructed block of the current block.
[0255] Using the chrominance component, the video decoding device generates the final mixing region of the chrominance block by sampling the final mixing region determined in the co-located luma block according to the color format. As another example, for the chrominance component, the above process may be performed to determine the mixing region of each sub-region centered on the geometric partitioning boundary.
[0256] The video decoding device calculates the mixing matrix by using the mixing region generated as described above, and then uses the mixing matrix and the initial prediction signal to generate the final prediction signal of the chrominance block.
[0257] Figure 17 is a flowchart of a method for encoding a current block by a video encoding device according to at least one embodiment of the present invention.
[0258] The video encoding device determines the geometric partitioning mode of the current block (S1700). For example, in terms of rate-distortion optimization, the video encoding device may determine the geometric partitioning mode of the current block.
[0259] The video encoding device partitions the current block into a first sub-region and a second sub-region centered on the geometric partitioning boundary according to the geometric partitioning mode (S1702).
[0260] The video encoding device obtains the prediction modes of the first sub-region and the second sub-region (S1704). The prediction modes of the first sub-region and the second sub-region may be an inter prediction mode or an intra prediction mode.
[0261] The video encoding device generates a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction modes of the first sub-region and the second sub-region (S1706).
[0262] The video encoding device determines a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and the initial mixing region (S1708). Here, the initial mixing region may be determined based on the size and geometric partitioning mode of the current block.
[0263] The video encoding device determines the mixing matrix in the final mixing region (S1710).
[0264] The video encoding device generates a first final prediction signal of a current block by performing weighted summation of a first prediction signal and a second prediction signal in a final mixing area by using a mixing matrix (S1712).
[0265] The video encoding device obtains a prediction mode of the current block (S1714). The prediction mode of the current block can be an inter prediction mode or an intra prediction mode.
[0266] The video encoding device generates a second final prediction signal of the current block according to the prediction mode (S1716).
[0267] The video encoding device determines a flag indicating the application of a geometric partitioning mode based on the first final prediction signal and the second final prediction signal (S1718).
[0268] From the perspective of rate-distortion optimization, the video encoding device can determine the foregoing flag. For example, when the first final prediction signal is the best, the video encoding device sets the flag to true. Alternatively, when the second final prediction signal is the best, the video encoding device sets the flag to false.
[0269] The video encoding device encodes the flag (S1720).
[0270] Depending on the value of the flag, the video encoding device subtracts the first final prediction signal or the second final prediction signal from the original block of the current block to generate a residual signal. Then, the video encoding device encodes the residual signal.
[0271] Although the steps in the respective flowcharts are described as being executed sequentially, these steps merely illustrate the technical ideas of some embodiments of the present invention. Therefore, those of ordinary skill in the art to which the present invention pertains can perform the steps by changing the order described in the respective drawings or by executing two or more steps in parallel. Therefore, the steps in the respective flowcharts are not limited to the order shown in the chronological order.
[0272] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in the present invention are labeled as "…… unit" to highlight the possibility of their independent implementation.
[0273] On the other hand, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-volatile recording medium, which can be read and executed by one or more processors. The non-volatile recording medium can include various types of recording devices that store data in a form readable by a computer system. For example, the non-volatile recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.
[0274] Although exemplary embodiments of the present invention have been described for illustrative purposes, those of ordinary skill in the art to which the present invention pertains should understand that various modifications, additions, and substitutions can be made without departing from the spirit and scope of the present invention. Accordingly, the embodiments of the present invention have been described for the sake of simplicity and clarity. The scope of the technical idea of the embodiments of the present invention is not limited by the examples. Accordingly, those of ordinary skill in the art to which the present invention pertains should understand that the scope of the present invention should not be limited by the embodiments clearly described above, but by the claims and their equivalents.
[0275] Reference Numerals
[0276] 120: Predictor
[0277] 540: Predictor
[0278] 602: Prediction Unit Determiner
[0279] 604: Prediction Technique Determiner
[0280] 606: Prediction Mode Determiner
[0281] 608: Prediction Executor.
[0282] Cross-Reference to Related Applications
[0283] This application claims the priority and benefits of Korean Patent Application No. 10-2022-0156601, filed on November 21, 2022, and Korean Patent Application No. 10-2023-0153800, filed on November 8, 2023, the entire contents of each of which are incorporated herein by reference.
Claims
1. A method for reconstructing a current block by a video decoding device, the method comprising: Decoding the geometric partitioning mode of the current block; Partitioning the current block into a first sub-region and a second sub-region centered on a geometric segmentation boundary according to the geometric partitioning mode; Generating a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction modes of the first sub-region and the second sub-region; Determining a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and an initial mixing region; Determining a mixing matrix in the final mixing region; And Generating a final prediction signal of the current block by performing weighted summation of the first prediction signal and the second prediction signal in the final mixing region by using the mixing matrix.
2. The method according to claim 1, further comprising: Determining an initial mixing region based on the size of the current block and the geometric partitioning mode of the current block.
3. The method according to claim 1, wherein, Determining the final mixing region includes: Determining whether the geometric segmentation boundary is a strong edge by comparing the sample values of a first initial prediction signal and a second initial prediction signal located at the geometric segmentation boundary; and Checking whether the geometric segmentation boundary is a strong edge.
4. The method according to claim 3, wherein, When the geometric segmentation boundary is not a strong edge, determining the final mixing region includes: Secondarily determining whether the geometric segmentation boundary is a strong edge by comparing the sample values of the first initial prediction signal and the second initial prediction signal located at a first inner boundary in the first sub-region.
5. The method according to claim 4, wherein When the geometric segmentation boundary is not a strong edge, determining the final mixing region includes: Thirdly determining whether the geometric segmentation boundary is a strong edge by comparing the sample values of the first initial prediction signal and the second initial prediction signal located at a second inner boundary in the second sub-region.
6. The method according to claim 4, further comprising: Checking whether the geometric segmentation boundary determined secondarily is a strong edge, wherein when the geometric segmentation boundary determined secondarily is a strong edge, determining the final mixing region further includes: Setting the size of the mixing region of the first sub-region to 0.
7. The method according to claim 4, further comprising: Checking whether the geometric segmentation boundary determined secondarily is a strong edge, wherein when the geometric segmentation boundary determined secondarily is not a strong edge, determining the final mixing region further includes: Determining the final mixing region by comparing the sample values of the first initial prediction signal and the second initial prediction signal, the first initial prediction signal and the second initial prediction signal being located in the first sub-region and being at a mixing candidate distance from the geometric segmentation boundary.
8. The method according to claim 5, further comprising: Checking whether the geometric segmentation boundary determined thirdly is a strong edge, wherein when the geometric segmentation boundary determined thirdly is a strong edge, determining the final mixing region further includes: Setting the size of the mixing region of the second sub-region to 0.
9. The method according to claim 3, wherein When the geometric segmentation boundary is a strong edge, determining the final mixing region further includes: Calculating a slope by using the sample values of the first initial prediction signal and the second initial prediction signal within the initial mixing region.
10. The method according to claim 10, wherein Determining the final mixing region further includes: Calculate the strength of the geometric segmentation boundary based on the slope; and Determine whether the geometric segmentation boundary is a strong edge by utilizing the strength of the geometric segmentation boundary.
11. The method according to claim 1, wherein, Determining the mixing matrix is characterized in that each weight in the mixing matrix has an integer value from 0 to 2 k where k is an integer value equal to or greater than 0.
12. The method according to claim 1, wherein Determining the mixing matrix includes: Organize the mixing matrix to include weights that increase and decrease around the geometric segmentation boundary.
13. A method for encoding a current block by a video encoding device, the method comprising: Determine the geometric partitioning pattern of the current block; Partition the current block into a first sub-region and a second sub-region centered on the geometric segmentation boundary according to the geometric partitioning pattern; Generate a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction patterns of the first sub-region and the second sub-region; Determine a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and the initial mixing region; Determine the mixing matrix in the final mixing region; And Generate a first final prediction signal of the current block by weighted summing the first prediction signal and the second prediction signal in the final mixing region by utilizing the mixing matrix.
14. The method according to claim 13, further comprising: Obtain the prediction pattern of the current block; And Generate a second final prediction signal of the current block based on the prediction pattern.
15. The method according to claim 14, further comprising: Determine a flag indicating the application of the geometric partitioning pattern based on the first final prediction signal and the second final prediction signal; And Encode the flag.
16. The method according to claim 13, further comprising: Obtain the prediction patterns of the first sub-region and the second sub-region; And Determine the initial mixing region based on the size of the current block and the geometric partitioning pattern of the current block.
17. A computer-readable recording medium storing a bitstream generated by a video encoding method, the video encoding method comprising: Determine the geometric partitioning pattern of the current block; Partition the current block into a first sub-region and a second sub-region centered on the geometric segmentation boundary according to the geometric partitioning pattern; Generate a first prediction signal for the first sub-region and a second prediction signal for the second sub-region according to the prediction patterns of the first sub-region and the second sub-region; Determine a final mixing region for weighted summation of the first prediction signal and the second prediction signal based on the first prediction signal, the second prediction signal, and the initial mixing region; Determine the mixing matrix in the final mixing region; And Generate a final prediction signal of the current block by weighted summing the first prediction signal and the second prediction signal in the final mixing region by utilizing the mixing matrix.
Citation Information
Patent Citations
Method and system for hand gesture-based control of a device
KR1020220156601A
Solar lead-acid battery charging circuit using voltage regulator and CDS cell
KR1020230153800A