Video encoding / decoding method and apparatus for predicting and modifying partition structure of coding tree unit

By using convolutional neural network to predict and modify the segmentation structure of the encoding tree unit, the video encoding method is optimized, and the encoding efficiency problem of high-resolution and high-frame rate video data is solved, and a more efficient encoding and decoding process is achieved.

CN120303931APending Publication Date: 2025-07-11HYUNDAI MOTOR CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380081324.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-17
Filing Date
2023-11-23
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing video encoding technologies are insufficient in coding efficiency when processing high-resolution and high-frame rate video data, and require more efficient encoding methods to reduce the hardware resources required for storage and transmission.

Method used

The convolutional neural network is used to predict and modify the segmentation structure of the encoding tree unit. By merging or segmenting the encoding unit, the encoding efficiency is improved, and the video data is analyzed using the convolutional neural network to determine the optimal segmentation method.

Benefits of technology

Improves the efficiency of video encoding and decoding, reduces the amount of bitstream data required for storage and transmission, and improves the image enhancement effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303931A_ABST
    Figure CN120303931A_ABST
Patent Text Reader

Abstract

A video encoding / decoding method and apparatus are provided. A video decoding method according to the present disclosure comprises the steps of: predicting a partition structure of a current coding tree unit; determining whether to merge the current coding units within the current coding tree unit using the division structure of the current coding tree unit, and based on the determination of whether to merge, merging the current coding units or skipping the merging of the current coding units; determining whether to divide the current coding unit based on whether the current coding unit has been merged, and dividing the current coding unit or skipping the division of the current coding unit to determine a division structure of the current coding tree unit based on the determination of whether to divide; and generating a prediction block of the current coding unit using the partition structure of the current coding tree unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] In some embodiments, the present disclosure relates to video encoding / decoding methods and apparatuses for predicting and modifying a segmentation structure of a coding tree unit. More specifically, the present disclosure relates to video encoding / decoding methods and apparatuses for predicting the segmentation structure of a coding tree unit using a convolutional neural network and modifying the predicted segmentation structure of the coding tree unit. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Since video data has a large amount of data compared to audio or still picture data, video data requires a large amount of hardware resources (including memory) to store or transmit the video data without processing for compression.

[0004] Therefore, an encoder is typically used to compress and store or transmit video data. A decoder receives the compressed video data, decompresses the received compressed video data, and plays the decompressed video data. Video compression techniques include H.264 / Advanced Video Coding (AVC), High Efficiency Video Coding (HEVC), and Versatile Video Coding (VVC), and VVC has an improved coding efficiency of about 30% or more compared to HEVC.

[0005] However, as the image size, resolution, and frame rate gradually increase, the amount of data to be encoded also increases. Therefore, there is a need to provide a new compression technique with higher coding efficiency and improved image enhancement effects than existing compression techniques.

[0006] A video encoding apparatus recursively divides a coding tree unit (CTU) to determine a segmentation structure of the coding tree unit. The segmentation structure of the coding tree unit is a structure of coding units (CUs) that constitute the coding tree unit. When determining the segmentation structure of the coding tree unit, the video encoding apparatus sends information about the segmentation structure of the coding tree unit to the video decoding apparatus. The video decoding apparatus uses the information about the segmentation structure of the coding tree unit to divide the coding tree unit into coding units and reconstruct samples of each coding unit. It is necessary for the video encoding apparatus to send only a part of the information about the segmentation structure of the coding tree unit to improve the coding efficiency.

[0007]

Disclosure of the Present Invention

[0008]

Technical Problem

[0009] The present disclosure seeks to provide, in some embodiments, methods and apparatuses for predicting the segmentation structure of a coding tree unit using reference information and a convolutional neural network.

[0010] The present disclosure also seeks to provide a method and apparatus for determining a segmentation structure of a final coding tree unit by modifying a predicted segmentation structure of a coding tree unit.

[0011] The present disclosure also seeks to provide a method and apparatus for improving video encoding / decoding efficiency.

[0012] The present disclosure also seeks to provide a recording medium storing a bitstream generated by the video encoding / decoding method or apparatus of the present disclosure.

[0013] The present disclosure also seeks to provide a method and apparatus for transmitting a bitstream generated by the video encoding / decoding method or apparatus of the present disclosure. Summary of the Invention

[0014] In at least one embodiment, the present disclosure provides a video decoding method, including: predicting a segmentation structure of a current coding tree unit, using the segmentation structure of the current coding tree unit to determine whether to merge or not merge a current coding unit within the current coding tree unit, and based on the determined merging or non-merging of the current coding unit, merging the current coding unit or omitting the merging of the current coding unit, determining whether to split or not split the current coding unit based on the determined merging or non-merging of the current coding unit, and determining the segmentation structure of the current coding tree unit by splitting the current coding unit or omitting the splitting of the current coding unit based on the determined splitting or non-splitting, and generating a prediction block of the current coding unit using the segmentation structure of the current coding tree unit.

[0015] In another embodiment, the present disclosure provides a video encoding method, including: predicting a segmentation structure of a current coding tree unit, using the segmentation structure of the current coding tree unit to determine whether to merge or not merge a current coding unit within the current coding tree unit, and based on the determined merging or non-merging of the current coding unit, merging the current coding unit or omitting the merging of the current coding unit, determining whether to split or not split the current coding unit based on the determined merging or non-merging of the current coding unit, and determining the segmentation structure of the current coding tree unit by splitting the current coding unit or omitting the splitting of the current coding unit based on the determined splitting or non-splitting, and generating a prediction block of the current coding unit using the segmentation structure of the current coding tree unit.

[0016] In addition, the present disclosure may provide a method for transmitting a bitstream generated by the video encoding method or apparatus according to the present disclosure.

[0017] In addition, the present disclosure may provide a recording medium storing a bitstream generated by the video encoding method or apparatus according to the present disclosure.

[0018] In addition, the present disclosure may provide a recording medium storing a bitstream received, decoded, and used for reconstructing a video by a video decoding apparatus according to the present disclosure.

[0019]

Beneficial Effects

[0020] As described above, in some embodiments, the present disclosure may provide a method and apparatus for predicting a segmentation structure of a coding tree unit using reference information and a convolutional neural network.

[0021] In addition, the present disclosure may provide a method and apparatus for determining a final segmentation structure of a coding tree unit by modifying the predicted segmentation structure of the coding tree unit.

[0022] In addition, the present disclosure may provide a method and apparatus for improving video encoding / decoding efficiency.

[0023] The effects obtainable by the present disclosure are not limited to the above effects, and those skilled in the art can clearly understand other unmentioned effects from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a block diagram of a video encoding apparatus capable of implementing the technology of the present disclosure.

[0025] Figure 2 illustrates a method of dividing a block using a quadtree plus binary tree plus ternary tree (QTBTTT) structure.

[0026] Figure 3a and Figure 3b illustrates a plurality of intra prediction modes including a wide-angle intra prediction mode.

[0027] Figure 4 illustrates adjacent blocks of a current block.

[0028] Figure 5 is a block diagram of a video decoding apparatus capable of implementing the technology of the present disclosure.

[0029] Figure 6 is a flowchart of a process of reconstructing a coded unit according to at least one embodiment of the present disclosure.

[0030] Figure 7 is a flowchart of a process of reconstructing a coded unit by predicting and modifying a segmentation structure of a coding tree unit according to at least one embodiment of the present disclosure.

[0031] Figure 8 is a flowchart of a process of determining a predicted block of a coding tree unit according to at least one embodiment of the present disclosure.

[0032] Figure 9It is a diagram showing a method for determining a prediction block of a coding tree unit according to at least one embodiment of the present disclosure.

[0033] Figure 10 It is a diagram showing a process of modifying a segmentation structure of a coding tree unit according to at least one embodiment of the present disclosure.

[0034] Figures 11a to 11g It is a diagram showing a method for merging and splitting coding units according to at least one embodiment of the present disclosure.

[0035] Figure 12 It is a flowchart of a video decoding process according to at least one embodiment of the present disclosure.

[0036] Figure 13 It is a flowchart for showing a video encoding process according to at least one embodiment of the present disclosure. Detailed Description of Specific Embodiments

[0037] Hereinafter, some embodiments of the present disclosure will be described in detail with reference to the accompanying illustrative drawings. In the following description, although the elements are shown in different drawings, the same reference numerals denote the same elements. In addition, in the following description of some embodiments, for the purpose of clarity and conciseness, detailed descriptions of related known components and functions may be omitted when it is considered that they would obscure the subject matter of the present disclosure.

[0038] Figure 1 It is a block diagram of a video encoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 1 the illustrated diagram, a video encoding device and components of the device will be described.

[0039] The encoding device may include a picture splitter 110, a predictor 120, a subtractor 130, a transformer 140, a quantizer 145, a rearrangement unit 150, an entropy encoder 155, an inverse quantizer 160, an inverse transformer 165, an adder 170, a loop filter unit 180, and a memory 190.

[0040] Each component of the encoding device may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the function of each component may be implemented as software and may also be implemented as a microprocessor to execute the function of the software corresponding to each component.

[0041] A video is composed of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed on each region. For example, a picture is divided into one or more tiles and / or slices. Here, one or more tiles can be defined as a tile group. Each tile and / or slice is divided into one or more coding tree units (CTUs). Additionally, each CTU is divided into one or more coding units (CUs) through a tree structure. Information applied to each coding unit (CU) is encoded as the syntax of the CU, and information commonly applied to the CUs included in a CTU is encoded as the syntax of the CTU. Furthermore, information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and information applied to all blocks constituting one or more pictures is encoded into the picture parameter set (PPS) or the picture header. Moreover, information commonly referred to by multiple pictures is encoded into the sequence parameter set (SPS). Additionally, information commonly referred to by one or more SPSs is encoded into the video parameter set (VPS). Furthermore, information commonly applied to a tile or a tile group can also be encoded as the syntax of the tile or tile group header. The syntax included in the SPS, PPS, slice header, tile or tile group header can be referred to as high-level syntax.

[0042] The picture splitter 110 determines the size of the coding tree unit (CTU). Information regarding the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and is delivered to the video decoding device.

[0043] The picture splitter 110 divides each picture constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively divides the CTUs by using a tree structure. The leaf nodes in the tree structure become coding units (CUs), which are the basic units for encoding.

[0044] The tree structure can be a quadtree (QT), in which a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size. The tree structure can also be a binary tree (BT), in which a higher node is divided into two lower nodes. The tree structure can also be a ternary tree (TT), in which a higher node is divided into three lower nodes at a ratio of 1:2:1. The tree structure can also be a structure in which two or more of the QT structure, BT structure, and TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, adding the binary tree ternary tree (BTTT) to the tree structure is referred to as a multi-type tree (MTT).

[0045] Figure 2 is a diagram for describing a method of dividing blocks by using the QTBTTT structure.

[0046] As Figure 2 shown, the CTU can first be split into QT structures. The quadtree splitting can be recursive until the size of the split block reaches the minimum block size (MinQTSize) of the leaf nodes allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is split into four lower-level nodes is encoded by the entropy encoder 155 and signaled to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further split into at least one of the BT structure or the TT structure. There can be multiple splitting directions in the BT structure and / or the TT structure. For example, there can be two directions, i.e., the direction in which the block of the corresponding node is split horizontally and the direction in which the block of the corresponding node is split vertically. As Figure 2 shown, when the MTT splitting starts, a second flag (mtt_split_flag) indicating whether the node is split, and if the node is split, a flag indicating the splitting direction (vertical or horizontal) and / or a flag indicating the splitting type (binary or ternary) are encoded by the entropy encoder 155 and signaled to the video decoding device.

[0047] Alternatively, before encoding the first flag (QT_split_flag) indicating whether each node is split into four lower-level nodes, a CU splitting flag (split_cu_flag) indicating whether the node is split can also be encoded. When the value of the CU splitting flag (split_cu_flag) indicates that each node is not split, the block of the corresponding node becomes a leaf node in the split tree structure and becomes a CU that is the basic unit of encoding. When the value of the CU splitting flag (split_cu_flag) indicates that each node is split, the video encoding device first starts encoding the first flag by the above scheme.

[0048] When QTBT is used as another example of a tree structure, there can be two types, i.e., a type in which the block of the corresponding node is split horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and a type in which the block of the corresponding node is split vertically into two blocks of the same size (i.e., symmetric vertical splitting). A splitting flag (split_flag) indicating whether each node of the BT structure is split into lower-level blocks and splitting type information indicating the splitting type are encoded by the entropy encoder 155 and delivered to the video decoding device. At the same time, a type in which the block of the corresponding node is split into two blocks that are not symmetric to each other can also be presented. The asymmetric form can include a form in which the block of the corresponding node is split into two rectangular blocks with a size ratio of 1:3, or can also include a form in which the block of the corresponding node is split in a diagonal direction.

[0049] The CU can have different sizes according to the QTBT or QTBTTT partitioning from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) is referred to as the "current block". Due to the adoption of QTBTTT partitioning, the shape of the current block can be a rectangular shape in addition to the square shape.

[0050] The predictor 120 predicts the current block to generate a predicted block. The predictor 120 includes an intra predictor 122 and an inter predictor 124.

[0051] Generally, each current block in a picture can be predictively encoded. Generally, the prediction of the current block can be performed by using intra prediction techniques (using data from the picture including the current block) or inter prediction techniques (using data from pictures encoded before the picture including the current block). Inter prediction includes both uni-directional prediction and bi-directional prediction.

[0052] The intra predictor 122 uses the pixels (reference pixels) located around the current block in the current picture including the current block to predict the pixels in the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as Figure 3a shown, the multiple intra prediction modes can include two non-directional modes (including the planar mode and the DC mode), and can include 65 directional modes. The adjacent pixels and the arithmetic formula to be used are differently defined according to each prediction mode.

[0053] To perform effective directional prediction on a current block having a rectangular shape, directional modes (intra prediction modes #67 to #80, #-1 to #-14) as shown by the dashed arrows in Figure 3b can be additionally used. The directional modes can be referred to as "wide-angle intra prediction modes". In Figure 3b , the arrows indicate the corresponding reference samples for prediction without indicating the prediction direction. The prediction direction is opposite to the direction shown by the arrows. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode that performs prediction in the direction opposite to a specific directional mode without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block can be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide-angle intra prediction modes with an angle less than 45 degrees (intra prediction modes #67 to #80) are available. When the current block has a rectangular shape with a width greater than the height, wide-angle intra prediction modes with an angle greater than -135 degrees are available.

[0054] The intra predictor 122 may determine an intra prediction to be used for encoding a current block. In some examples, the intra predictor 122 may encode the current block by using multiple intra prediction modes and may also select an appropriate intra prediction mode to be used from the tested modes. For example, the intra predictor 122 may calculate rate-distortion values by using rate-distortion analysis for multiple tested intra prediction modes and may also select an intra prediction mode having the best rate-distortion characteristics from the tested modes.

[0055] The intra predictor 122 selects one intra prediction mode from multiple intra prediction modes and predicts the current block by using neighboring pixels (reference pixels) determined according to the selected intra prediction mode and arithmetic formulas. Information about the selected intra prediction mode is encoded by the entropy encoder 155 and is delivered to the video decoding device.

[0056] The inter predictor 124 generates a prediction block for the current block by using motion compensation processing. The inter predictor 124 searches for a block most similar to the current block in a reference picture that was encoded and decoded earlier than the current picture and generates a prediction block for the current block by using the searched block. Additionally, a motion vector (MV) is generated, which corresponds to the displacement between the current block in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. Motion information including information about the reference picture and information about the motion vector used for predicting the current block is encoded by the entropy encoder 155 and is delivered to the video decoding device.

[0057] The inter predictor 124 may also perform interpolation on the reference picture or reference block to improve prediction accuracy. In other words, subsamples between two consecutive integer samples are interpolated by applying filter coefficients to multiple consecutive integer samples including the two integer samples. When performing the process of searching for a block most similar to the current block for the interpolated reference picture, the motion vector may be represented not with integer sampling unit precision but with fractional unit precision. The precision or resolution of the motion vector may be set differently for each target region to be encoded (e.g., units such as slices, tiles, CTUs, CUs, etc.). When such an adaptive motion vector resolution (AMVR) is applied, information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution may be information indicating the precision of the motion vector difference to be described below.

[0058] Meanwhile, the inter-frame predictor 124 can perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference pictures and two motion vectors representing the positions of the blocks most similar to the current block in each reference picture are used. The inter-frame predictor 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1), respectively. The inter-frame predictor 124 also searches for the block most similar to the current block in the corresponding reference picture to generate a first reference block and a second reference block. Additionally, a predicted block for the current block is generated by averaging or weighted averaging the first reference block and the second reference block. Additionally, motion information including information about the two reference pictures used for predicting the current block and information about the two motion vectors is delivered to the entropy encoder 155. Here, reference picture list 0 may be composed of pictures among the pre-reconstructed pictures that are before the current picture in the display order, and reference picture list 1 may be composed of pictures among the pre-reconstructed pictures that are after the current picture in the display order. However, although not particularly limited thereto, pre-reconstructed pictures that are after the current picture in the display order may additionally be included in reference picture list 0. In contrast, pre-reconstructed pictures that are before the current picture may also be additionally included in reference picture list 1.

[0059] To minimize the amount of bits consumed for encoding motion information, various methods can be used.

[0060] For example, when the reference picture and motion vector of the current block are the same as those of a neighboring block, information identifying the neighboring block is encoded to deliver the motion information of the current block to the video decoding device. This method is called the merge mode.

[0061] In the merge mode, the inter-frame predictor 124 selects a predetermined number of merge candidates (hereinafter referred to as "merge candidates") from the neighboring blocks of the current block.

[0062] As Figure 4 shown, all or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture can be used as neighboring blocks for deriving merge candidates. In addition, blocks other than the current picture in which the current block is located within the reference picture (which may be the same as or different from the reference picture used for predicting the current block) can also be used as merge candidates. For example, a block at the same position as the current block within the reference picture or a block adjacent to the block at the same position can additionally be used as a merge candidate. If the number of merge candidates selected by the method described above is less than the preset number, a zero vector is added to the merge candidates.

[0063] The inter-frame predictor 124 configures a merge list including a predetermined number of merge candidates by using neighboring blocks. A merge candidate to be used as motion information of the current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the entropy encoder 155 and delivered to the video decoding device.

[0064] Merge skip mode is a special case of merge mode. After quantization, when all transform coefficients for entropy coding are close to zero, only neighbor block selection information is transmitted without transmitting residual signals. By using merge skip mode, relatively high coding efficiency can be achieved for images with slight motion, still images, screen content images, etc.

[0065] Hereinafter, the merge mode and the merge skip mode are collectively referred to as the merge / skip mode.

[0066] Another method for encoding motion information is the Advanced Motion Vector Prediction (AMVP) mode.

[0067] In the AMVP mode, the inter-frame predictor 124 derives a motion vector predictor candidate for a motion vector of the current block by using neighboring blocks of the current block. Figure 4 All or some of the left block A0, lower left block A1, upper block B0, upper right block B1, and upper left block B2 adjacent to the current block in the current picture shown are used as neighboring blocks for deriving motion vector prediction value candidates. In addition, blocks other than the current picture where the current block is located, which are located in a reference picture (which may be the same as or different from the reference picture used to predict the current block), may also be used as neighboring blocks for deriving motion vector prediction value candidates. For example, a block located at the same position as the current block in the reference picture or a block adjacent to a block located at the same position may be used. If the number of motion vector candidates selected by the method described above is less than the preset number, a zero vector is added to the motion vector candidates.

[0068] The inter predictor 124 derives motion vector predictor candidates by using motion vectors of neighboring blocks, and determines a motion vector predictor for a motion vector of a current block by using the motion vector predictor candidates. In addition, a motion vector difference is calculated by subtracting the motion vector predictor from the motion vector of the current block.

[0069] The motion vector prediction value can be obtained by applying a predefined function (e.g., median and average operations, etc.) to the motion vector prediction value candidates. In this case, the video decoding device also knows the predefined function. Additionally, since the neighboring blocks used to derive the motion vector prediction value candidates are blocks that have already been encoded and decoded, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information for identifying the motion vector prediction value candidates. Thus, in this case, the information about the motion vector difference and the information about the reference picture used to predict the current block are encoded.

[0070] Meanwhile, the motion vector prediction value can also be determined by a scheme of selecting any one of the motion vector prediction value candidates. In this case, the information for identifying the selected motion vector prediction value candidate is additionally encoded jointly with the information about the motion vector difference and the information about the reference picture used to predict the current block.

[0071] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra predictor 122 or the inter predictor 124 from the current block.

[0072] The transformer 140 converts the residual signal in the residual block with pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 140 can perform the transformation on the residual signal in the residual block by using the total size of the residual block as the transform unit, or can also divide the residual block into multiple sub-blocks and perform the transformation by using these sub-blocks as the transform unit. Alternatively, the residual block is divided into two sub-blocks, a transform region and a non-transform region, to perform the transformation on the residual signal by using only the transform region sub-block as the transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates that only the sub-block is transformed, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag) are encoded by the entropy encoder 155 and signaled to the video decoding device. Additionally, the size of the transform region sub-block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_quad_flag) distinguishing the corresponding segmentation is additionally encoded by the entropy encoder 155 and signaled to the video decoding device.

[0073] Meanwhile, the transformer 140 can perform transformations on the residual blocks separately in the horizontal and vertical directions. For the transformation, various types of transformation functions or transformation matrices can be used. For example, a pair of transformation functions for horizontal transformation and vertical transformation can be defined as a multiple transformation set (MTS). The transformer 140 can select a transformation function pair with the highest transformation efficiency in the MTS and can transform the residual blocks in each of the horizontal and vertical directions. Information (mts_idx) about the transformation function pair in the MTS is encoded by the entropy encoder 155 and signaled to the video decoding device.

[0074] The quantizer 145 quantizes the transform coefficients output from the transformer 140 using quantization parameters and outputs the quantized transform coefficients to the entropy encoder 155. The quantizer 145 can also immediately quantize the relevant residual blocks without transformation for any block or frame. The quantizer 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients arranged in a two-dimensional layout can be encoded and signaled to the video decoding device.

[0075] The rearrangement unit 150 can perform rearrangement of coefficient values for the quantized residual values.

[0076] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the DC coefficient to the high-frequency domain coefficients by using a zig-zag scan or a diagonal scan to output the 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, vertical scanning that scans the 2D coefficient array along the column direction and horizontal scanning that scans the 2D block type coefficients along the row direction can also be used instead of the zig-zag scan. In other words, according to the size of the transform unit and the intra prediction mode, the scanning method to be used can be determined among the zig-zag scan, diagonal scan, vertical scan, and horizontal scan.

[0077] The entropy encoder 155 generates a bitstream by encoding the sequence of 1D quantized transform coefficients output from the rearrangement unit 150 using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc.

[0078] In addition, the entropy encoder 155 encodes information related to block partitioning (such as CTU size, CTU partition flag, QT partition flag, MTT partition type, MTT partition direction, etc.) to allow the video decoding device to partition blocks equivalently to the video encoding device. In addition, the entropy encoder 155 encodes information regarding the prediction type indicating whether the current block is encoded by intra prediction or inter prediction. The entropy encoder 155 encodes intra prediction information (i.e., information regarding the intra prediction mode) or inter prediction information (in the case of the merge mode, merge index, and in the case of the AMVP mode, information regarding the reference picture index and motion vector difference) according to the prediction type. In addition, the entropy encoder 155 encodes information related to quantization (i.e., information regarding the quantization parameter and information regarding the quantization matrix).

[0079] The inverse quantizer 160 inverse quantizes the quantized transform coefficients output from the quantizer 145 to generate transform coefficients. The inverse transformer 165 transforms the transform coefficients output from the inverse quantizer 160 from the frequency domain to the spatial domain to reconstruct the residual block.

[0080] The adder 170 adds the reconstructed residual block and the prediction block generated by the predictor 120 to reconstruct the current block. When performing intra prediction on the next sequence block, the pixels in the reconstructed current block can be used as reference pixels.

[0081] The loop filter unit 180 performs filtering on the reconstructed pixels in order to reduce block effects, ringing effects, blurring effects, etc. that occur due to block-based prediction and transform / quantization. The loop filter unit 180 as a loop filter may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0082] The deblocking filter 182 filters the boundaries between the reconstructed blocks to remove block effects that occur due to block unit encoding / decoding, and the SAO filter 184 and the ALF 186 perform additional filtering on the deblocked filtered video. The SAO filter 184 and the ALF 186 are filters for compensating the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The SAO filter 184 applies an offset on a CTU basis to enhance subjective image quality and coding efficiency. On the other hand, the ALF 186 performs block unit filtering and applies different filters to compensate for distortion by distinguishing the boundaries of the corresponding blocks and the degree of variation. Information regarding the filter coefficients to be used for the ALF may be encoded and signaled to the video decoding device.

[0083] The reconstructed blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 186 are stored in the memory 190. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of blocks within a picture to be encoded later.

[0084] The video encoding device may store the bitstream of the encoded video data in a non - transitory storage medium or transmit the bitstream to the video decoding device via a communication network.

[0085] Figure 5 is a functional block diagram of a video decoding device that can implement the technology of the present disclosure. Hereinafter, with reference to Figure 5 , the video decoding device and the components of the device will be described.

[0086] The video decoding device may include an entropy decoder 510, a rearrangement unit 515, an inverse quantizer 520, an inverse transform unit 530, a predictor 540, an adder 550, a loop filter unit 560, and a memory 570.

[0087] Similar to Figure 1 the video encoding device, each component of the video decoding device may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and the microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0088] The entropy decoder 510 decodes the bitstream generated by the video encoding device to extract information related to block partitioning to determine the current block to be decoded, and extracts the prediction information and information about the residual signal required to reconstruct the current block.

[0089] The entropy decoder 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), and divides the picture into CTUs of the determined size. Additionally, the CTU is determined as the highest layer of the tree structure, i.e., the root node, and the partitioning information of the CTU can be extracted to partition the CTU using the tree structure.

[0090] For example, when partitioning the CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the partitioning of the QT is extracted to divide each node into four lower - layer nodes. Additionally, for the nodes corresponding to the leaf nodes of the QT, the second flag (mtt_split_flag) related to the partitioning of the MTT, the partitioning direction (vertical / horizontal), and / or the partitioning type (binary / ternary) are extracted to divide the corresponding leaf nodes into the MTT structure. As a result, each node below the leaf nodes of the QT is recursively divided into the BT or TT structure.

[0091] As another example, when a CTU is split by using the QTBTTT structure, a CU split flag (split_cu_flag) indicating whether the CU is split is extracted. When the corresponding block is split, a first flag (QT_split_flag) can also be extracted. During the splitting process, for each node, zero or more recursive MTT splits can occur after zero or more recursive QT splits. For example, for a CTU, the MTT split can occur immediately, or conversely, only multiple QT splits may occur.

[0092] As another example, when a CTU is split by using the QTBT structure, a first flag (QT_split_flag) related to the split of the QT is extracted to split each node into four lower-layer nodes. In addition, a split flag (split_flag) indicating whether the node corresponding to the leaf node of the QT is further split into a BT and split direction information are extracted.

[0093] Meanwhile, when the entropy decoder 510 determines the current block to be decoded by using the splitting of the tree structure, the entropy decoder 510 extracts information about the prediction type indicating whether the current block is intra prediction or inter prediction. When the prediction type information indicates intra prediction, the entropy decoder 510 extracts the syntax element for the intra prediction information (intra prediction mode) of the current block. When the prediction type information indicates inter prediction, the entropy decoder 510 extracts the information representing the syntax element for the inter prediction information (i.e., the motion vector and the reference picture referred to by the motion vector).

[0094] In addition, the entropy decoder 510 extracts quantization-related information and extracts information about the quantized transform coefficients of the current block as information about the residual signal.

[0095] The rearrangement unit 515 can change the sequence of the 1D quantized transform coefficients entropy decoded by the entropy decoder 510 back into a 2D coefficient array (i.e., a block) in an order opposite to the coefficient scan order performed by the video coding device.

[0096] The inverse quantizer 520 inverse quantizes the quantized transform coefficients and inverse quantizes the quantized transform coefficients by using the quantization parameter. The inverse quantizer 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in 2D. The inverse quantizer 520 can perform inverse quantization by applying a matrix of quantization coefficients (scaling values) from the video coding device to the 2D array of quantized transform coefficients.

[0097] The inverse transformer 530 generates a residual block for the current block by reconstructing the residual signal by inverse-transforming the inverse quantized transform coefficients from the frequency domain to the spatial domain.

[0098] In addition, when the inverse transformer 530 performs inverse transformation on a partial area (sub-block) of the transform block, the inverse transformer 530 extracts the flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or the position information (cu_sbt_pos_flag) of the sub-block. The inverse transformer 530 also inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to reconstruct the residual signal, and fills the un-inverse-transformed area with the value "0" as the residual signal to generate the final residual block for the current block.

[0099] In addition, when MTS is applied, the inverse transformer 530 determines the transform function or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device. The inverse transformer 530 also performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.

[0100] The predictor 540 may include an intra predictor 542 and an inter predictor 544. The intra predictor 542 is activated when the prediction type of the current block is intra prediction, and the inter predictor 544 is activated when the prediction type of the current block is inter prediction.

[0101] The intra predictor 542 determines the intra prediction mode of the current block among multiple intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoder 510. The intra predictor 542 also predicts the current block by using the neighboring reference pixels of the current block according to the intra prediction mode.

[0102] The inter predictor 544 determines the motion vector of the current block and the reference picture referred to by the motion vector by using the syntax element for the inter prediction mode extracted from the entropy decoder 510, and predicts the current block by using the motion vector and the reference picture.

[0103] The adder 550 reconstructs the current block by adding the residual block output from the inverse transformer 530 and the prediction block output from the inter predictor 544 or the intra predictor 542. The pixels in the reconstructed current block are used as reference pixels for intra prediction of the blocks to be decoded later.

[0104] The loop filter unit 560 serving as a loop filter may include a deblocking filter 562, an SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the reconstructed blocks to remove the blocking artifacts that occur due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels that occurs due to lossy coding. The filter coefficients of the ALF are determined by using the information on the filter coefficients decoded from the bitstream.

[0105] The reconstructed blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in a picture are reconstructed, the reconstructed picture can be used as a reference picture for inter prediction of the blocks within the pictures to be encoded later.

[0106] The following embodiments may be performed by the intra prediction unit 122, the transform unit 140, and the inverse transform unit 165 in a video coding device. They may also be performed by the inverse transform unit 530 and the intra prediction unit 542 in a video decoding device. Hereinafter, the video decoding device is described as performing the corresponding processing steps, but the video coding device may also perform the same steps.

[0107] Figure 6 is a flowchart of a process of a reconstructed coding unit according to at least one embodiment of the present disclosure. Averagely dividing a picture may generate coding tree units (CTUs). For example, a coding tree unit may be a 128×128 area in the upper left corner of a picture. A coding tree unit may be divided into one or more coding units (CUs), and the sample values may be reconstructed by the coding units. Each time a coding unit is determined, the sample values of the coding unit may be reconstructed. Alternatively, after all the coding units in the coding tree unit are determined, the sample values of the coding units may be reconstructed while scanning the coding units in a scan order.

[0108] Reference Figure 6, the video decoding device determines whether the current coding tree unit exists within an intra slice (S610). In an intra slice, intra prediction can be performed on all coding units within the slice. When it is determined that the current coding tree unit exists within the intra slice (S610 - Yes), the video decoding device analyzes the information regarding the segmentation structure of the coding tree unit (S620). The segmentation structure of the coding tree unit is the structure of the coding units that make up the coding tree unit. The information regarding the segmentation structure of the coding tree unit is used to determine the segmentation structure of the coding tree unit. The information regarding the segmentation structure of the coding tree unit may include information regarding the segmentation or non - segmentation of the determined coding tree unit, information regarding the segmentation method of the coding tree unit, and information regarding the segmentation direction of the coding tree unit. The video decoding device uses the determined segmentation structure of the coding tree unit to reconstruct the coding unit (S630). The information regarding the parsed segmentation structure of the coding tree unit can be used to determine the segmentation structure of the coding tree unit.

[0109] When it is determined that the current coding tree unit does not exist within the intra slice (S610 - No), the video decoding device analyzes the information regarding the segmentation structure of the coding tree unit or predicts the segmentation structure of the coding tree unit (S640). If the current coding tree unit does not exist within the intra slice, then the current coding tree unit may exist within an inter slice. In an inter slice, intra prediction or inter prediction can be performed on the coding units within the slice. If the segmentation structure of the coding tree unit is to be predicted, the video decoding device can use reference information and a convolutional neural network to predict the segmentation structure of the coding tree unit. The video decoding device uses the determined segmentation structure of the coding tree unit to reconstruct the coding unit (S650). The information regarding the parsed segmentation structure of the coding tree unit can be used to determine the segmentation structure of the coding tree unit. Alternatively, the predicted segmentation structure of the coding tree unit can be used to determine the segmentation structure of the coding tree unit.

[0110] Figure 7 is a flowchart of a process for reconstructing a coding unit by predicting and modifying the segmentation structure of a coding tree unit according to at least one embodiment of the present disclosure.

[0111] Reference Figure 7, the video decoding device constructs reference information (S710). The reference information may include information about reconstructed samples of coding tree units adjacent to the top of the current coding tree unit and reconstructed samples of coding tree units adjacent to the left of the current coding tree unit. The reference information may include information obtained by two-dimensionally packing information about motion vectors, quantization parameters, and prediction modes of coding tree units adjacent to the top of the current coding tree unit. The reference information may include information obtained by two-dimensionally packing information about motion vectors, quantization parameters, and prediction modes of coding tree units adjacent to the left of the current coding tree unit. The reference information may include information about regions spatially identical to the current coding tree unit in a reference picture and adjacent regions of such spatially identical regions, etc. Information about regions spatially identical to the current coding tree unit in a reference picture and adjacent regions of such spatially identical regions may include information about temporal parameters or information about segmentation parameters. The reference information may include information about the prediction block of the current coding tree unit.

[0112] The video decoding device predicts the segmentation structure of a coding tree unit (S720). The reference information and a convolutional neural network may be used to predict the segmentation structure of the coding tree unit. The convolutional neural network is an artificial neural network for analyzing videos. The convolutional neural network can extract features from pictures and process them. The video decoding device modifies the predicted segmentation structure of the coding tree unit (S730). The predicted segmentation structure of the coding tree unit may be modified by scanning each coding unit from the predicted segmentation structure of the coding tree unit and merging or splitting each coding unit. Merge flags and split flags may be used to merge or split coding units. Modifying the predicted segmentation structure of the coding tree unit determines the segmentation structure of the final coding tree unit. The video decoding device reconstructs coding units (S740). Each coding unit is reconstructed from the segmentation structure of the final coding tree unit, and the sample values of each coding unit are reconstructed.

[0113] Figure 8 is a flowchart of a process for determining a prediction block of a coding tree unit according to at least one embodiment.

[0114] Reference Figure 8, the reference information may include information about the prediction block of the current coding tree unit. The video decoding device determines the motion information of the current coding tree unit (S810). The motion information of the current coding tree unit can be determined by referring to the motion information of the coding tree unit adjacent to the top of the current coding tree unit and the motion information of the coding tree unit adjacent to the left of the current coding tree unit. The motion information may include information such as the index of the reference picture and the motion vector. The motion information of the coding units in the coding tree units adjacent to the top of the current coding tree unit and the motion information of the coding units in the coding tree units adjacent to the left of the current coding tree unit can be listed. Among the listed motion information items, the motion information of the current coding tree unit can be determined.

[0115] The motion information determined to be that of the current coding tree unit may be the motion information of the coding unit closest in distance to the upper left sample of the current coding tree unit. When there are multiple coding units closest in distance to the upper left sample of the current coding tree unit, the motion information of the coding unit with the largest size can be determined as the motion information of the current coding tree unit. The motion information of the coding unit adjacent to the top of the current coding tree unit can be determined as the motion information of the current coding tree unit. The motion information of the coding unit adjacent to the left of the current coding tree unit can be determined as the motion information of the current coding tree unit.

[0116] The video decoding device uses the motion information of the current coding tree unit to generate a prediction block for the current coding tree unit (S820). The prediction block can be generated to have the same size as the current coding tree unit. Each motion information item can generate one prediction block. Therefore, when multiple motion information items of the current coding tree unit are determined, multiple prediction blocks of the current coding tree unit can be generated. The video decoding device performs a weighted sum of the prediction blocks of the current coding tree unit (S830). Performing a weighted sum of the prediction blocks of the current coding tree unit can generate the final prediction block of the current coding tree unit. The weights assigned to the prediction blocks of the current coding tree unit can all be the same. The weights assigned to the prediction blocks of the current coding tree unit can be any value.

[0117] Figure 9 is a diagram showing a method for determining a prediction block of a coding tree unit according to at least one embodiment of the present disclosure.

[0118] Reference Figure 9 , the video decoding device may use the motion vectors of the coding tree unit adjacent to the top of the current coding tree unit and the motion vectors of the coding tree unit adjacent to the left of the current coding tree unit to determine the motion vector of the current coding tree unit in the current picture. The video decoding device may use the determined motion vector of the current coding tree unit to generate a prediction block for the current coding tree unit.

[0119] The video decoding device may determine an average motion vector by averaging motion vectors in each reference picture based on a reference picture index. In this process, motion vectors in each reference picture that are significantly different from the median motion vector may be excluded. The video decoding device may use the reference picture corresponding to the reference picture index with the smallest index difference from the current picture to determine the motion information of the current coded tree block. The video decoding device may determine the average motion vector obtained by averaging the motion vectors in the reference picture as the motion vector of the current coded tree block.

[0120] The video decoding device may determine a first motion vector by averaging the motion vectors of a first reference picture, and determine the index of the first reference picture as the motion information of the current coded tree block. The video decoding device may use the first motion vector and the index of the first reference picture to generate a first prediction block. The video decoding device may determine a second motion vector by averaging the motion vectors of a second reference picture, and determine the index of the second reference picture as the motion information of the current coded tree block. The video decoding device may use the second motion vector and the index of the second reference picture to generate a second prediction block. The video decoding device may perform a weighted average on the first prediction block and the second prediction block to generate a final prediction block. The weights assigned to the first prediction block and the second prediction block may be the same. The weights assigned to the first prediction block and the second prediction block may be any values. The weights assigned to the first prediction block and the second prediction block may be inversely proportional to the temporal distance between the current picture and the first reference picture and the temporal distance between the current picture and the second reference picture. Information about the final prediction block is included in the reference information.

[0121] Figure 10 It is a diagram showing a process of modifying a segmentation structure of a coded tree unit according to at least one embodiment of the present disclosure.

[0122] Reference Figure 10 , the video decoding device uses the predicted segmentation structure of the coded tree unit to sequentially search for coded units (S1010). The coded units to be merged or segmented may be searched in a zigzag scan order. When a coded unit is merged or segmented so as to modify the segmentation structure of the modified coded tree unit, the index of each coded unit is also modified. In this case, the coded unit may be searched based on the modified index. If the coded unit is no longer segmented or merged, the coded unit with the next index may be searched.

[0123] The video decoding device uses a merge flag to merge coding units (S1020). If the merge flag (e.g., merge_flag) has a first value (e.g., 0), the coding units may not be merged. If the merge flag (e.g., merge_flag) is a second value (e.g., 1), the coding units may be merged. When merging coding units, the current coding unit may be merged with all coding units in the parent unit of the current coding unit. If any one of the predetermined conditions is satisfied, the merge_flag may be deduced as the first value (e.g., 0). The first condition among the predetermined conditions is the condition of whether there are split coding units in the lower-level coding units of the parent unit of the current coding unit. The second condition is the condition of whether the coding units having the same depth as the current coding unit in the lower-level coding units of the parent unit of the current coding unit are not merged. The third condition is the condition of whether the current coding unit is generated by splitting.

[0124] The video decoding device uses a split flag to split coding units (S1030). When the split flag (e.g., split_flag) has a first value (e.g., 0), the coding units may not be split. If the split flag (e.g., split_flag) is a second value (e.g., 1), the coding units may be split. When the coding units are split, the current coding unit may be split into two or more coding units. If the split_flag is a second value (e.g., 1), the information about the splitting method and the information about the splitting direction may be parsed. The information about the splitting method may include the information about quadtree splitting, binary tree splitting, and ternary tree splitting. The information about the splitting direction may include the information about the vertical direction and the horizontal direction, etc. If any one of the predetermined conditions is satisfied, the split_flag may be deduced as the first value (e.g., 0). The first condition among the predetermined conditions is the condition of whether the merge_flag of the current coding unit is a second value (e.g., 1). The second condition is the condition of whether the width, height, and area of the current coding unit are equal to the lower threshold. The video decoding device may first determine whether to merge or not merge the current coding unit, and then determine whether to split or not split the current coding unit. Alternatively, the video decoding device may preferentially determine whether to split or not split the current coding unit, and then determine whether to merge or not merge the current coding unit.

[0125] Figures 11a to 11g is a diagram showing a method of merging and splitting coding units according to at least one embodiment of the present disclosure.

[0126] Reference Figure 11a, the video decoding device determines the splitting structure of the predicted coding tree unit. In the splitting structure of the predicted coding tree unit, three coding units in the upper left parent unit can be assigned indexes 1, 2, and 3 in order from left to right. The coding unit adjacent to the right side of the upper left parent unit can be assigned index 4. The coding unit adjacent to the lower left in the parent unit can be assigned index 5. Similarly, indexes can be assigned to the remaining coding units.

[0127] Reference Figure 11b , the video decoding device can obtain the merge_flag for the coding unit assigned index 1. Since the merge_flag is the second value (e.g., 1), the video decoding device can merge the coding unit assigned index 1, the coding unit assigned index 2, and the coding unit assigned index 3 within the parent unit. The merged coding unit can be assigned index 1. Since the merge_flag has the second value (e.g., 1), the split_flag can be deduced as the first value (e.g., 0). Since the split_flag is deduced as the first value (e.g., 0), the video decoding device can not split the coding unit initially assigned index 1.

[0128] Reference Figure 11c , the video decoding device can obtain the merge_flag of the merged coding unit assigned index 1 in Figure 11b . Since the merge_flag is the first value (e.g., 0), merging can not be performed. The video decoding device can obtain the split_flag of the merged coding unit assigned index 1. Since the split_flag is the second value (e.g., 1), and the information about the splitting method indicates a binary tree split in the horizontal direction, the merged coding unit assigned index 1 can be split into two coding units in the horizontal direction. Among these two coding units, the top coding unit can be assigned index 1, and the bottom coding unit can be assigned index 2. The coding unit adjacent to the right side of these two coding units can be assigned index 3.

[0129] Reference Figure 11d , since the top coding unit assigned index 1 in Figure 11c is a coding unit generated by splitting, the merge_flag can be deduced as the first value (e.g., 0). Since the merge_flag is the first value (e.g., 0), merging can not be performed. The video decoding device can obtain the split_flag of the top coding unit assigned index 1. Since the split_flag is the first value (e.g., 0), splitting can not be performed.

[0130] Reference Figure 11e , since the assignedFigure 11c The bottom coding unit with index 2 in is a coding unit generated by splitting, so the merge_flag can be deduced to be the first value (e.g., 0). Since the merge_flag is the first value (e.g., 0), merging may not be performed. The video decoding device can obtain the split_flag of the bottom coding unit assigned with index 2. Since the split_flag is the first value (e.g., 0), splitting may not be performed.

[0131] Reference Figure 11f , since there is a split coding unit in the lower-level coding units of the parent unit of the coding unit with index 3 in Figure 11c , the merge_flag can be deduced to be the first value (e.g., 0). In the lower-level coding units of the parent unit of the coding unit assigned with index 3, the split coding units are the top coding unit assigned with index 1 and the bottom coding unit assigned with index 2. Since the merge_flag is the first value (e.g., 0), merging may not be performed. The video decoding device can obtain the split_flag of the coding unit assigned with index 3. Since the split_flag is the second value (e.g., 1) and the information about the splitting method indicates a ternary tree split in the vertical direction, the coding unit assigned with index 3 can be split into three coding units in the vertical direction. Among these three coding units, the left coding unit can be assigned index 3, the middle coding unit can be assigned index 4, and the right coding unit can be assigned index 5.

[0132] Reference Figure 11g , the left coding unit assigned with Figure 11f index 3 in is a coding unit generated by splitting, so the merge_flag can be deduced to be the first value (e.g., 0). Since the merge_flag is the first value (e.g., 0), merging may not be performed. The video decoding device can obtain the split_flag of the left coding unit assigned with index 3. Since the split_flag is the first value (e.g., 0), splitting may not be performed.

[0133] Figure 12 is a flowchart of a video decoding process according to at least one embodiment of the present disclosure.

[0134] Reference Figure 12 , the video decoding device predicts the splitting structure of the current coding tree unit (S1210). The process of predicting the splitting structure of the current coding tree unit may include obtaining reference information and using the reference information and a neural network to predict the splitting structure of the current coding tree unit.

[0135] The reference information may include at least one of information on reconstructed samples of coding tree units adjacent to the top of the current coding tree unit and reconstructed samples of coding tree units adjacent to the left of the current coding tree unit, information obtained by two-dimensionally packing information on motion vectors, quantization parameters, and prediction modes of coding tree units adjacent to the top and left of the current coding tree unit, information on regions spatially identical to the current coding tree unit in a reference picture and adjacent regions of the spatially identical region, and information on a prediction block of the current coding tree unit.

[0136] The video decoding apparatus determines whether to merge or not merge the current coding unit within the current coding tree unit using the segmentation structure of the current coding tree unit, and merges the current coding unit or omits the merging of the current coding unit based on the determined merge or non-merge (S1220). The process of merging the current coding unit or the process of omitting the merging of the current coding unit may include omitting the merging of the current coding unit based on satisfying at least one predetermined condition. The predetermined condition may include a condition of whether there is a segmented coding unit among the subordinate coding units of the parent unit of the current coding unit, a condition of whether the coding units having the same depth as the current coding unit among the subordinate coding units of the parent unit of the current coding unit are not merged, and a condition of whether the current coding unit is generated by segmentation.

[0137] The video decoding apparatus determines whether to segment or not segment the current coding unit based on the determined merge or non-merge of the current coding unit, and determines the segmentation structure of the current coding tree unit by segmenting the current coding unit based on the determined segmentation or non-segmentation or by omitting the segmentation of the current coding unit (S1230). The process of determining the segmentation structure of the current coding tree unit may include determining the segmentation structure of the current coding tree unit by omitting the segmentation of the current coding unit based on satisfying at least one predetermined condition. The predetermined condition may include a condition of whether the current coding unit is merged and a condition of whether the width, height, and area of the current coding unit are equal to predetermined values. The video decoding apparatus uses the segmentation structure of the current coding tree unit to generate a prediction block of the current coding unit (S1240).

[0138] Figure 13 is a flowchart illustrating a video coding process according to at least one embodiment of the present disclosure.

[0139] The video coding apparatus predicts the segmentation structure of the current coding tree unit (S1310). The process of predicting the segmentation structure of the current coding tree unit may include obtaining reference information and using the reference information and a neural network to predict the segmentation structure of the current coding tree unit.

[0140] The reference information may include at least one of information on reconstructed samples of a coding tree unit adjacent to the top of the current coding tree unit and reconstructed samples of a coding tree unit adjacent to the left of the current coding tree unit, information obtained by two-dimensionally packing information on motion vectors, quantization parameters, and prediction modes of coding tree units adjacent to the top and left of the current coding tree unit, information on a region spatially identical to the current coding tree unit in a reference picture and an adjacent region of the spatially identical region, and information on a prediction block of the current coding tree unit.

[0141] The video coding device may utilize the partitioning structure of the current coding tree unit to determine whether to merge or not merge the current coding unit within the current coding tree unit, and may merge the current coding unit or omit the merging of the current coding unit based on the determined merge or non-merge (S1320). The process of merging the current coding unit or omitting its merging may include omitting the merging of the current coding unit based on satisfying at least one predetermined condition. The predetermined condition may include a condition of whether there is a partitioned coding unit among the lower-level coding units of the parent unit of the current coding unit, a condition of whether the coding units having the same depth as the current coding unit among the lower-level coding units of the parent unit of the current coding unit are not merged, and a condition of whether the current coding unit is generated by partitioning.

[0142] The video coding device determines whether to partition or not partition the current coding unit based on the determined merge or non-merge of the current coding unit, and determines the partitioning structure of the current coding tree unit by partitioning the current coding unit based on the determined partition or non-partition or by omitting the partitioning of the current coding unit (S1330). The process of determining the partitioning structure of the current coding tree unit may include determining the partitioning structure of the current coding tree unit by omitting the partitioning of the current coding unit based on satisfying at least one predetermined condition. The predetermined condition may include a condition of whether the current coding unit is merged and a condition of whether the width, height, and area of the current coding unit are equal to predetermined values. The video coding device uses the partitioning structure of the current coding tree unit to generate a prediction block of the current coding unit (S1340).

[0143] Although the steps in the corresponding flowchart are described as being executed sequentially, these steps merely illustrate the technical concept of some embodiments of the present disclosure. Therefore, those of ordinary skill in the art to which the present disclosure pertains may execute the steps by changing the order described in the corresponding drawings or by executing two or more steps in parallel. Therefore, the steps in the corresponding flowchart are not limited to the shown time-series order.

[0144] It should be understood that the above description presents illustrative embodiments that can be implemented in various other ways. The functions described in some embodiments can be implemented by hardware, software, firmware, and / or combinations thereof. It should also be understood that the functional components described in this disclosure are marked with “… unit” to emphasize the possibility of their independent implementation.

[0145] Meanwhile, the various methods or functions described in some embodiments can be implemented as instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. The non-transitory recording medium can include various types of recording devices that store data in a form readable by a computer system. For example, the non-transitory recording medium can include storage media such as erasable programmable read-only memory (EPROM), flash drives, optical disk drives, magnetic hard disk drives, and solid-state drives (SSD), etc.

[0146] Although the embodiments of the present disclosure have been described for illustrative purposes, those of ordinary skill in the art to which the present disclosure pertains should understand that various modifications, additions, and substitutions are possible without departing from the concept and scope of the present disclosure. Therefore, the embodiments of the present disclosure have been described for the sake of brevity and clarity. The scope of the technical concept of the embodiments of the present disclosure is not limited by the illustrations. Therefore, those of ordinary skill in the art to which the present disclosure pertains should understand that the scope of the present disclosure should not be limited by the embodiments described above in detail, but rather by the claims and their equivalents.

[0147] Cross-reference to related applications

[0148] This application claims the priority of Korean Patent Application No. 10-2022-0160938, filed on November 25, 2022, and Korean Patent Application No. 10-2023-0159810, filed on November 17, 2023. The entire disclosure of each patent is incorporated herein by reference in its entirety.

Claims

1. A video decoding method, comprising: Predicting a segmentation structure of a current coding tree unit; Using the segmentation structure of the current coding tree unit to determine whether to merge or not merge a current coding unit within the current coding tree unit, and merging the current coding unit or omitting the merging of the current coding unit based on the determined merging or non-merging of the current coding unit; Based on the determined merging or non-merging of the current coding unit, determining whether to split or not split the current coding unit, and determining the segmentation structure of the current coding tree unit by splitting the current coding unit or omitting the splitting of the current coding unit based on the determined splitting or non-splitting; And Using the segmentation structure of the current coding tree unit to generate a prediction block of the current coding unit.

2. The video decoding method according to claim 1, wherein, Predicting the segmentation structure of the current coding tree unit includes: Obtaining reference information; and Using the reference information and a neural network to predict the segmentation structure of the current coding tree unit.

3. The video decoding method according to claim 2, wherein, The reference information includes at least one of the following items: Information about reconstructed samples of a coding tree unit adjacent to the top of the current coding tree unit and reconstructed samples of a coding tree unit adjacent to the left of the current coding tree unit, Information obtained by two-dimensionally packing information about motion vectors, quantization parameters, and prediction modes of coding tree units adjacent to the top and the left of the current coding tree unit, Information about a region spatially identical to the current coding tree unit in a reference picture and adjacent regions of the spatially identical region, and Information about a prediction block of the current coding tree unit.

4. The video decoding method according to claim 1, wherein, Merging the current coding unit or omitting the merging of the current coding unit includes: Based on meeting at least one predetermined condition, omitting the merging of the current coding unit.

5. The video decoding method according to claim 4, wherein, The predetermined condition includes: A condition of whether there is a split coding unit among the lower-level coding units of the parent unit of the current coding unit, A condition of whether a coding unit having the same depth as the current coding unit among the lower-level coding units of the parent unit of the current coding unit is not merged, and A condition of whether the current coding unit is generated by splitting.

6. The video decoding method according to claim 1, wherein, Determining the segmentation structure of the current coding tree unit includes: Determining the segmentation structure of the current coding tree unit by omitting the splitting of the current coding unit based on meeting at least one predetermined condition.

7. The video decoding method according to claim 6, wherein, The predetermined condition includes: A condition of whether the current coding unit is merged, and A condition of whether the width, height, and area of the current coding unit are equal to predetermined values.

8. A video coding method, comprising: Predicting a segmentation structure of a current coding tree unit; Using the segmentation structure of the current coding tree unit to determine whether to merge or not merge a current coding unit within the current coding tree unit, and merging the current coding unit or omitting the merging of the current coding unit based on the determined merging or non-merging of the current coding unit; Determine whether to split or not split the current coding unit based on the determined merge or non-merge of the current coding unit, and determine the split structure of the current coding tree unit by splitting the current coding unit or omitting the split of the current coding unit based on the determined split or non-split; And Generate a prediction block of the current coding unit using the split structure of the current coding tree unit.

9. The video encoding method according to claim 8, wherein, Predicting the split structure of the current coding tree unit includes: Determine reference information; and Use the reference information and a neural network to predict the split structure of the current coding tree unit.

10. The video encoding method according to claim 9, wherein, The reference information includes at least one of the following items: Information about the reconstructed samples of the coding tree unit adjacent to the top of the current coding tree unit and the reconstructed samples of the coding tree unit adjacent to the left of the current coding tree unit, Information obtained by two-dimensionally packing information about the motion vectors, quantization parameters, and prediction modes of the coding tree units adjacent to the top and the left of the current coding tree unit, Information about the region spatially identical to the current coding tree unit in the reference picture and the adjacent regions of the spatially identical region, and Information about the prediction block of the current coding tree unit.

11. The video encoding method according to claim 8, wherein, Merging the current coding unit or omitting the merge of the current coding unit includes: Based on meeting at least one predetermined condition, omit the merge of the current coding unit.

12. The video encoding method according to claim 11, wherein, The predetermined conditions include: The condition of whether there is a split coding unit among the lower-level coding units of the parent unit of the current coding unit, The condition of whether the coding units having the same depth as the current coding unit among the lower-level coding units of the parent unit of the current coding unit are not merged, and The condition of whether the current coding unit is generated by splitting.

13. The video encoding method according to claim 8, wherein, Determining the split structure of the current coding tree unit includes: Determine the split structure of the current coding tree unit by omitting the split of the current coding unit based on meeting at least one predetermined condition.

14. The video encoding method according to claim 13, wherein, The predetermined conditions include: The condition of whether the current coding unit is merged, and The condition of whether the width, height, and area of the current coding unit are equal to predetermined values.

15. A computer-readable recording medium storing a bitstream generated by a video coding method, the video coding method including: Predict the split structure of the current coding tree unit; Use the split structure of the current coding tree unit to determine whether to merge or not merge the current coding unit within the current coding tree unit, and merge the current coding unit or omit the merge of the current coding unit based on the determined merge or non-merge of the current coding unit; Determine whether to split or not split the current coding unit based on the determined merge or non-merge of the current coding unit, and determine the split structure of the current coding tree unit by splitting the current coding unit or omitting the split of the current coding unit based on the determined split or non-split; And Generate a prediction block of the current coding unit using the split structure of the current coding tree unit.

Citation Information

Patent Citations

  • Biosensor comprising polymer-dispersed reduced graphene oxide, Diagnosis sensor for prostate cancer comprising the same and manufacturing method thereof

    KR1020220160938A

  • Resin producing method

    KR1020230159810A