Method and non-transitory computer-readable medium for encoding or decoding video
By introducing a new structure into video encoding technology to encode and decode the encoding unit syntax, the problem of complexity of block division structure and syntax structure in the prior art is solved, and a more efficient encoding and decoding process is realized.
Patent Information
- Application Number
- CN202210852847.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-11
- Filing Date
- 2019-01-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-01-04
AI Technical Summary
In the existing video encoding technology, the complexity of the block division structure and syntax structure has led to the increased difficulty of encoding and decoding the encoding unit syntax, especially in HEVC, the relevant syntax of CU, PU and TU must be encoded and decoded respectively.
A new structure is proposed for encoding or decoding the syntax of the encoding unit, determining the encoding unit through block division, and decoding its prediction syntax and transformation syntax, including skip flags and encoding unit cbf, and reconstructing the encoding unit based on the prediction syntax and transformation syntax.
The encoding and decoding process of encoding unit syntax is simplified, encoding efficiency is improved, and dependence on complex block division structures and syntax structures is reduced.
Smart Images

Figure CN115037930B_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application with the original application number 201980017079.X (International Application No.: PCT / KR2019 / 000136, filing date: January 4, 2019, invention title: Method and apparatus for encoding or decoding an image). Technical Field
[0002] The present disclosure relates to the encoding and decoding of video. In one aspect, the present disclosure relates to the encoding or decoding syntax for an encoding unit that is a basic unit of encoding. Background Art
[0003] Since the amount of video data is larger than the amount of voice data or still image data, storing or transmitting video data without compression requires a large amount of hardware resources including a memory. Therefore, when storing or transmitting video data, an encoder is used to compress the video data for storage or transmission. Then, a decoder receives the compressed video data and decompresses and reproduces the video data. Compression techniques for such video include H.264 / AVC and High Efficiency Video Coding (HEVC), and HEVC was established in early 2013 and has about 40% higher encoding efficiency than H.264 / AVC.
[0004] Figure 1 FIG. illustrates block partitioning in HEVC.
[0005] In HEVC, a picture is divided into a plurality of Coding Tree Units (CTUs) having a square shape, and each CTU is recursively divided into a plurality of Coding Units (CUs) having a square shape according to a quadtree structure. When determining a CU as a basic unit of encoding, the CU is divided into one or more Prediction Units (PUs) and prediction is performed on a per-PU basis. The CU is divided into PUs by selecting one of a plurality of partitioning types having good encoding efficiency. The encoding device encodes the prediction syntax for each PU so that the decoding device can predict each PU in the same manner as the encoding device.
[0006] In addition, the CU is divided into one or more Transform Units (TUs) according to a quadtree structure, and the residual signal, which is the difference between actual pixels and predicted pixels, is transformed using the size of the TU. The syntax for the transform is encoded on a per-TU basis and is sent to the decoding device.
[0007] As described above, HEVC has a complex block partitioning structure such as dividing a CTU into CUs, dividing a CU into PUs, and dividing a CU into TUs. Therefore, CUs, PUs, and TUs can be blocks of different sizes. In this block partitioning structure of HEVC, the relevant syntaxes of CUs and PUs and TUs in the CUs must be encoded separately. In HEVC, the syntax of CUs is initially encoded. Then, each CU calls each PU to encode the syntax for each PU, and also calls each TU to encode the syntax for each TU.
[0008] To cope with the complexity of the block partitioning structure and the syntax structure, a block partitioning technique of dividing a CTU into CUs and then using each CU as a PU and a TU has been recently discussed. In this new partitioning structure, when a CU is determined, prediction and transformation are performed in the size of the CU without additional partitioning. That is, CUs, PUs, and TUs are the same blocks. The introduction of the new partitioning structure requires a new structure for encoding the syntax for CUs. Summary of the Invention
[0009] Technical Problem
[0010] To meet this requirement, one aspect of the present invention proposes a new structure for encoding or decoding the syntax of coding units.
[0011] Technical Solution
[0012] According to one aspect of the present disclosure, there is provided a method for decoding a video, the method including the steps of: determining, by block partitioning, a coding unit to be decoded; decoding a prediction syntax of the coding unit, the prediction syntax including a skip flag indicating whether the coding unit is encoded in a skip mode; after decoding the prediction syntax, decoding a transform syntax including a transform / quantization skip flag and a coding unit cbf, where the transform / quantization skip flag indicates whether at least part of inverse transformation, inverse quantization, and in-loop filtering is skipped, and the coding unit cbf indicates whether all coefficients in a luminance block and two chrominance blocks constituting the coding unit are all zero, and reconstructing the coding unit based on the prediction syntax and the transform syntax.
[0013] According to another aspect of the present disclosure, there is provided a video decoding device, which includes a decoder and a reconstructor. The decoder is configured to: determine an encoding unit to be decoded through block partitioning; decode a prediction syntax of the encoding unit, where the prediction syntax includes a skip flag indicating whether the encoding unit is encoded in a skip mode; and after decoding the prediction syntax, decode a transform syntax of the encoding unit including a transform / quantization skip flag and an encoding unit cbf, where the transform / quantization skip flag indicates whether at least part of inverse transform, inverse quantization, and in-loop filtering is skipped, and the encoding unit cbf indicates whether all coefficients in the luminance block and two chrominance blocks constituting the encoding unit are all zero; the reconstructor is configured to reconstruct the encoding unit based on the prediction syntax and the transform syntax. Description of the Drawings
[0014] Figure 1 FIG. is an illustration of block partitioning in HEVC.
[0015] Figure 2 FIG. is an exemplary block diagram of a video encoding device capable of implementing the technology of the present disclosure.
[0016] Figure 3 FIG. is an example diagram of block partitioning using a QTBT structure.
[0017] Figure 4 FIG. is an example diagram of various intra prediction modes.
[0018] Figure 5 FIG. is an example diagram of peripheral blocks of a current CU.
[0019] Figure 6 FIG. is an exemplary block diagram of a video decoding device capable of implementing the technology of the present disclosure.
[0020] Figure 7 FIG. is an exemplary flowchart for decoding CU syntax according to the present disclosure.
[0021] Figure 8 FIG. is another exemplary flowchart for decoding CU syntax according to the present disclosure.
[0022] Figure 9 FIG. is an exemplary flowchart for illustrating the process of decoding a transform syntax for corresponding luminance and chrominance components.
[0023] Figure 10 FIG. is an exemplary flowchart for illustrating a method of decoding cbf for three components constituting a CU. Detailed Description of the Invention
[0024] In the following, some embodiments of the present invention will be described in detail with reference to the accompanying drawings. It should be noted that when adding reference numerals to the constituent elements in the corresponding drawings, the same reference numerals designate the same elements, although these elements are shown in different drawings. In addition, in the following description of the present invention, when a detailed description of known functions and configurations incorporated herein would make the subject matter of the present invention rather unclear, the detailed description will be omitted.
[0025] Figure 2 is an exemplary block diagram of a video encoding device capable of implementing the technology of the present disclosure.
[0026] The video encoding device includes a block divider 210, a predictor 220, a subtractor 230, a transformer 240, a quantizer 245, an encoder 250, an inverse quantizer 260, an inverse transformer 265, an adder 270, an in-loop filter 280, and a memory 290. Each element of the video encoding device can be implemented as a hardware chip, or can be implemented as software, and one or more microprocessors can be implemented to execute the functions of the software corresponding to the respective elements.
[0027] A video is composed of a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed for each region. For example, a picture is divided into one or more slices and / or tiles, and each slice or tile is divided into one or more coding tree units (CTUs). In addition, each CTU is divided into one or more coding units (CUs) through a tree structure. The information applied to each CU is encoded as the syntax of the CU, and the information commonly applied to the CUs included in one CTU is encoded as the syntax of the CTU. The information commonly applied to all blocks in one slice is encoded as the syntax of the slice, and the information applied to all blocks constituting one picture is encoded in the picture parameter set (PPS). In addition, the information commonly referred to by a plurality of pictures is encoded in the sequence parameter set (SPS). In addition, the information commonly referred to by one or more SPSs is encoded in the video parameter set (VPS).
[0028] The block divider 210 determines the size of a coding tree unit (CTU). Information about the CTU size (CTU size) is encoded as the syntax of the SPS or PPS and is sent to the video decoding device. The block divider 210 divides each picture constituting the video into a plurality of CTUs of a predetermined size, and then recursively divides the CTUs using a tree structure. A leaf node in the tree structure is used as a coding unit (CU), and the CU is a basic unit of coding. The tree structure may be a quadtree (QT) structure in which a node (or a parent node) is divided into four child nodes (or child nodes) of the same size, a binary tree (BT) structure in which a node is divided into two child nodes, a ternary tree (TT) structure in which a node is divided into three child nodes at a ratio of 1:2:1, or a structure adopting one or more of the QT structure, the BT structure, and the TT structure. For example, a quadtree plus binary tree (QTBT) structure may be used, or a quadtree plus ternary tree (QTBTTT) structure may be used.
[0029] Figure 3 is an example diagram of block partitioning using the QTBT structure. In Figure 3 , (a) illustrates the partitioning of a block according to the QTBT structure, and (b) represents the partitioning according to the tree structure. In Figure 3 , the solid line represents the partitioning according to the QT structure, and the dashed line represents the partitioning according to the BT structure. In Figure 3 of (b), regarding the symbols of the layers, the layer expression without parentheses represents the layer of the QT, and the layer expression in parentheses represents the layer of the BT. In the BT structure represented by the dashed line, the numbers are partitioning type information.
[0030] As Figure 3 shows, the CTU can be initially divided according to the QT structure. The QT division can be repeated until the size of the divided block reaches the minimum block size MinQTSize of the leaf node allowed in the QT. A first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four nodes in the lower layer is encoded by the encoder 250 and is signaled to the video decoding device.
[0031] When the leaf node of the QT is not greater than the maximum block size (MaxBTSize) of the root node allowed in the BT, it can be further divided into the BT structure. The BT can have multiple splitting types. For example, in some examples, there may be two splitting types, namely, the type of splitting a block horizontally into two blocks of the same size (i.e., symmetric horizontal splitting) and the type of splitting a block vertically into two blocks of the same size (i.e., symmetric vertical splitting). The second flag (BT_split_flag) indicating whether each node of the BT structure is split into lower-layer blocks and the splitting type information indicating the splitting type are encoded by the encoder 250 and sent to the video decoding device. There may be additional types of splitting a node's block into two asymmetric blocks. The asymmetric splitting type can include the type of splitting a block into two rectangular blocks with a size ratio of 1:3 or the type of splitting a node's block in the diagonal direction.
[0032] Alternatively, the QTBTTT structure can be used. In the QTBTTT structure, the CTU can be initially divided into the QT structure, and then the leaf node of the QT can be divided into one or more of the BT structure or the TT structure. The TT structure can also have multiple splitting types. For example, regarding splitting, there may be two splitting types: one type is to horizontally split the block of the corresponding node into three blocks at a ratio of 1:2:1 (i.e., symmetric horizontal splitting), and the other type is to vertically split at a ratio of 1:2:1 (i.e., symmetric vertical splitting). In the case of QTBTTT, not only the flag indicating whether each node is split into lower-layer blocks and the splitting type information (or splitting direction information) indicating the splitting type (or splitting direction) but also the supplementary information for distinguishing whether the splitting structure is the BT structure or the TT structure can be signaled to the video decoding device.
[0033] According to the QTBT or QTBTTT splitting of the CTU, the CUs can have various sizes. In the present disclosure, the CUs are not further split for prediction or transformation. That is, the CUs, PUs, and TUs are blocks of the same size and are located at the same position.
[0034] The predictor 220 predicts the CU to generate a prediction block. The predictor 220 predicts the luminance component and the chrominance component that constitute the CU respectively.
[0035] Generally, the CUs within a picture can each be predictively encoded. Generally, intra-frame prediction techniques or inter-frame prediction techniques can be used to implement the prediction of the current block. Intra-frame prediction techniques use the data in the picture containing the CU, and inter-frame prediction techniques use the data of the pictures encoded before the picture containing the CU. Inter-frame prediction includes both uni-directional prediction and bi-directional prediction. For this purpose, the predictor 220 includes an intra-frame predictor 222 and an inter-frame predictor 224.
[0036] The intra predictor 222 uses the pixels (reference samples) in the current picture including the CU that are located around the CU to predict the pixels in the CU. Depending on the prediction direction, there are multiple intra prediction modes. For example, as Figure 4 shown, the multiple intra prediction modes may include a non - directional mode and 65 directional modes, and the non - directional mode may include a planar mode and a DC mode. Depending on each prediction mode, the formula and peripheral pixels to be used are defined differently.
[0037] The intra predictor 222 may determine the intra prediction mode to be used when encoding the CU. In some examples, the intra predictor 222 may encode the CU using multiple intra prediction modes and select an appropriate intra prediction mode to be used from the tested modes. For example, the intra predictor 222 may calculate rate - distortion values using rate - distortion analysis of multiple tested intra prediction modes and may select the intra mode with the best rate - distortion characteristics from among the tested modes.
[0038] The intra predictor 222 selects one intra prediction mode from among the multiple intra prediction modes and uses the formula determined according to the selected intra prediction mode and neighboring pixels (reference pixels) to predict the CU. The syntax for indicating the selected intra prediction mode is encoded by the encoder 250 and sent to the video decoding device. The intra prediction mode selected for predicting the luminance component in the CU may be used to predict the chrominance component. However, the present disclosure is not limited thereto. For example, the intra prediction mode selected for the luminance component and multiple intra prediction modes for the chrominance component may be configured as candidates, and one of the candidates may be used as the intra prediction mode for the chrominance component. In this case, the syntax for the intra prediction mode corresponding to the chrominance component is signaled separately.
[0039] The inter predictor 224 generates a prediction block for the CU through motion compensation. The inter predictor searches for the block most similar to the CU in a reference picture that was encoded and decoded earlier than the current picture and uses the searched - for block to generate a prediction block for the CU. Then, the inter predictor generates a motion vector corresponding to the displacement between the CU in the current picture and the prediction block in the reference picture. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. Information including information about the reference picture and information about the motion vector used for predicting the CU is encoded by the encoder 250 and sent to the video decoding device. Generally, the information about the reference picture means a reference picture index for identifying the reference picture for inter prediction among multiple reference pictures, and the information about the motion vector means the motion vector difference between the actual motion vector and the predicted motion vector of the CU.
[0040] Other methods can be used to minimize the number of bits required to encode motion information. For example, when the reference picture and motion vector of the current block are the same as those of neighboring blocks, the motion information about the current block can be sent to the decoding device by encoding the information used to identify the neighboring blocks. This method is called "merge mode".
[0041] In merge mode, the inter-prediction unit 224 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from among the neighboring blocks of the current block. As Figure 5 illustrated, all or part of the neighboring blocks used to derive the merge candidates, i.e., the left block L, the above block A, the upper right block AR, the lower left block BL, and the upper left block AL adjacent to the current block in the current picture, can be used. Additionally, blocks located in a reference picture (which may be the same as or different from the reference picture used to predict the current block) rather than in the current picture where the current block is located can be used as merge candidates. For example, a collocated block at the same position as the current block in the reference picture or a block adjacent to the collocated block can also be used as a merge candidate. The inter-prediction unit 224 uses such neighboring blocks to configure a merge list including a predetermined number of merge candidates. A merge candidate to be used as the motion information about the current block is selected from among the merge candidates included in the merge list, and a merge index for identifying the selected candidate is generated. The generated merge index is encoded by the encoder 250 and sent to the decoding device.
[0042] The subtractor 230 subtracts the predicted pixels in the predicted block generated by the intra-prediction unit 222 or the inter-prediction unit 224 from the actual pixels in the CU to generate a residual block.
[0043] The transformer 240 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transformer 240 uses a transform unit of the CU size to transform the residual signal in the residual block. The quantizer 245 quantizes the transform coefficients output from the transformer 240 and outputs the quantized transform coefficients to the encoder 250. Although the transformation and quantization of the residual signal have been described as always being performed, the present disclosure is not limited thereto. Either one or more of the transformation and quantization can be selectively skipped. For example, only one of the transformation and quantization can be skipped, or both the transformation and quantization can be skipped.
[0044] The encoder 250 encodes information such as the CTU size, QT split flag, BT split flag, and split type associated with block splitting so that the video decoding device can split blocks in the same manner as in the video encoding device.
[0045] In addition, the encoder 250 encodes information required for the video decoding device to reconstruct the CU and sends the information to the video decoding device. In the present disclosure, the encoder 250 initially encodes the prediction syntax required for the predicted CU. The prediction syntax encoded by the encoder 250 includes a skip flag that indicates whether the CU is encoded in the skip mode. Here, the skip mode is a special case of the merge mode and is different from the merge mode in that after encoding the merge index (merge_idx), no information about the CU is encoded. Therefore, in the skip mode, the transform syntax is not encoded, and all coefficients in the CU are set to 0. When the CU is encoded in the skip mode, the video decoding device uses the motion vector and reference picture of the merge candidate indicated by the merge index as the motion vector and reference picture of the current CU to generate a prediction block. Since the residual signals are all set to 0, the prediction block is reconstructed as the CU. In addition, it is possible to apply deblocking filtering or SAO filtering to the CU encoded in the skip mode through syntax signaled at a higher level than the CU (e.g., CTU, slice, PPS, etc.). For example, when the syntax (slice_sao_flag) indicating whether to apply SAO is signaled slice by slice, it is determined whether to apply SAO to the CU encoded in the skip mode in the corresponding slice according to the syntax (slice_sao_flag). Alternatively, when the CU is encoded in the skip mode, at least some in-loop filtering can be skipped. For example, the transquant_skip_flag described later can be automatically set to 1, and the transform, quantization, and at least some in-loop filtering for the CU can be skipped.
[0046] When the CU is not encoded in the skip mode, the prediction syntax includes prediction type information (pred_mode_flag) indicating whether the CU is encoded by intra prediction or inter prediction and intra prediction information (i.e., information about the intra prediction mode) or inter prediction information (information about the reference picture and motion vector) according to the prediction type. The inter prediction information includes a merge flag (merge_flag) indicating whether the reference picture and motion vector of the CU are encoded in the merge mode, and includes the merge index (merge_idx) when merge_flag is 1, or picture information and motion vector difference information when merge_flag is 0. In addition, prediction motion vector information may be additionally included.
[0047] After encoding the prediction syntax, the encoder 250 encodes the transform syntax required for the transform CU. The transform syntax includes information related to selective skipping of transform and quantization and information about the coefficients in the CU. Here, when there are two or more transform methods, the transform syntax includes information indicating the type of transform to be used for the CU. Here, the types of transforms applied to the horizontal axis direction and the vertical axis direction can be included separately.
[0048] The information related to selective skipping of transform and quantization can separately include a transform / quantization skip flag (transquant_skip_flag) indicating whether to skip the transform, quantization, and at least some in-loop filtering for the CU, and a transform skip flag (transform_skip_flag) indicating whether to skip the transform for each of the luminance component and the chrominance component constituting the CU. The transform skip flag (transform_skip_flag) is encoded independently for each of the components constituting the CU. However, the present disclosure is not limited thereto, and the transform skip flag may be encoded only once for one CU. In this case, when the transform_skip_flag is 1, the transform of the CU is skipped, that is, the transforms of both the luminance component and the chrominance component constituting the CU are skipped.
[0049] The information related to selective skipping of transform and quantization can be represented by a syntax (transquant_idx). In this case, the syntax can have four values. For example, transquant_idx = 0 means skipping both transform and quantization, transquant_idx = 1 means performing only quantization, transquant_idx = 2 means performing only transform, and transquant_idx = 3 means performing both transform and quantization. The syntax can be binarized using a fixed length (FL) binarization method so that these four values are the same number of bits. For example, it can be binarized as shown in Table 1.
[0050] [Table 1]
[0051] transquant_idx Transformation Quantization Binarization (ID) 0 Disable Disable 00 1 Disable Enable 01 2 Enable Disable 10 3 Enable Enable 11
[0052] Alternatively, the syntax can be binarized using a truncated unary (TU) binarization method so that fewer bits are allocated to the value with a higher occurrence probability. For example, the probability of performing both transform and quantization is usually the highest, followed by the probability of performing only quantization when the transform is skipped. Therefore, the syntax can be binarized as shown in Table 2.
[0053] [Table 2]
[0054] transquant_idx Transformation Quantization Binarization (TU) 0 Disable Disable 111 1 Disable Enable 10 2 Enable Disable 110 3 Enable Enable 0
[0055] Information about coefficients in a CU includes coding block flags (cbf) indicating whether there are non-zero coefficients in the luminance component and two chrominance components of the CU, and syntax for indicating coefficient values. Here, a "coefficient" can be a quantized transform coefficient (when both transformation and quantization are performed), or can be a quantized residual signal obtained by skipping transformation (when transformation is skipped), or a residual signal (when both transformation and quantization are skipped).
[0056] In the present disclosure, the structure or order in which the encoder 250 encodes the prediction syntax and transformation syntax for a CU is the same as the structure or order in which the decoder 610 of the video decoding device described later decodes the encoded prediction syntax and transformation syntax. Since the structure or order in which the decoder 610 decodes the syntax can be clearly understood to understand the structure or order in which the encoder 250 encodes the syntax, details of the syntax encoding structure or order of the encoder 250 are omitted to avoid redundant description.
[0057] The inverse quantizer 260 inverse quantizes the quantized transform coefficients output from the quantizer 245 to generate transform coefficients. The inverse transformer 265 transforms the transform coefficients output from the inverse quantizer 260 from the frequency domain to the spatial domain and reconstructs the residual block.
[0058] The adder 270 adds the reconstructed residual block to the prediction block generated by the predictor 220 to reconstruct the CU. The pixels in the reconstructed CU are used as reference samples for performing intra prediction of the next block in sequence.
[0059] The in-loop filter 280 filters the reconstructed pixels to reduce block artifacts, ringing artifacts, and blurring artifacts caused by block-based prediction and transform / quantization. The in-loop filter 280 may include a deblocking filter 282 and an SAO filter 284. The deblocking filter 180 filters the boundaries between the reconstructed blocks to remove block artifacts caused by block-by-block encoding / decoding, and the SAO filter 284 performs additional filtering on the deblocked picture. The SAO filter 284 is used to compensate for the difference between the reconstructed pixels and the original pixels caused by lossy coding. Since the deblocking filter and SAO are filtering techniques defined in the HEVC standard technology, further detailed description thereof is omitted.
[0060] The reconstructed blocks filtered by the deblocking filter 282 and the SAO filter 284 are stored in the memory 290. Once all the blocks in a picture are reconstructed, the reconstructed picture is used as a reference picture for performing inter prediction on the blocks in the picture to be encoded.
[0061] Figure 6 is an exemplary block diagram of a video decoding device capable of implementing the technology of the present disclosure.
[0062] The video decoding device includes a decoder 610 and a reconstructor 600. The reconstructor 600 includes an inverse quantizer 620, an inverse transformer 630, a predictor 640, an adder 650, an in-loop filter 660, and a memory 670. Similar to Figure 2 the video encoding device described above, each element of the video decoding device can be implemented as a hardware chip or can be implemented as software, and one or more microprocessors can be implemented to execute the functions of the software corresponding to each element.
[0063] The decoder 610 decodes the bitstream received from the video encoding device. The decoder 610 determines the CU to be decoded by decoding the information related to block partitioning. The decoder 610 extracts the information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), determines the size of the CTU, and divides the picture into CTUs of the determined size. Then, the decoder determines that the CTU is the top layer (i.e., the root node) of the tree structure and extracts the partitioning information of the CTU, thereby using the tree structure to partition the CTU. For example, when the CTU is partitioned using the QTBT structure, the first flag (QT_split_flag) related to QT partitioning is extracted to divide each node into four lower-layer nodes. For the node corresponding to the leaf node of QT, the partitioning type (partitioning direction) information and the second flag (BT_split_flag) related to BT partitioning are extracted, and the corresponding leaf node is partitioned according to the BT structure. Another example is that when the CTU is partitioned using the QTBTTT structure, the first flag (QT_split_flag) related to QT partitioning is extracted, and each node is divided into four lower-layer nodes. In addition, for the node corresponding to the leaf node of QT, the split_flag indicating whether the node is further partitioned into BT or TT, the additional information for distinguishing the BT structure or the TT structure, and the partitioning type (or partitioning direction) information are extracted. Thus, each node below the leaf node of QT is recursively partitioned into a BT or TT structure.
[0064] The decoder 610 also decodes the prediction syntax and the transform syntax necessary for reconstructing the CU according to the bitstream. In this operation, the decoder decodes the prediction syntax first and then decodes the transform syntax. Subsequently, the structure of the decoder 610 for decoding the prediction syntax and the transform syntax will be described with reference to Figure 7 the following drawings.
[0065] The inverse quantizer 620 inverse quantizes the coefficients derived using the transform syntax. The inverse transformer 630 inverse transforms the inverse quantized coefficients from the frequency domain to the spatial domain to reconstruct the residual signal, thereby generating a residual block for the CU. One or more of the inverse quantization or inverse transformation may be skipped according to information related to the selective skipping of the transform and quantization included in the transform syntax decoded by the decoder.
[0066] The predictor 640 generates a prediction block for the CU using the prediction syntax. The predictor 640 includes an intra predictor 642 and an inter predictor 644. The intra predictor 642 is activated when the prediction type of the CU is intra prediction, and the inter predictor 644 is activated when the prediction type of the CU is inter prediction.
[0067] The intra predictor 642 determines the intra prediction mode of the CU among a plurality of intra prediction modes using the intra prediction information extracted from the decoder 610, and predicts the CU using the reference samples around the CU according to the intra prediction mode.
[0068] The inter predictor 644 uses the inter prediction information extracted from the decoder 610 to determine the motion vector of the CU and the reference picture for the motion vector, and uses the motion vector and the reference picture to predict the CU. For example, when merge_flag is 1, after configuring the merge list and extracting merge_idx in the same manner as the image encoding device, the inter predictor 644 sets the motion vector and reference picture of the current CU to the reference picture and motion vector of the block indicated by merge_idx among the merge candidates included in the merge list. On the other hand, when merge_flag is 0, the reference picture information and motion vector difference information are extracted from the bitstream to determine the reference picture and motion vector of the current CU. In this case, the predicted motion vector information may be additionally extracted.
[0069] The adder 650 adds the residual block output from the inverse transformer to the prediction block output from the inter predictor or intra predictor to reconstruct the CU. The pixels in the reconstructed CU are used as reference samples for intra prediction of the blocks to be decoded later.
[0070] By sequentially reconstructing the CUs, the CTU composed of the CUs and the picture composed of the CTUs are reconstructed.
[0071] The in-loop filter 660 includes a deblocking filter 662 and a SAO filter 664. The deblocking filter 662 performs deblocking filtering on the boundaries between reconstructed blocks to remove block artifacts generated due to block-by-block decoding. The SAO filter 664 performs additional filtering on the reconstructed blocks after deblocking filtering to compensate for the difference between the reconstructed pixels and the original pixels caused by lossy coding. The reconstructed blocks filtered by the deblocking filter 662 and the SAO filter 664 are stored in the memory 670. When all the blocks in a picture are reconstructed, the reconstructed picture is used as a reference picture for inter prediction of the blocks in the subsequent picture to be decoded.
[0072] Hereinafter, the structure or process of the decoder 610 of the video decoding device for decoding the CU syntax is described in detail.
[0073] As described above, the video decoding device initially decodes the prediction syntax of the coding unit including the skip flag. After decoding all the prediction syntax, the video decoding device decodes the transform syntax including the transquant_skip_flag and the coding unit cbf (cbf_cu). Here, cbf_cu is a syntax indicating whether all the coefficients in two chrominance blocks and one luminance block constituting the coding unit are 0. When the skip flag is 0, that is, when the prediction type of the current CU is not in the skip mode, the transform syntax is decoded.
[0074] Figure 7 is an exemplary flowchart for decoding the CU syntax according to the present disclosure.
[0075] The video decoding device initially decodes the skip flag among other prediction syntax (S702). When the skip flag is 1, as described above, the syntax in the CU except for the merge_idx is not included in the bitstream. Therefore, when the skip flag is 1 (S704), the video decoding device decodes the merge_idx according to the bitstream (S706), and ends the syntax decoding for the CU. When the CU has been encoded in the skip mode, it is possible to determine whether to apply deblocking filtering or SAO filtering to the CU by the syntax signaled at a higher level (e.g., CTU, slice, PPS, etc.) than the CU. Alternatively, for the CU that has been encoded in the skip mode, at least some in-loop filtering can be skipped. For example, the transquant_skip_flag can be automatically set to 1, so at least some in-loop filtering can be skipped.
[0076] When skip_flag is 0 (S704), the prediction syntax for inter prediction or intra prediction is decoded. Initially, the video decoding device decodes prediction type information (pred_mode_flag) indicating whether the prediction type of the current CU is intra prediction or inter prediction (S708). When pred_mode_flag indicates intra prediction (e.g., pred_mode_flag = 1) (S710), the video decoding device decodes intra prediction information indicating the intra prediction mode of the current CU (S712). On the other hand, when pred_mode_flag indicates inter prediction (e.g., pred_mode_flag = 0) (S710), the video decoding device decodes merge_flag (S714). Then, when merge_flag indicates the merge mode (e.g., merge_flag = 1) (S716), merge_idx is decoded (S718). When merge_flag indicates that the mode is not the merge mode (e.g., merge_flag = 0) (S716), the prediction syntax for normal inter prediction (i.e., reference picture information and motion vector difference) is decoded (S720). In this case, prediction motion vector information may additionally be decoded. After decoding all the prediction syntax of the current CU in this way, the video decoding device decodes the transform syntax of the current CU.
[0077] The video decoding device first decodes cbf_cu indicating whether the coefficients in the two chrominance blocks and one luminance block constituting the CU are all 0. However, when the prediction type of the current CU is intra prediction, cbf_cu may be automatically set to 1 without being decoded according to the bitstream (S736). Additionally, when the prediction type of the current CU is the skip mode, cbf_cu may be automatically set to 0. When the prediction type is the merge mode, cbf_cu may be automatically set to 1 (S732). When the prediction type of the current CU is inter prediction rather than the merge mode, cbf_cu is decoded according to the bitstream (S734).
[0078] After decoding cbf_cu, the video decoding device decodes transquant_skip_flag according to the value of cbf_cu. For example, when cbf_cu is 0 (S738), this means that there are no non-zero luminance components and non-zero chrominance components in the CU, so the syntax decoding for the CU is terminated without further decoding of the transform syntax. When cbf_cu is 1 (S738), this means that there are luminance components or chrominance components with non-zero values in the CU, so transquant_skip_flag is decoded (S740). Subsequently, the transform syntax for each of the luminance components and chrominance components in the CU is decoded (S742).
[0079] Alternatively, the decoding order of cbf_cu and transquant_skip_flag can be changed. Figure 8 is another exemplary flowchart for decoding CU syntax according to the present disclosure.
[0080] Since Figure 8 the operations S802 to S820 for decoding the prediction syntax in Figure 7 are the same as S702 to S720 for decoding the prediction syntax in
[0081] only the operations for decoding the transform syntax of the CU will be described below.
[0082] When the prediction type of the current CU is not in the skip mode, after decoding the prediction syntax, the video decoding device decodes transquant_skip_flag according to the bitstream (S832). After decoding transquant_skip_flag, cbf_cu is decoded. For example, when the prediction type of the current CU is intra prediction (pred_mode_flag = 1) or the current CU is in the merge mode (merge_flag = 1) (S834), cbf_cu is not decoded according to the bitstream and is automatically set to 1 (S836). On the other hand, when the prediction type of the current CU is inter prediction (pred_mode_flag = 0) and the current CU is not in the merge mode (merge_flag = 0) (S834), cbf_cu is decoded according to the bitstream (S838).
[0082] Thereafter, when cbf_cu is 1, the video decoding device decodes the transform syntax for each of the luminance components and chrominance components in the CU (S842).
[0083] In the above embodiments, transquant_skip_flag is a syntax that indicates whether to skip the transformation, quantization, and at least some in-loop filtering of a CU. Here, "at least some in-loop filtering" may include both deblocking filtering and SAO filtering. Alternatively, it may mean in-loop filtering other than deblocking filtering. In this case, the video decoding device decodes the transquant_skip_flag. When transquant_skip_flag is 1, the transformation, quantization, and SAO filtering of the CU are skipped. When transquant_skip_flag is 1, the video decoding device may further decode a deblocking filter flag that indicates whether to perform deblocking filtering on the CU. In this case, when deblocking_filter_flag is 0, all transformation, quantization, deblocking filtering, and SAO filtering of the CU are skipped. On the other hand, when deblocking_filter_flag is 1, the transformation, quantization, and SAO filtering of the CU are skipped, and deblocking filtering is performed.
[0084] Pixels in a CU are composed of three color components: one luminance component (Y) and two chrominance components (Cb, Cr). Hereinafter, a block composed of luminance components is referred to as a luminance block, and a block composed of chrominance components is referred to as a chrominance block. A CU is composed of one chrominance block and two chrominance blocks. For each of the luminance block and the chrominance block that constitute a CU, the video decoding device described in the present disclosure performs an operation of decoding a transformation syntax for obtaining coefficients in the CU. For example, this operation corresponds to Figure 7 operation S742 or Figure 8 operation S842. However, it is obvious that it is not necessary to combine Figure 7 and Figure 8 to decode the transformation syntax for each component, and the decoding can be applied to a CU syntax structure other than the syntax structures in Figure 7 and Figure 8 .
[0085] Figure 9 is an exemplary flowchart for decoding the transformation syntax for the corresponding luminance component and chrominance component.
[0086] The video decoding device decodes a first chrominance cbf (e.g., cbf_cb) indicating whether there is at least one non - zero coefficient in a first chrominance block (e.g., a chrominance block composed of the Cb component) constituting a CU, a second chrominance cbf (e.g., cbf_cr) indicating whether there is at least one non - zero coefficient in a second chrominance block (e.g., a chrominance block composed of the Cr component), and a luminance block (e.g., cbf_luma) indicating whether there is at least one non - zero coefficient in the luminance block constituting the CU according to the bitstream (S902). The decoding order of these three component cbf's is not restricted, but for example, these three component cbf's can be decoded in the order of cbf_cb, cbf_cr, and cbf_luma.
[0087] When the first two cbf's decoded in S902 are 0 (all coefficients in the block corresponding to each cbf are 0), the last cbf may not be decoded. Refer to Figure 7 or Figure 8 , when cbf_cu is 1, cbf_cb, cbf_cr, and cbf_luma are decoded. Therefore, when the cbf's of two components are both 0, the cbf of the other component must be 1.
[0088] After decoding the cbf's of the three components, S904 to S910 are performed for each component. As an example, for the luminance component, the video decoding device determines whether cbf_luma = 1 (S904). When cbf_luma = 0, this means that there are no non - zero coefficients in the luminance block, so all values in the luminance block are set to 0.
[0089] On the other hand, when cbf_luma = 1, the video decoding device decodes a transform_skip_flag indicating whether to perform a transform on the luminance block. Refer to Figure 7 or Figure 8 , when the decoded transquant_skip_flag is 1, the transform and quantization of the CU (all components in the CU) are skipped. Therefore, when transquant_skip_flag is 1, it is not necessary to decode the transform_skip_flag for the luminance component. On the other hand, when transquant_skip_flag = 0, this means that the transform and quantization of all components in the CU are not always skipped. Therefore, when transquant_skip_flag is 0, the video decoding device decodes a transform_skip_flag indicating whether to perform a transform on the luminance component (S906, S908).
[0090] Then, the video decoding device decodes the coefficient values of the luminance component according to the bitstream (S910).
[0091] It has been described that the cbf' of the three components constituting the CU is decoded in S902. Hereinafter, another embodiment of decoding the cbf' of the three components constituting the CU will be described with reference to Figure 10 In this embodiment, a syntax called cbf_chroma indicating whether all coefficients in two chroma blocks constituting the CU are 0 is also defined.
[0092] Figure 10 is an exemplary flowchart for decoding the cbf of the three components constituting the CU.
[0093] The video decoding device decodes the cbf_chroma indicating whether all coefficients in two chroma blocks constituting the CU are 0 (S1002).
[0094] When cbf_chroma is 0 (S1004), the cbf' of the two chroma blocks (i.e., cbf_cb and cbf_cr) are both set to 0 (S1006). Referring to Figure 7 or Figure 8 , when cbf_cu is 1, the cbf' of the three components is decoded. Therefore, when cbf_chroma is 0, cbf_luma is automatically set to 1 (S1006).
[0095] On the other hand, when cbf_chroma is 1, the video decoding device decodes cbf_cb (S1008). cbf_chroma = 1 means that there are non-zero coefficients in at least one of the two chroma blocks. Therefore, when the cbf of one of the two chroma blocks is 0, the cbf of the other chroma block should be 1. Therefore, when the decoded cbf_cb is 0, cbf_cr is automatically set to 1 (S1010, S1014). On the other hand, when cbf_cb is 1, cbf_cr is decoded according to the bitstream (S1012). Here, the decoding order of cbf_cb and cbf_cf can be changed. When the decoding of cbf_cb and cbf_cr is completed, the video decoding device decodes cbf_luma (S1016).
[0096] The operations after decoding cbf_cb, cbf_cr, and cbf_luma are the same as those in Figure 9 .
[0097] Although exemplary embodiments have been described for illustrative purposes, those skilled in the art should appreciate that various modifications and changes can be made without departing from the spirit and scope of the embodiments. The exemplary embodiments have been described for the sake of brevity. Accordingly, those skilled in the art will understand that the scope of the embodiments is not limited by the embodiments explicitly described above, but is included in the claims and their equivalents.
[0098] Cross - reference to related applications
[0099] This application claims priority to Patent Application No. 10 - 2018 - 0001728, filed in Korea on January 5, 2018, and Patent Application No. 10 - 2018 - 0066664, filed in Korea on June 11, 2018.
Claims
1. A method for decoding a video, the method comprises the following steps: Determine the coding unit to be decoded by partitioning the coding tree units in the picture in a tree structure; Decode the prediction syntax related to the prediction of the coding unit, wherein, perform the prediction of the coding unit according to the size of the coding unit, wherein, the prediction syntax includes a skip flag, and the skip flag indicates whether to code the coding unit in a skip mode, and wherein, the prediction syntax further includes prediction type information, and the prediction type information indicates whether the coding unit is intra-frame predicted or inter-frame predicted when the skip flag does not indicate that the coding unit is coded in the skip mode; After decoding the prediction syntax, decode the transform syntax related to the transform of the coding unit, wherein, perform the transform of the coding unit according to the size of the coding unit, and wherein, the transform syntax includes a coding unit coding block flag cbf for indicating whether all coefficients in the luminance block and two chrominance blocks constituting the coding unit are all zero; and Reconstruct the coding unit based on the prediction syntax and the transform syntax, wherein, the transform syntax is not decoded until the decoding of the prediction syntax is completed, and when the skip flag indicates that the coding unit is coded in the skip mode, the transform syntax is not decoded, wherein, when the skip flag does not indicate that the coding unit is coded in the skip mode, decoding the transform syntax includes: When the prediction type of the coding unit is inter-frame prediction and not in the merge mode, decode the coding unit cbf according to the bitstream, and When the prediction type of the coding unit is intra-frame prediction or the merge mode, infer that the coding unit cbf is a value of 1 without decoding the coding unit cbf according to the bitstream, and the value of 1 indicates that at least part of the luminance block and the two chrominance blocks have non-zero coefficients.
2. The method according to claim 1, wherein, Reconstructing the coding unit comprises the following steps: Use the size of the coding unit as the prediction unit size, and generate a prediction block of the coding unit according to the prediction syntax; Use the size of the coding unit as the transform unit size, and generate a residual block of the coding unit according to the transform syntax; and Reconstruct the coding unit by adding the prediction block and the residual block.
3. The method according to claim 1, wherein, Decoding the transform syntax comprises the following steps: Decode the coding unit cbf; and When the coding unit cbf indicates that at least one of the coefficients in the luminance block and the two chrominance blocks constituting the coding unit is non-zero, decode the transform / quantization skip flag, wherein, the transform / quantization skip flag indicates whether to skip at least part of the in-loop filtering, inverse transform and inverse quantization.
4. The method according to claim 3, wherein, The transform / quantization skip flag indicates whether inverse transform, inverse quantization, and in-loop filtering other than deblocking filtering are skipped for the coding unit.
5. The method according to claim 4, wherein, decoding the transform syntax comprises the steps of: when the transform / quantization skip flag indicates that inverse transform, inverse quantization, and in-loop filtering other than deblocking filtering are skipped for the coding unit, decoding a deblocking filter flag indicating whether deblocking filtering is skipped for the reconstructed coding unit.
6. The method according to claim 1, wherein, when the coding unit cbf indicates that at least one of the coefficients in the luminance block and the two chrominance blocks constituting the coding unit is non-zero, decoding the transform syntax comprises the steps of: decoding one of the chrominance cbf and the luminance cbf; and decoding the other of the chrominance cbf and the luminance cbf according to the value of the decoded one of the chrominance cbf and the luminance cbf, wherein the chrominance cbf indicates whether all coefficients in the two chrominance blocks are all zero, and the luminance cbf indicates whether all coefficients in the luminance block are all zero.
7. The method according to claim 6, the method further comprises the steps of: when the chrominance cbf indicates that there is at least one non-zero coefficient in the two chrominance blocks, decoding a first sub-chrominance cbf indicating whether all coefficients in the first chrominance block among the two chrominance blocks are all zero; and when the first sub-chrominance cbf indicates that there is at least one non-zero coefficient in the first chrominance block, decoding a second sub-chrominance cbf indicating whether all coefficients in the second chrominance block among the two chrominance blocks are all zero.
8. A video coding method, the video coding method comprises the steps of: determining a coding unit to be coded by partitioning coding tree units in a picture in a tree structure; encoding a prediction syntax related to the prediction of the coding unit, wherein, the prediction of the coding unit is performed according to the size of the coding unit, wherein the prediction syntax includes a skip flag indicating whether the coding unit is coded in a skip mode, and wherein the prediction syntax further includes prediction type information, and the prediction type information indicates whether the coding unit is intra-predicted or inter-predicted when the skip flag does not indicate that the coding unit is coded in the skip mode; after encoding the prediction syntax, encoding a transform syntax related to the transform of the current block, wherein the transform of the coding unit is performed according to the size of the coding unit, and wherein the transform syntax includes a coding unit coding block flag cbf for indicating whether all coefficients in the luminance block and the two chrominance blocks constituting the coding unit are all zero, wherein, the transform syntax is not encoded until the encoding of the prediction syntax is completed, and when the skip flag indicates that the coding unit is coded in the skip mode, the transform syntax is not encoded. Wherein, when the skip flag does not indicate that the coding unit is coded in the skip mode, coding the transform syntax includes: When the prediction type of the coding unit is inter prediction and not the merge mode, coding the coding unit cbf, and Wherein, when the prediction type of the coding unit is intra prediction or the merge mode, the coding unit cbf is not coded such that the video decoding device infers that the coding unit cbf is a value 1, and the value 1 indicates that at least part of the luminance block and the two chrominance blocks have non-zero coefficients.
9. A method for providing video data to a video decoding device, the method comprising the steps of: Encoding the video data into a bitstream; And Sending the bitstream to the video decoding device, Wherein, encoding the video data includes the steps of: Determining coding units to be coded by partitioning coding tree units in a picture in a tree structure; Coding prediction syntax related to prediction of the coding unit, Wherein, the prediction of the coding unit is performed according to the size of the coding unit, Wherein, the prediction syntax includes a skip flag indicating whether the coding unit is coded in the skip mode, and Wherein, the prediction syntax further includes prediction type information which indicates whether the coding unit is intra predicted or inter predicted when the skip flag does not indicate that the coding unit is coded in the skip mode; and After coding the prediction syntax, coding transform syntax related to transformation of a current block, wherein the transformation of the coding unit is performed according to the size of the coding unit, and Wherein, the transform syntax includes a coding unit coding block flag cbf for indicating whether all coefficients in the luminance block and two chrominance blocks constituting the coding unit are all zero, Wherein, the transform syntax is not coded until the coding of the prediction syntax is completed, and when the skip flag indicates that the coding unit is coded in the skip mode, the transform syntax is not coded, Wherein, when the skip flag does not indicate that the coding unit is coded in the skip mode, coding the transform syntax includes: When the prediction type of the coding unit is inter prediction and not the merge mode, coding the coding unit cbf, and Wherein, when the prediction type of the coding unit is intra prediction or the merge mode, the coding unit cbf is not coded such that the video decoding device infers that the coding unit cbf is a value 1, and the value 1 indicates that at least part of the luminance block and the two chrominance blocks have non-zero coefficients.
Citation Information
Patent Citations
Super-block for high performance video coding
CN102835107A
Image decoding apparatus, image encoding apparatus, and data structure of encoded data
CN103460694A