Method for decoding and encoding a video sequence and method for providing encoded data

By resampling and inter-frame prediction of frames with various sampling formats in the video sequence, the problem of low encoding efficiency in the prior art is solved, and lower memory consumption and delay are achieved.

CN115104307BActive Publication Date: 2025-06-03HYUNDAI MOTOR CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180014578.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-02
Filing Date
2021-02-17
Publication Date
2025-06-03
Estimated Expiration
2041-02-17

AI Technical Summary

Technical Problem

The existing video encoding technology is not very encoding efficient when processing frames with luminance signals and chrominance signals of various sampling formats, resulting in increased memory consumption and delay.

Method used

Inter prediction is performed on the current picture by resampling and referring to the luminance and chrominance signals of the reference picture to match the chrominance signal resolution and format of the current picture and reference picture.

Benefits of technology

Improves the encoding efficiency of video encoding and decoding, reduces memory consumption and delay, and is suitable for high resolution and high frame rate video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104307B_ABST
    Figure CN115104307B_ABST
Patent Text Reader

Abstract

The present disclosure relates to video encoding and decoding based on resampled chrominance signals. This embodiment provides a video encoding / decoding method, which is used to perform inter-frame prediction on a current picture by resampling and referring to the luminance signal and chrominance signal of a reference picture, so as to improve the encoding efficiency in the video encoding and decoding of frames with luminance signals and chrominance signals of various sampling formats in a video sequence.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to Korean Patent Application Nos. 10-2020-0018864, filed on February 17, 2020, 10-2020-0025838, filed in Korea on March 2, 2020, and 10-2021-0021015, filed in Korea on February 17, 2021, the disclosures of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to image (video) encoding and decoding, and more particularly, to a video encoding / decoding method for performing inter prediction on a current picture by resampling and referring to luminance signals and chrominance signals of reference pictures having various sampling formats. Background Art

[0004] The statements in this section merely provide background information related to the present disclosure and may not constitute prior art.

[0005] Since the volume of video data is larger than that of voice data or still image data, storing or transmitting video data without compression processing requires a large amount of hardware resources including memory.

[0006] Therefore, when storing or transmitting video data, an encoder is typically used to compress the video data for storage or transmission. Then, a decoder receives the compressed video data and decompresses and reproduces the video data. Such video compression techniques include H.264 / AVC and High Efficiency Video Coding (HEVC), which improve the coding efficiency by about 40% compared to H.264 / AVC.

[0007] However, video size, resolution, and frame rate are gradually increasing, and thus, the amount of data to be encoded also increases. Therefore, new compression techniques with better coding efficiency and higher image quality than existing compression techniques are needed.

[0008] In image (video) encoding / decoding, referring to previously decoded images can encode / decode a current picture to improve coding efficiency. For example, in the case of video encoding / decoding of frames having luminance signals and chrominance signals with various chroma formats (e.g., 4:4:4, 4:2:2, and 4:2:0) in a video sequence, the luminance signals and chrominance signals of the current picture may have a different resolution from those of the reference picture. In such a case, a method for encoding / decoding the current picture by correcting the resolution of the reference picture is needed. Summary of the Invention

[0009] In order to improve the coding efficiency in video encoding and decoding of frames with luminance signals and chrominance signals having various sampling formats in a video sequence, the present disclosure performs inter-frame prediction on a current picture by resampling and referring to the luminance signal and chrominance signal of a reference picture. The object of the present disclosure is to provide a video encoding / decoding method for reducing memory consumption and latency during the encoding / decoding process.

[0010] According to one aspect of the present disclosure, there is provided a video decoding method for a current block in a current picture including a chrominance signal having a resolution and chrominance format separated from a luminance signal, the video decoding method including: obtaining size information and a chrominance format of the current picture, and generating a resolution of the chrominance signal of the current picture according to the size information and the chrominance format; obtaining a decoded residual signal and inter-frame prediction information about the current block, where the inter-frame prediction information includes a reference picture index and a motion vector; obtaining a resolution and a chrominance format of the chrominance signal of a reference picture specified by the reference picture index; when the resolution or chrominance format of the chrominance signal of the current picture is different from the resolution or chrominance format of the chrominance signal of the reference picture, resampling the chrominance signal of the reference picture to match the resolution and chrominance format of the chrominance signal of the reference picture and the resolution and chrominance format of the chrominance signal of the current picture; generating a prediction signal for the current block based on the inter-frame prediction information; and generating a reconstructed block by adding the prediction signal and the residual signal.

[0011] According to another aspect of the present disclosure, there is provided a video encoding method for a current block in a current picture including a chrominance signal having a resolution and chrominance format separated from a luminance signal, the video encoding method including: obtaining size information and a chrominance format of the current picture, and generating a resolution of the chrominance signal of the current picture according to the size information and the chrominance format; obtaining inter-frame prediction information about the current block, where the inter-frame prediction information includes a reference picture index and a motion vector; obtaining a resolution and a chrominance format of the chrominance signal of a reference picture specified by the reference picture index; when the resolution or chrominance format of the chrominance signal of the current picture is different from the resolution or chrominance format of the chrominance signal of the reference picture, resampling the chrominance signal of the reference picture to match the resolution and chrominance format of the chrominance signal of the reference picture and the resolution or chrominance format of the chrominance signal of the current picture; generating a prediction signal for the current block based on the inter-frame prediction information; and generating a residual signal by subtracting the prediction signal from the current block.

[0012] As described above, according to the present embodiment, the encoding efficiency can be improved by providing a video encoding / decoding method for performing inter prediction on a current picture by resampling and referring to the luminance signal and chrominance signal of a reference picture in video encoding and decoding of frames having various sampling formats for the luminance signal and chrominance signal in a video sequence.

[0013] In addition, according to the present embodiment, by providing a video encoding / decoding method for reducing memory consumption and latency in the encoding / decoding process, the bit rate of various contents such as game broadcasts, 360-degree video streams, VR / AR videos, and online lectures can be reduced, the burden on the network and energy consumption of a reproduction device performing video decoding can be alleviated, and fast decoding can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 An exemplary block diagram of a video encoding device for implementing the technology of the present invention.

[0015] Figure 2 A diagram for describing a method of dividing blocks by using a QTBTTT structure.

[0016] FIGS. 3A and 3B are schematic diagrams showing a plurality of intra prediction modes including a wide-angle intra prediction mode.

[0017] Figure 4 An exemplary diagram of adjacent blocks for a current block.

[0018] Figure 5 An exemplary block diagram of a video decoding device for implementing the technology of the present invention.

[0019] Figure 6 Shows a video encoding device including a resampling block according to an embodiment of the present disclosure.

[0020] Figure 7 Shows a video decoding device including a resampling block according to an embodiment of the present disclosure.

[0021] Figure 8 An exemplary diagram showing a frame to which RPR is applied.

[0022] Figure 9 An exemplary diagram showing a reference frame having various chroma formats according to an embodiment of the present disclosure.

[0023] Figure 10 Another exemplary diagram showing a reference frame having various chroma formats according to an embodiment of the present disclosure.

[0024] Figure 11 Another exemplary diagram showing a reference frame having various chroma formats according to an embodiment of the present disclosure.

[0025] Figure 12 is a schematic flowchart of a video decoding method according to an embodiment of the present invention.

[0026] Figure 13 is a block diagram of a video decoding apparatus using ACT according to an embodiment of the present disclosure.

[0027] Figure 14 is a conceptual exemplary diagram showing a video decoding process using ACT according to an embodiment of the present disclosure.

[0028] Figure 15 is a conceptual exemplary diagram showing MRL according to an embodiment of the present disclosure.

[0029] Figure 16 is an exemplary diagram showing the space of a residual signal and a reference sample in ACT according to an embodiment of the present disclosure.

[0030] Figure 17 is an exemplary diagram showing a pipeline for adding a reference sample during intra prediction in the YCgCo space according to an embodiment of the present disclosure.

[0031] Figure 18 is a schematic exemplary diagram showing a decoding process according to an embodiment of the present disclosure.

[0032] Figure 19 is an exemplary diagram showing a process of decoding a single-core bitstream using multiple cores according to an embodiment of the present disclosure.

[0033] Figure 20 is an exemplary diagram showing the operation of a hardware video decoding apparatus according to an embodiment of the present disclosure. Detailed Description of the Embodiments

[0034] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the exemplary drawings. When reference numerals denote components of each drawing, it should be noted that although the same components are shown in different drawings, the same components may be denoted by the same reference numerals. In addition, when describing the embodiments, detailed descriptions of known related configurations and functions may be omitted to avoid unnecessarily obscuring the subject matter of the embodiments.

[0035] Figure 1 is an exemplary block diagram of a video encoding apparatus for implementing the technology of the present disclosure. Hereinafter, with reference to Figure 1 the description of, a video encoding apparatus and sub-components of the apparatus will be described.

[0036] The encoding device may be configured to include a picture partitioning unit 110, a prediction unit 120, a subtractor 130, a transformation unit 140, a quantization unit 145, a rearrangement unit 150, an entropy encoding unit 155, a dequantization unit 160, an inverse transformation unit 165, an adder 170, a loop filter unit 180, and a memory 190.

[0037] Each component of the encoding device may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and a microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0038] A video is composed of one or more sequences including a plurality of pictures. Each picture is divided into a plurality of regions, and encoding is performed on each region. For example, a picture is divided into one or more slices or / and tiles. Here, one or more slices may be defined as a slice group. Each slice or / and tile is divided into one or more coding tree units (CTUs). In addition, each CTU is divided into one or more coding units (CUs) by a tree structure. The information applied to each CU is encoded as the syntax of the CU, and the information commonly applied to the CUs included in a CTU is encoded as the syntax of the CTU. In addition, the information commonly applied to all blocks in a slice is encoded as the syntax of the slice header, and the information applied to all blocks constituting one or more pictures is encoded as a picture parameter set (PPS) or a picture header. In addition, the information commonly referred to by multiple pictures is encoded into a sequence parameter set (SPS). In addition, the information commonly referred to by one or more SPSs is encoded into a video parameter set (VPS). Further, the information commonly applied to a slice or a slice group may also be encoded as the syntax of the slice or slice group header. The syntax included in the SPS, PPS, slice header, slice or slice group header may be referred to as high-level syntax.

[0039] The picture partitioning unit 110 determines the size of a coding tree unit (CTU). The information about the size of the CTU (CTU size) is encoded as the syntax of the SPS or PPS and transmitted to the video decoding device.

[0040] The picture partitioning unit 110 divides each picture constituting the video into a plurality of coding tree units (CTUs) having a predetermined size, and then recursively divides the CTUs using a tree structure. The leaf nodes in the tree structure become coding units (CUs) that are the basic units of encoding.

[0041] The tree structure can be a quadtree (QT) in which a higher node (or parent node) is divided into four lower nodes (or child nodes) of the same size, a binary tree (BT) in which a higher node is divided into two lower nodes, a ternary tree (TT) in which a higher node is divided into three lower nodes at a ratio of 1:2:1, or a structure in which two or more of the QT structure, BT structure, and TT structure are mixed. For example, a quadtree plus binary tree (QTBT) structure can be used or a quadtree plus binary tree ternary tree (QTBTTT) structure can be used. Here, BTTT is added to the tree structure to be called a multi-type tree (MTT).

[0042] Figure 2 is a diagram for describing a method of dividing a block by using the QTBTTT structure.

[0043] As Figure 2 shown, the CTU can first be divided into a QT structure. The quadtree division can be recursive until the size of the divided block reaches the minimum block size (MinQTSize) of the leaf node allowed in the QT. The first flag (QT_split_flag) indicating whether each node of the QT structure is divided into four lower nodes is encoded by the entropy coding unit 155 and signaled to the video decoding device. When the leaf node of the QT is not larger than the maximum block size (MaxBTSize) of the root node allowed in the BT, the leaf node can be further divided into at least one of the BT structure or the TT structure. There can be multiple division directions in the BT structure and / or the TT structure. For example, there can be two directions, that is, the direction in which the block of the corresponding node is horizontally divided and the direction in which the block of the corresponding node is vertically divided. As Figure 2 shown, when the MTT division starts, the second flag (mtt_split_flag) indicating whether the node is divided and the flag indicating the division direction (vertical or horizontal) and / or the flag indicating the division type (binary or ternary) if the node is divided are encoded by the entropy coding unit 155 and signaled to the video decoding device.

[0044] Optionally, before encoding the first flag (QT_split_flag) indicating whether each node is divided into four lower nodes, the CU division flag (split_cu_flag) indicating whether the node is divided can also be encoded. When the value of the CU division flag (split_cu_flag) indicates that each node is not divided, the block of the corresponding node becomes a leaf node in the division tree structure and becomes a decoding unit (CU) that is the basic unit of encoding. When the value of the CU division flag (split_cu_flag) indicates that each node is divided, the video encoding device first starts encoding the first flag by the above scheme.

[0045] When the QTBT is used as another instance of the tree structure, there can be two types, that is, the type in which the block of the corresponding node is horizontally divided into two blocks of the same size (i.e., symmetric horizontal division) and the type in which the block of the corresponding node is vertically divided into two blocks of the same size (i.e., symmetric vertical division). The split flag indicating whether each node of the BT structure is divided into lower-layer blocks and the split type information indicating the split type are encoded by the entropy encoding unit 155 and transmitted to the video decoding device. At the same time, the type in which the block of the corresponding node is divided into two blocks in an asymmetric form with respect to each other can be additionally presented. The asymmetric form can include the form in which the block of the corresponding node is divided into two rectangular blocks with a size ratio of 1:3, or can also include the form in which the block of the corresponding node is divided in the diagonal direction.

[0046] The CU can have different sizes according to the QTBT or QTBTTT division from the CTU. Hereinafter, the block corresponding to the CU to be encoded or decoded (i.e., the leaf node of the QTBTTT) will be referred to as the "current block". Due to the adoption of the QTBTTT division, the shape of the current block can be a rectangular shape in addition to the square shape.

[0047] The prediction unit 120 predicts the current block to generate a prediction block. The prediction unit 120 includes an intra prediction unit 122 and an inter prediction unit 124.

[0048] Generally, each of the current blocks in the picture can be predictively coded. Generally, the prediction of the current block can be performed by using an intra prediction technique (using data from the picture including the current block) or an inter prediction technique (using data from the picture compiled before the picture including the current block). The inter prediction includes both uni-directional prediction and bi-directional prediction.

[0049] The intra prediction unit 122 predicts the pixels in the current block by using the pixels (reference pixels) located at adjacent positions of the current block in the current picture including the current block. According to the prediction direction, there are multiple intra prediction modes. For example, as shown in FIG. 3A, the multiple intra prediction modes can include two non-directional modes (including the planar mode and the DC mode) and 65 directional modes. The adjacent pixels and the arithmetic equations to be used are defined differently according to each prediction mode.

[0050] To perform effective directional prediction on a current block having a rectangular shape, directional patterns (##67 to #80, intra prediction mode #-1 to #-14) illustrated by the dashed arrows in FIG. 3B may additionally be used. The directional pattern may be referred to as a “wide-angle intra prediction mode”. In FIG. 3B, the arrows indicate the corresponding reference samples for prediction and do not represent the prediction direction. The prediction direction is opposite to the direction indicated by the arrows. When the current block has a rectangular shape, the wide-angle intra prediction mode is a mode that performs prediction in a direction opposite to a specific directional pattern without additional bit transmission. In this case, in the wide-angle intra prediction mode, some wide-angle intra prediction modes available for the current block may be determined by the ratio of the width and height of the current block having a rectangular shape. For example, when the current block has a rectangular shape with a height less than the width, wide-angle intra prediction modes (intra prediction mode #67 to #80) having an angle less than 45 degrees are available, and when the current block has a rectangular shape with a width greater than the height, wide-angle intra prediction modes having an angle greater than -135 degrees are available.

[0051] The intra prediction unit 122 may determine the intra prediction to be used for encoding the current block. In some examples, the intra prediction unit 122 may encode the current block by using multiple intra prediction modes and also select an appropriate intra prediction mode to be used from the test modes. For example, the intra prediction unit 122 may calculate rate-distortion values by using rate-distortion analysis for multiple tested intra prediction modes and also select an intra prediction mode having the best rate-distortion characteristics among the tested modes.

[0052] The intra prediction unit 122 selects one intra prediction mode from multiple intra prediction modes and predicts the current block by using adjacent pixels (reference pixels) and an arithmetic equation determined according to the selected intra prediction mode. Information about the selected intra prediction mode is encoded by the entropy coding unit 155 and transmitted to the video decoding device.

[0053] The inter prediction unit 124 generates a prediction block for the current block by using motion compensation processing. The inter prediction unit 124 searches for the block most similar to the current block in a reference picture that is encoded and decoded earlier than the current picture and generates a prediction block for the current block by using the searched block. Additionally, a motion vector (MV) corresponding to the displacement between the current block in the current picture and the prediction block in the reference picture is generated. Generally, motion estimation is performed on the luminance component, and the motion vector calculated based on the luminance component is used for both the luminance component and the chrominance component. Motion information including information about the reference picture and information about the motion vector used for predicting the current block is encoded by the entropy coding unit 155 and transmitted to the video decoding device.

[0054] The inter-frame prediction unit 124 may also perform interpolation for a reference picture or a reference block to increase the prediction accuracy. That is, sub-sampling between two consecutive integer samples is interpolated by applying filter coefficients to a plurality of consecutive integer samples including two integer samples. When performing the process of searching for a block most similar to the current block for the interpolated reference picture, the motion vector representation may have a decimal unit accuracy instead of an integer sample unit accuracy. The accuracy or resolution of the motion vector may be set differently for each target region to be encoded (for example, units such as slices, tiles, CTUs, CUs, etc.). When this adaptive motion vector resolution (AMVR) is applied, information about the motion vector resolution to be applied to each target region should be signaled. For example, when the target region is a CU, information about the motion vector resolution applied to each CU is signaled. The information about the motion vector resolution may be information indicating the accuracy of the differential motion vector described below.

[0055] Meanwhile, the inter-frame prediction unit 124 may perform inter-frame prediction by using bidirectional prediction. In the case of bidirectional prediction, two reference pictures and two motion vectors representing the positions of the blocks most similar to the current block in each reference picture are used. The inter-frame prediction unit 124 selects a first reference picture and a second reference picture from reference picture list 0 (RefPicList0) and reference picture list 1 (RefPicList1) respectively, and searches for the block most similar to the current block in each reference picture to generate a first reference block and a second reference block. In addition, a predicted block of the current block is generated by averaging or weighted averaging the first reference block and the second reference block. In addition, motion information including information about the two reference pictures used for predicting the current block and information about the two motion vectors is transmitted to the encoding unit 150. Here, reference picture list 0 may be composed of pictures among the pre-recovered pictures that are before the current picture in the display order, and reference picture list 1 may be composed of pictures among the pre-recovered pictures that are after the current picture in the display order. However, although not particularly limited thereto, pre-recovered pictures after the current picture in the display order may be additionally included in reference picture list 0, and conversely, pre-recovered pictures before the current picture may be additionally included in reference picture list 1.

[0056] Various methods may be used to minimize the number of bits consumed for encoding motion information.

[0057] For example, when the reference picture and the motion vector of the current block are the same as those of an adjacent block, information identifying the adjacent block is encoded to transmit the motion information of the current block to the video decoding device. This method is called the merge mode.

[0058] In the merge mode, the inter prediction unit 124 selects a predetermined number of merge candidate blocks (hereinafter referred to as "merge candidates") from the neighboring blocks of the current block.

[0059] As neighboring blocks for deriving merge candidates, all or some of the left block L, top block A, top-right block AR, bottom-left block BL, and top-left block AL adjacent to the current block in the current picture can be used as Figure 4 illustrated. Further, blocks other than the current picture in which the current block is located within the reference picture (which may be the same as or different from the reference picture used for predicting the current block) can also be used as merge candidates. For example, a block at the same location as the current block within the reference picture or a block adjacent to the block at the same location can also be used as a merge candidate.

[0060] The inter prediction unit 124 configures a merge list including a predetermined number of merge candidates by using the neighboring blocks. A merge candidate to be used as the motion information of the current block is selected from the merge candidates included in the merge list, and merge index information for identifying the selected candidate is generated. The generated merge index information is encoded by the encoding unit 150 and transmitted to the video decoding device.

[0061] Another method for encoding motion information is the advanced motion vector prediction (AMVP) mode.

[0062] In the AMVP mode, the inter prediction unit 124 derives prediction motion vector candidates for the motion vector of the current block by using the neighboring blocks of the current block. As neighboring blocks for deriving prediction motion vector candidates, all or some of the left block L, top block A, top-right block AR, bottom-left block BL, and top-left block AL adjacent to the current block in the current picture as Figure 4 illustrated can be used. In addition, blocks other than the current picture in which the current block is located within the reference picture (which may be the same as or different from the reference picture used for predicting the current block) can also be used as neighboring blocks for deriving prediction motion vector candidates. For example, a block at the same location as the current block within the reference picture or a block adjacent to the block at the same location can be used.

[0063] The inter prediction unit 124 derives prediction motion vector candidates by using the motion vectors of the neighboring blocks, and determines a prediction motion vector for the motion vector of the current block by using the prediction motion vector candidates. In addition, a differential motion vector is calculated by subtracting the prediction motion vector from the motion vector of the current block.

[0064] The predicted motion vector can be obtained by applying a predefined function (e.g., central value and average value calculation, etc.) to the predicted motion vector candidates. In this case, the video decoding device also knows the predefined function. Additionally, since the neighboring blocks used to derive the predicted motion vector candidates are blocks for which encoding and decoding have been completed, the video decoding device may also already know the motion vectors of the neighboring blocks. Therefore, the video encoding device does not need to encode the information used to identify the predicted motion vector candidates. Thus, in this case, the information about the differential motion vector and the information about the reference picture used to predict the current block are encoded.

[0065] Meanwhile, the predicted motion vector can also be determined by a scheme of selecting any one of the predicted motion vector candidates. In this case, the information used to identify the selected predicted motion vector candidate is additionally encoded together with the information about the differential motion vector and the information about the reference picture used to predict the current block.

[0066] The subtractor 130 generates a residual block by subtracting the predicted block generated by the intra prediction unit 122 or the inter prediction unit 124 from the current block.

[0067] The transform unit 140 transforms the residual signal in the residual block having pixel values in the spatial domain into transform coefficients in the frequency domain. The transform unit 140 can transform the residual signal in the residual block by using the total size of the residual block as the transform unit, or can also divide the residual block into a plurality of sub-blocks and perform the transform by using the sub-blocks as the transform unit. Optionally, the residual block is divided into two sub-blocks as a transform region and a non-transform region to transform the residual signal by only using the transform region sub-block as the transform unit. Here, the transform region sub-block can be one of two rectangular blocks having a size ratio of 1:1 based on the horizontal axis (or vertical axis). In this case, a flag (cu_sbt_flag) indicates that only the sub-block is transformed, and the direction (vertical / horizontal) information (cu_sbt_horizontal_flag) and / or the position information (cu_sbt_pos_flag) are encoded by the entropy encoding unit 155 and signaled to the video decoding device. Additionally, the size of the transform region sub-block can have a size ratio of 1:3 based on the horizontal axis (or vertical axis), in which case the flag (cu_sbt_quad_flag) for dividing the corresponding division is additionally encoded by the entropy encoding unit 155 and signaled to the video decoding device.

[0068] Meanwhile, the transform unit 140 can perform transforms on the residual blocks separately in the horizontal and vertical directions. For the transforms, different types of transform functions or transform matrices can be used. For example, a pair of transform functions for horizontal and vertical transforms can be defined as a multiple transform set (MTS). The transform unit 140 can select a pair of transform functions with the highest transform efficiency in the MTS and transform the residual blocks in each of the horizontal and vertical directions. Information (mts_idx) about the pair of transform functions in the MTS is encoded by the entropy coding unit 155 and signaled to the video decoding device.

[0069] The quantization unit 145 quantizes the transform coefficients output from the transform unit 140b using quantization parameters and outputs the quantized transform coefficients to the entropy coding unit 155. The quantization unit 145 can also immediately quantize the relevant residual blocks without performing a transform for any block or frame. The quantization unit 145 can also apply different quantization coefficients (scaling values) according to the positions of the transform coefficients in the transform block. The quantization matrix applied to the quantized transform coefficients set in a two-dimensional manner can be encoded and sent to the video decoding device.

[0070] The rearrangement unit 150 can perform rearrangement of the coefficient values of the quantized residual values.

[0071] The rearrangement unit 150 can change the 2D coefficient array into a 1D coefficient sequence by using coefficient scanning. For example, the rearrangement unit 150 can scan the DC coefficient to the high-frequency domain coefficients by using zigzag scanning or diagonal scanning to output a 1D coefficient sequence. According to the size of the transform unit and the intra prediction mode, vertical scanning that scans the 2D coefficient array in the column direction and horizontal scanning that scans the 2D block type coefficients in the row direction can also be used instead of zigzag scanning. That is, according to the size of the transform unit and the intra prediction mode, the scanning method to be used can be determined among zigzag scanning, diagonal scanning, vertical scanning, and horizontal scanning.

[0072] The entropy coding unit 155 generates a bitstream by encoding the sequence of ID quantized transform coefficients output from the rearrangement unit 150 by using various coding schemes including context-based adaptive binary arithmetic coding (CABAC), exponential Golomb, etc.

[0073] In addition, the entropy coding unit 155 encodes information related to block partitioning (such as CTU size, CTU partitioning flag, QT partitioning flag, MTT partitioning type, MTT partitioning direction, etc.) to allow the video decoding device to partition blocks equivalently to the video encoding device. In addition, the entropy coding unit 155 encodes information on the prediction type indicating whether the current block is encoded by intra prediction or inter prediction, and encodes intra prediction information (i.e., information on the intra prediction mode) or inter prediction information (in the case of the merge mode, merge index, and in the case of the AMVP mode, information on the reference picture index and differential motion vector) according to the prediction type. In addition, the entropy coding unit 155 encodes information related to quantization (i.e., information on the quantization parameter and information on the quantization matrix).

[0074] The dequantization unit 160 dequantizes the quantized transform coefficients output from the quantization unit 145 to generate transform coefficients. The inverse transform unit 165 transforms the transform coefficients output from the dequantization unit 160 from the frequency domain to the spatial domain to recover the residual block.

[0075] The addition unit 170 adds the recovered residual block to the prediction block generated by the prediction unit 120 to recover the current block. When performing intra prediction on the next-order block, the pixels in the recovered current block are used as reference pixels.

[0076] The loop filter unit 180 performs filtering on the recovered pixels to reduce block artifacts, ringing artifacts, blurring artifacts, etc. that occur due to block-based prediction and transform / quantization. The filter unit 180 as the loop filter may include all or some of a deblocking filter 182, a sample adaptive offset (SAO) filter 184, and an adaptive loop filter (ALF) 186.

[0077] The deblocking filter 180 filters the boundaries between the recovered blocks to remove block artifacts that occur due to block-based encoding / decoding, and the SAO filter 184 and the ALF 184 perform additional filtering on the deblocked filtered video. The SAO filter 184 and the ALF 184 are filters for compensating the difference between the recovered pixels and the original pixels that occurs due to lossy encoding. The SAO filter 184 applies an offset as a CTU unit to enhance the visual image quality and coding efficiency. In contrast, the ALF 184 performs block unit filtering and applies different filters by dividing the degree of the boundaries and variation amounts of the corresponding blocks to compensate for distortion. Information on the filter coefficients to be used for the ALF may be encoded and sent to the video decoding device.

[0078] The restored blocks filtered by the deblocking filter 182, the SAO filter 184, and the ALF 184 are stored in the memory 190. When all the blocks in a picture are restored, the restored picture can be used as a reference picture for inter prediction of blocks within a picture to be encoded later.

[0079] Figure 5 is an example functional block diagram of a video decoding apparatus in which the technology of the present invention can be implemented. Hereinafter, with reference to Figure 5 ,the video decoding apparatus and sub-components of the apparatus will be described.

[0080] The video decoding apparatus may be configured to include an entropy decoding unit 510, a rearrangement unit 515, a dequantization unit 520, an inverse transformation unit 530, a prediction unit 540, an adder 550, a loop filter unit 560, and a memory 570.

[0081] Similar to Figure 1 the video encoding apparatus, each component of the video decoding apparatus may be implemented as hardware or software or implemented as a combination of hardware and software. In addition, the functions of each component may be implemented as software, and a microprocessor may also be implemented to execute the functions of the software corresponding to each component.

[0082] The entropy decoding unit 510 decodes the bitstream generated by the video encoding apparatus to determine the current block to be decoded, and extracts information related to block partitioning, and extracts prediction information and information about the residual signal required to restore the current block.

[0083] The entropy decoding unit 510 determines the size of the CTU by extracting information about the CTU size from the sequence parameter set (SPS) or the picture parameter set (PPS), and divides the picture into CTUs with the determined size. Additionally, the CTU is determined as the highest layer of the tree structure, i.e., the root node, and the partitioning information of the CTU is extracted to partition the CTU using the tree structure.

[0084] For example, when partitioning the CTU using the QTBTTT structure, first, the first flag (QT_split_flag) related to the partitioning of the QT is extracted to divide each node into four lower-layer nodes. Additionally, for the nodes corresponding to the leaf nodes of the QT, the second flag (MTT_split_flag) and the partitioning direction (vertical / horizontal) and / or partitioning type (binary / ternary) related to the partitioning of the MTT are extracted to divide the corresponding leaf nodes into the MTT structure. As a result, each node below the leaf nodes of the QT is recursively divided into the BT or TT structure.

[0085] As another example, when dividing a CTU by using a QTBTTT structure, a CU division flag (split_cu_flag) indicating whether a CU is divided is extracted, and when the corresponding block is divided, a first flag (QT_split_flag) can also be extracted. During the division process, for each node, zero or more recursive MTT divisions can occur after zero or more recursive QT divisions. For example, with respect to a CTU, an MTT division can occur immediately, or conversely, only multiple QT divisions can occur.

[0086] As another example, when dividing a CTU by using a QTBT structure, a first flag (QT_split_flag) related to the division of a QT is extracted to divide each node into four lower-layer nodes. Additionally, a division flag (split_flag) indicating whether the node corresponding to the leaf node of the QT is further divided into a BT and division direction information are extracted.

[0087] Meanwhile, when the entropy decoding unit 510 determines a current block to be decoded by using a tree-structured division, the entropy decoding unit 510 extracts information about a prediction type indicating whether the current block is intra-predicted or inter-predicted. When the prediction type information indicates intra-prediction, the entropy decoding unit 510 extracts a syntax element of the intra-prediction information (intra-prediction mode) of the current block. When the prediction type information indicates inter-prediction, the entropy decoding unit 510 extracts information representing syntax elements for inter-prediction information (i.e., a motion vector and a reference picture to which the motion vector refers).

[0088] In addition, the entropy decoding unit 510 extracts quantization-related information and information about quantized transform coefficients of the current block as information about a residual signal.

[0089] The rearrangement unit 515 can again change a sequence of 1D quantized transform coefficients entropy-decoded by the entropy decoding unit 510 into a 2D coefficient array (i.e., a block) in an order opposite to the coefficient scan order performed by the video coding device.

[0090] The dequantization unit 520 dequantizes the quantized transform coefficients and dequantizes the quantized transform coefficients by using a quantization parameter. The dequantization unit 520 can also apply different quantization coefficients (scaling values) to the quantized transform coefficients arranged in a 2D manner. The dequantization unit 520 can perform dequantization by applying a matrix (scaling value) of quantization coefficients from the video coding device to a 2D array of the quantized transform coefficients.

[0091] The inverse transform unit 530 generates a residual block for the current block by inversely transforming the dequantized transform coefficients from the frequency domain to the spatial domain to recover a residual signal.

[0092] In addition, when the inverse transform unit 530 inverse-transforms a partial region (sub-block) of a transform block, the inverse transform unit 530 extracts a flag (cu_sbt_flag) indicating that only the sub-block of the transform block is transformed, direction (vertical / horizontal) information (cu_sbt_horizontal_flag) of the sub-block, and / or position information (cu_sbt_pos_flag) of the sub-block, and inverse-transforms the transform coefficients of the corresponding sub-block from the frequency domain to the spatial domain to recover the residual signal, and fills the un-inverse-transformed region with "0" values as the residual signal to generate a final residual block for the current block.

[0093] In addition, when applying MTS, the inverse transform unit 530 determines a transform index or transform matrix to be applied in each of the horizontal and vertical directions by using the MTS information (mts_idx) signaled from the video coding device, and performs inverse transformation on the transform coefficients in the transform block in the horizontal and vertical directions by using the determined transform function.

[0094] The prediction unit 540 may include an intra prediction unit 542 and an inter prediction unit 544. The intra prediction unit 542 is activated when the prediction type of the current block is intra prediction, and the inter prediction unit 544 is activated when the prediction type of the current block is inter prediction.

[0095] The intra prediction unit 542 determines the intra prediction mode of the current block among a plurality of intra prediction modes according to the syntax element of the intra prediction mode extracted from the entropy decoding unit 510, and predicts the current block by using the adjacent reference pixels of the current block according to the intra prediction mode.

[0096] The inter prediction unit 544 determines the motion vector of the current block and the reference picture for the motion vector reference by using the syntax element for the inter prediction mode extracted from the entropy decoding unit 510.

[0097] The adder 550 restores the current block by adding the residual block output from the inverse transform unit to the prediction block output from the inter prediction unit or the intra prediction unit. When performing intra prediction on a block to be decoded later, the pixels in the restored current block are used as reference pixels.

[0098] The loop filter unit 560 as a loop filter may include a deblocking filter 562, a SAO filter 564, and an ALF 566. The deblocking filter 562 performs deblocking filtering on the boundaries between the restored blocks to remove block artifacts generated due to block-based unit decoding. The SAO filter 564 and the ALF 566 perform additional filtering on the restored blocks after deblocking filtering to compensate for the difference between the restored pixels and the original pixels generated due to lossy coding. The filter coefficients of the ALF are determined by using the information on the filter coefficients decoded from the bitstream.

[0099] The restored blocks filtered by the deblocking filter 562, the SAO filter 564, and the ALF 566 are stored in the memory 570. When all the blocks in a picture are restored, the restored picture can be used as a reference picture for inter prediction of blocks within a picture to be encoded later.

[0100] This embodiment relates to the image (video) encoding and decoding as described above. More specifically, this embodiment provides a video encoding / decoding method, which performs inter prediction on a current picture by resampling and referring to the luminance signal and chrominance signal of a reference picture, so as to improve the encoding efficiency in the video encoding and decoding of frames with luminance signals and chrominance signals of various sampling formats in a video sequence.

[0101] In the following description according to the present disclosure, it is assumed that the memories 190 and 570 of the video encoding / decoding device as shown in Figure 1 and Figure 5 include a decoded picture buffer (DPB).

[0102] As Figure 1 and Figure 5 illustrate, the video encoding / decoding device maintains the same size for all pictures in a video sequence. Thus, the size information of the picture is indicated in the SPS, which is a high-level syntax defining the encoding parameters of the video sequence. For example, in the SPS, pic_width_in_luma_samples and pic_height_in_luma_samples are used as syntax elements regarding the picture size.

[0103] Figure 6 Shows a video encoding device including a resampling block according to an embodiment of the present disclosure.

[0104] Figure 7 Shows a video decoding device including a resampling block according to an embodiment of the present disclosure.

[0105] Figure 6 and Figure 7 The resampling blocks of the video encoding / decoding device as shown in change the width and height of the pictures included in a video sequence. In this case, applying the same resampling rate to a picture without distinguishing luminance / chrominance is referred to as reference picture resampling (RPR).

[0106] In Figure 6In an example, the input video frame may be downsampled by the downsampler 105. The output video frame obtained by encoding and then decoding a previous input frame is stored in the DPB in the memory 190. Here, when encoding the next frame to be encoded at a sampling rate different from the input size of the previous video frame, the video encoding device may upsample or downsample the frames stored in the DPB so that the frames match the rate using the resampler 195.

[0107] In Figure 7 an example, the reference pictures in the DPB in the memory 570 are stored without being resampled. However, when the sizes of the current picture and the reference picture are different, the reference pictures in the DPB are upsampled or downsampled using the resampler 575 so that the reference pictures have the same size as the current picture, and then the video decoding device may perform decoding. During the motion estimation and motion compensation processes, the motion vectors may also be scaled according to the size ratio and the order difference of the reference pictures.

[0108] To implement this RPR function, the video encoding device transmits the following syntax.

[0109] First, the video encoding device uses the SPS transmission syntax pic_width_max_in_luma_samples and pic_height_max_in_luma_samples. Here, pic_width_max_in_luma_samples represents the maximum width of the frame to be encoded in units of luma samples, and pic_height_max_in_luma_samples represents the maximum height of the frame to be encoded in units of luma samples. Its value must be a non-zero integer multiple of Max(8, MinCbSizeY). MinCbSizeY indicates the minimum size of the luma blocks constituting the video picture.

[0110] In addition, the video encoding device specifies the size of each video frame to be decoded by using the PPS transmission syntax pic_width_in_luma_samples and pic_height_in_luma_samples.

[0111] Furthermore, the video encoding device transmits a syntax called res_change_in_clvs_allowed_flag (or ref_pic_recessing_enabled_flag) on the SPS. When the value of this syntax is 1, it indicates that the spatial resolution of the video picture can be changed, and 0 indicates that the spatial resolution is always fixed. Thus, when res_change_in_clvs_allowed_flag is 0, pic_width_in_luma_samples and pic_height_in_luma_samples can be set to be the same as pic_width_max_in_luma_samples and pic_height_max_in_luma_samples.

[0112] To represent the size of the picture to which RPR is applied, the RPR scaling factors between the reference picture and the current picture can be used. The RPR scaling factors include pps_scaling_win_left_offset, pps_scaling_win_right_offset, pps_scaling_win_top_offset, and pps_scaling_win_bottom_offset.

[0113] When the RPR function is used, the conventional encoding techniques used to encode blocks within a video frame can be changed.

[0114] For example, for a reference picture to which RPR is applied, the video decoding / encoding device can set a limit such that the temporal motion vector predictor (TMVP) is not used. This is because when the sizes and scaling factors of the co-located reference pictures used to extract TMVP are different, the positions of the co-located blocks can be different. In addition, decoder-side motion vector refinement (DMVR) and prediction refinement with optical flow (PROF) may be restricted. When RPR is applied, the video decoding / encoding device can perform motion estimation after applying an additional down / up sampling filter to the reference picture. In this case, down / up sampling filters depending on the inter-frame prediction mode of the block (e.g., when affine prediction is used) can be used.

[0115] Meanwhile, RPR is based on the assumption that all pictures in the sequence have the same chroma format. For example, as in Figure 8As shown, in the case of Frame 1, the downsampling rate of the chrominance sampling is determined according to the downsampling rate of the luminance sampling. Thus, various chrominance formats are not provided for each picture in a video stream. Hereinafter, a resampling method for chrominance signals of video pictures having various chrominance formats (or sampling formats) such as 4:4:4, 4:2:2, and 4:2:0 in a video stream according to the present embodiment will be described.

[0116] Figures 9 to 11 is an exemplary diagram showing a reference frame having various chrominance formats according to an embodiment of the present disclosure.

[0117] As Figure 9 shown, the video coding device downsamples the luminance signal Y with respect to a reference picture to which RPR is applied, but maintains the original resolution of the chrominance signals U and V or resamples the chrominance signals U and V at different rates. When the chrominance signals maintain the original resolution, they have the same sampling rate as the downsampled luminance signal and thus have a 4:4:4 format. When a downsampling rate different from the downsampling rate of the luminance signal is provided to the chrominance signals, that is, when resampling is applied only in the vertical or horizontal direction, a 4:2:2 format can be provided for the chrominance signals.

[0118] As Figure 10 shown, the video coding device can provide a 4:4:4 format to the chrominance signals U and V by upsampling the chrominance signals U and V while keeping the luminance signal Y unchanged. Even in this case, if an upsampling rate different from the upsampling rate of the luminance signal is provided to the chrominance signals, a 4:2:2 format can be provided for the chrominance signals.

[0119] For the method of using RPR in the 4:4:4 format for 4:2:0 video pictures, as Figure 9 and Figure 10 shown, a reference picture in 4:2:0 in a 4:4:4 video picture can be used for RPR, as Figure 11 shown.

[0120] Even when independent resampling of the luminance signal and the chrominance signal is supported, the resolution of the luminance signal of the reference picture must be greater than or equal to the resolution of the chrominance signal.

[0121] To support independent resampling of the luminance signal and the chrominance signal, the video coding device may add the following syntax elements.

[0122] First, the video encoding device uses the SPS transport syntax pic_width_max_in_chroma_samples and pic_height_max_in_chroma_samples. Here, pic_width_max_in_chroma_samples represents the maximum width of the frame to be encoded in terms of chroma samples, and pic_height_max_in_chroma_samples represents the maximum height of the frame to be encoded in terms of chroma samples. Their values must be non-zero and an integer multiple of Max(8, MinCbSizeC). MinCbSizeC represents the minimum size of the chroma blocks that make up the video picture.

[0123] Optionally, the video encoding device can derive pic_width_max_in_chroma_samples and pic_height_max_in_chroma_samples from pic_height_max_in_luma_samples and pic_width_max_in_luma_samples using a scaling value indicating the sampling rate of the chroma signal.

[0124] In addition, the video encoding device specifies the size of each video frame to be decoded by using the PPS transport syntax pic_width_in_chroma_samples and pic_chroma_in_luma_samples.

[0125] Optionally, the video encoding device can derive pic_width_in_chroma_samples and pic_chroma_in_luma_samples from pic_width_in_luma_samples and pic_height_in_luma_samples using a scaling value indicating the sampling rate of the chroma signal.

[0126] The video encoding device can optionally transmit additionally the chroma format for a picture. For example, in the case of a change in image resolution, the video encoding device can transmit chroma_format_idc_for_res_change on the SPS to add the chroma format. Optionally, additional chroma_format_idc_for_res_change can be transmitted on the PPS. Optionally, additional chroma_format_idc_for_res_change can be transmitted on the picture header.

[0127] Furthermore, the video coding device transmits a syntax called res_change_in_clvs_allowed_flag (or ref_pic_recessing_enabled_flag) on the SPS. When the value of this syntax is 1, it indicates that the spatial resolution of the video picture can change, and 0 indicates that the spatial resolution is always fixed. In this case, the res_change_in_clvs_allowed_flag value can be configured to indicate the luma and chroma signals separately. Optionally, a single flag can be used to control both the luma and chroma signals.

[0128] The video coding device can generate the resolution of the chroma signal of a picture from pic_width_in_chroma_samples, pic_chroma_in_luma_samples, the chroma format, and the scaling factor of the luma signal.

[0129] Meanwhile, the video coding device can use the same filter as the filter used for downsampling / upsampling the luma signal to downsample / upsample the chroma signal. For example, the DCT interpolation filter (DCTIF) used for conventional reference picture upsampling can be used, or a Gaussian filter, a Lancoz filter, etc. can be used. In another embodiment of the present disclosure, the video coding device can use different filters to downsample / upsample the chroma signal. For example, a filter with a smaller number of taps than the filter used for the luma signal can be applied to the chroma signal.

[0130] In addition, during the motion estimation and motion compensation processes, the motion vector can also be scaled according to the size ratio of the reference picture.

[0131] When different types of resampling are used for the luma and chroma signals, the conventional coding techniques used to encode blocks in the video picture can also be changed.

[0132] For 4:4:4 video signals, the video coding device can perform encoding after converting from the YUV or RGB format to another color space (such as YCgCo) by applying an adaptive color transform (ACT). For pictures to which RPR is applied, the video coding device can limit the use of this ACT. Details of the ACT will be described later.

[0133] In addition, only when using a dual-tree structure for blocks that independently partition luminance and chrominance signals, the video coding device enables separate resampling of the chrominance signal. Optionally, to apply separate resampling to the luminance and chrominance signals, the video coding device uses a dual-tree structure. In this case, when applying separate resampling to the luminance and chrominance signals, the no_qtbtt_dual_tree_intra_constraint_flag (a syntax related to the use of the dual tree) is always set to 0. In addition, the qtbtt_dual_tree_intra_flag is always set to 1.

[0134] In another embodiment of the present disclosure, when using a luminance / chrominance dual-tree structure, the video coding device may perform resampling on the luminance signal. Optionally, to apply resampling to the luminance signal, the video coding device uses a dual-tree structure.

[0135] In another embodiment of the present disclosure, when using a luminance / chrominance dual-tree structure, the video coding device may perform resampling on the chrominance signal. Optionally, to apply resampling to the chrominance signal, the video coding device uses a dual-tree structure.

[0136] In another embodiment of the present disclosure, when using a single-tree structure, the video coding device adjusts the resampling rate of the chrominance signal according to the rate of the luminance signal.

[0137] In another embodiment of the present disclosure, when the partitioning structure of the I-frame in the current GOP is a dual-tree structure, the video coding device may perform separate resampling on the chrominance signals of all pictures of the reference I-frame or all pictures in the current GOP.

[0138] In another embodiment of the present disclosure, in the case of a single-tree structure where the blocks of the luminance and chrominance signals are not independently partitioned, the video coding device may perform separate resampling on the chrominance signal.

[0139] Meanwhile, with respect to the reference picture to which chrominance signal resampling is applied, the video decoding / encoding device may impose a restriction such that TMVP is not used. This is because when the sizes and scaling factors of the co-located reference pictures used for extracting TMVP are different, the positions of the co-located blocks may be different. Further, DMVR and PROF may also be restricted. When applying chrominance signal resampling, the video decoding / encoding device may perform motion estimation after applying an additional down / up sampling filter to the reference picture. In this case, a down / up sampling filter depending on the inter-frame prediction mode of the block (e.g., when using affine prediction) may be applied.

[0140] In this embodiment, the chrominance signal resampling performed by the video coding device as described above may be equivalently applied to the video decoding device.

[0141] Figure 12 is a schematic flowchart of a video decoding method according to an embodiment of the present invention.

[0142] Figure 12 An example of performs an inter prediction process by a video decoding device with respect to a current block in a current picture including a chrominance signal having a resolution and a chrominance format separated from the resolution and chrominance format of a luminance signal.

[0143] The video decoding device obtains size information and a chrominance format of the current picture (S1200). Here, the size information includes the maximum width and maximum height of the luminance signal (i.e., the luminance signal may have), the width and height of the luminance signal of the current picture, the maximum width and maximum height of the chrominance signal (i.e., the chrominance signal may have), the width and height of the chrominance signal of the current picture, a flag indicating a change in the chrominance format, a flag indicating a change in the resolution, and a scaling factor of the luminance signal.

[0144] The video decoding device generates a resolution for the chrominance signal of the current picture from the size information and the chrominance format of the current picture (S1202). The resolution of the chrominance signal of the current picture can be generated from the width and height of the chrominance signal of the current picture, the chrominance format, the scaling factor of the luminance signal, and the like.

[0145] The video decoding device obtains a decoded residual signal and inter prediction information of the current block (S1204). Here, the inter prediction information includes a reference picture index and a motion vector.

[0146] The video decoding device obtains the resolution and chrominance format of the chrominance signal of a reference picture specified by the reference picture index (S1206).

[0147] The video decoding device checks whether the resolution and chrominance format of the chrominance signal of the current picture match the resolution and chrominance format of the chrominance signal of the reference picture (S1208).

[0148] When the resolution or chrominance format of the chrominance signal of the current picture is different from the resolution or chrominance format of the reference picture, the video decoding device applies resampling for correcting the chrominance signal to the chrominance signal of a reference block included in the reference picture so that the resolution and chrominance format of the chrominance signal of the reference block match the resolution and chrominance format of the chrominance signal of the current block (S1210).

[0149] The video decoding device generates a prediction signal of the current block based on the inter prediction information (S1212).

[0150] When separate resampling is applied to the chrominance signal of the reference picture, the video decoding device may consider the resampling to adjust the motion vector with respect to the chrominance signal of the reference picture.

[0151] The video decoding device generates a reconstructed block by adding a prediction signal and a residual signal (S1214).

[0152] As described above, according to the present embodiment, the encoding efficiency can be improved by providing a video encoding / decoding method for performing inter-frame prediction on a current picture by resampling and referring to the luminance signal and chrominance signal of a reference picture in video encoding and decoding of frames of luminance signals and chrominance signals having various sampling formats in a video sequence.

[0153] Hereinafter, a method for reducing the memory consumption and latency of a video encoding / decoding device will be described.

[0154] In Screen Content Coding (SCC), the video encoding / decoding device adaptively transforms a residual signal from an RGB or YUV color space into a YCgCo space using the ACT as described above. For each Transform Unit (TU), the video encoding / decoding device can adaptively select one of two color spaces using an ACT flag. When the ACT flag is 1, the residual signal is encoded in the YCgCo space, and when the ACT flag is 0, the TU residual signal is encoded in the original color space.

[0155] In the case of a video with a chrominance format sampling rate of 4:4:4, the video encoding / decoding device can use the ACT.

[0156] Figure 13 is a block diagram of a video decoding device using the ACT according to an embodiment of the present disclosure.

[0157] As Figure 13 shown, the video decoding device can perform a color space transformation in the residual signal area. For example, in order to transform the residual signal in the YCgCo space back to the original chrominance space after the inverse transformation, an inverse ACT unit 535 is used as an additional decoding module.

[0158] The video encoding / decoding device uses one CU as a unit for transform processing unless the maximum transform size is smaller than the width or height of one CU. Thus, when using the ACT flag, a flag for selecting a color space for one CU can be specified. Since the residual signal is additionally transformed, the video encoding / decoding device can use the ACT when there is at least one non-zero transform coefficient for a CU encoded by inter-frame prediction and IBC (Intra Block Copy). In addition, the video encoding / decoding device can use the ACT only when the same prediction mode (i.e., DM mode) is selected for the chrominance signal and luminance signal regarding an intra-frame prediction CU.

[0159] The video encoding / decoding apparatus may use forward and backward YCgCo color transformation matrices as transformation matrices for color space transformation. Additionally, an adjusted QP may be applied to the transformed residual signal to compensate for changes in the dynamic range of the residual signal before and after color transformation.

[0160] ACT uses all three color components of the residual signal during the forward / backward color transformation process. Therefore, the video encoding / decoding apparatus does not use ACT in the following two cases where three color components cannot be used.

[0161] First, when the luminance and chrominance color elements are encoded into separate tree structures, i.e., when the luminance and chrominance samples in a CTU are partitioned into different structures, and thus the CUs of the luminance tree include only the luminance component and the CUs of the chrominance tree include only two chrominance components, the video encoding / decoding apparatus does not use ACT.

[0162] Second, when intra-subdivision prediction (ISP) is applied to the luminance component, the video encoding / decoding apparatus does not use ACT. Here, ISP refers to a technique of dividing a CU into two or four sub-rectangles in the horizontal or vertical direction according to the size of a block during intra prediction.

[0163] Figure 14 is a conceptual exemplary diagram illustrating a video decoding process using ACT according to an embodiment of the present disclosure.

[0164] Since inverse ACT requires all three components, it requires a memory for storing values intermediate between the inverse transformation and inverse ACT. Since the video encoding / decoding apparatus supports a transformation of up to 64 samples, backward ACT requires up to 64×64×3 color elements.

[0165] Figure 15 is a conceptual exemplary diagram illustrating the MRL according to an embodiment of the present disclosure.

[0166] During the intra prediction process, the video encoding / decoding apparatus may use more reference lines by using multiple reference lines (MRL). When MRL is applied, the video encoding / decoding apparatus uses the samples of two lines (reference line 1 and reference line 3 in the example of Figure 15 ) added to the upper side and the left side to perform intra prediction on a block-by-block basis. To select a reference line when MRL is applied, an index (mrl_idx) indicating the reference line may be signaled. When a non-zero reference line index is signaled, the planar and DC modes are excluded from the intra prediction mode.

[0167] Hereinafter, a method for reducing the memory consumption and delay required for video encoding / decoding processing when ACT is applied to an intra prediction block will be described. By Figure 16The region indicated by the YUV color space referred to in the example represents the reference samples in reference line 0 (mrl_idx = 0) for performing intra prediction of the current block. Additionally, the region above the region and the region to the left of the region indicate the reference samples of the region having non-zero reference lines.

[0168] The reference samples are expressed in the original YUV color space. However, when ACT is applied to the current block, the residual signal in the block is expressed in the YCgCo color space. Therefore, in order to reconstruct the original signal, a process of inverse-transforming the color space to the original YUV color space and then adding the reference sampling expressed in the YUV color space to reconstruct the original signal is required, as Figure 14 shown. When adding the reference samples, the video encoding / decoding device may use a clipping operation to limit the pixel values to sample values in the range from 0 to 255.

[0169] As Figure 14 shown, two or more transform processes (or inverse transform processes) are required before performing the addition operation on the reference samples, and thus memory consumption and latency may occur. In this case, the reference samples are pre-restored and stored in a memory buffer, and thus can be used last in the entire pipeline.

[0170] Figure 17 is an exemplary diagram of a pipeline for adding reference samples during intra prediction in the YCgCo space according to an embodiment of the present disclosure.

[0171] Figure 17 The example of

[0172] shows a method in which the reference samples to which ACT is applied are stored in a buffer, then added to the residual signal of the inverse transform in the YCgCo space, and finally the residual signal is restored using inverse ACT. Although this method additionally requires ACT for the reference samples, it has the following advantages: it can be performed independently / parallelly / with the conventional inverse ACT.

[0173] Regarding the methods shown in Figure 14 and Figure 17 if MRL is applied during intra prediction, the memory capacity required by the video encoding / decoding device will increase. Therefore, in intra prediction of the block to which ACT is applied, the video encoding / decoding device only uses the first line buffer of MRL (the reference line closest to the block, i.e., mrl_idx = 0).

[0174] In this embodiment, the video encoding / decoding device can determine whether to apply ACT and MRL by setting sps_act_enabled_flag and sps_mrl_enabled_flag on the SPS. When using ACT by setting sps_act_enabled_flag to 1, the video encoding / decoding device sets sps_mrl_enabled_flag to 0 to use only the first row buffer of MRL. Optionally, when sps_mrl_enabled_flag is set to 0, the video encoding / decoding device uses sps_act_enabled_flag. Optionally, when sps_mrl_enabled_flag is 1, the video encoding / decoding device sets sps_act_enabled_flag to 0.

[0175] In another embodiment of the present disclosure, instead of using a preset YCgCo transform matrix for ACT, the video encoding / decoding device can adaptively derive the coefficients of the transform matrix using the boundary pixels on the upper edge and left edge of the current CU, and then apply the coefficients to the current CU. Here, as the boundary samples for deriving the coefficients of the transform matrix, only the samples of the buffer for MRL can be used. For example, the video encoding / decoding device can use the samples in the first, second, and fourth rows located on the upper side and the left side or some of the samples.

[0176] In this embodiment, when the size of the transform kernel is greater than 16×16, the video encoding / decoding device performs the transform after applying zeroing to the luminance sampling. Here, zeroing refers to a method of replacing all the transform coefficients of the sub-blocks except the upper left sub-block with 0.

[0177] Even when transforming the chrominance block, the video encoding / decoding device can set the zeroing area. That is, the video encoding / decoding device can set the transform coefficients of the sub-blocks except the upper left block to 0, and then perform the transform on the chrominance block.

[0178] The zeroing area can be set differently according to the sampling format of the chrominance signal. For example, when the chrominance sample format of the current picture is 4:4:4, the video encoding / decoding device applies zero values to the blocks except the upper left block and then transforms the chrominance signal. At the same time, in the cases of 4:2:2 and 4:2:0 formats, zeroing may not be applied.

[0179] In another embodiment of the present invention, when the chrominance sampling format of the current picture is 4:4:4, 4:2:2, and 4:2:0, the video encoding / decoding device can apply zeroing to the blocks except the upper left block and then transform the chrominance signal.

[0180] In another embodiment of the present disclosure, the video encoding / decoding device may determine the zero-out region of a chrominance sampling block as a region obtained by reducing the zero-out region of the luminance signal by half in both the horizontal and vertical directions.

[0181] In another embodiment of the present disclosure, when the chrominance sample format is 4:2:0 and 4:2:2, the video encoding / decoding device may determine the zero-out region of a luminance block as a region obtained by reducing the zero-out region of the luminance signal by half in both the horizontal and vertical directions.

[0182] Zero-out can be applied to compensate for a considerable increase in the amount of computation as the size of the transform kernel increases. When zero-out is applied, the complexity is reduced, but there is some loss in terms of compression efficiency. Specifically, when zero-out is forcibly applied to an image with a lower QP (i.e., a higher bitrate), the complexity is lower, but the data loss rate increases, and the usage frequency of a 32×32 or larger transform kernel is significantly reduced, while the usage frequency of a transform kernel with a size smaller than 32 increases. Therefore, the video encoding / decoding device may not utilize the gain in decoding efficiency that can be obtained when using a large transform kernel.

[0183] Hereinafter, to solve such a problem, a method for controlling zero-out is proposed.

[0184] In the present embodiment, the video encoding / decoding device may control zero-out by adaptively sending use_tr_zero_out_flag in units of SPS, PPS, picture header, slice header, CU, or TU.

[0185] In another embodiment of the present disclosure, no_tr_zero_out_constraint_flag may be added to the common constraint information syntax. When this flag is 1, the video encoding / decoding device may restrict the use of zero-out.

[0186] In another embodiment of the present disclosure, zero-out may be controlled according to the profile / level. Since the level is a general measurement indicating the performance level of the video encoding / decoding device, zero-out can be used only below a specific level. For example, the video encoding / decoding device may set a level limit such that zero-out can be used only at level 3 or below.

[0187] In the present embodiment, when applying low-frequency non-separable transform (LFNST), the video encoding / decoding device sets the zero-out regions of luminance and chrominance samples to be the same. Here, LFNST is applied between the transformer 140 and the quantizer 145 in the case of the video encoding device, and is applied between the inverse quantizer 520 and the inverse transformer 530 in the case of the video decoding device to reduce the amount of computation.

[0188] In another embodiment of the present disclosure, when the LFNST is applied to chrominance samples, the video encoding / decoding device may set the LFNST zeroing region of the chrominance samples to be different from the LFNST zeroing region of the luminance samples. For example, the LFNST zeroing region of the chrominance samples may be set to half of the horizontal and vertical lengths of the LFNST zeroing region of the luminance samples.

[0189] In addition, the zeroing region may be determined according to the sampling format of the chrominance signal. When the sampling format of the chrominance signal in the current picture is 4:4:4, the video encoding / decoding device sets the same zeroing region as the luminance signal as the zeroing region of the chrominance signal. When the sampling format is 4:2:0, a different zeroing region may be set. For example, the zeroing region of the chrominance samples may be set to half of the horizontal and vertical lengths of the zeroing region of the luminance samples.

[0190] In another embodiment of the present disclosure, when the sampling format of the chrominance signal is 4:2:0, the video encoding / decoding device may set the same zeroing region as the zeroing region of the luminance signal as the zeroing region of the chrominance signal, and when the sampling format of the chrominance signal is 4:4:4, set twice the horizontal and vertical lengths of the zeroing region of the luminance samples as the zeroing region of the chrominance signal.

[0191] Meanwhile, when the video decoding device is implemented as hardware (H / W), the pipeline may be configured such that parallel processing of each encoding technique can be performed. At this time, the size of the transform block determines the maximum block size of the pipeline, which may become the bottleneck of the entire pipeline. This is because other hardware blocks constituting the pipeline can be designed by dividing the block into arbitrary small blocks, but it is difficult to apply such a division method to the transform block. Since the video encoding / decoding device uses a transform of a maximum of 64×64 blocks, the hardware video decoding device may need to have a pipeline with a minimum size of 64×64 blocks.

[0192] Using the virtual pipeline data unit (VPDU), when the blocks divided within a 64×64 block are completely within one VPDU or span other VPDUs, the video encoding / decoding device follows the limitation that the corresponding block must completely occupy the VPDU. Therefore, in terms of reducing the delay and memory consumption in the entire pipeline, it is important to reduce the delay required for block transformation.

[0193] Hereinafter, a method of decoding an encoded bitstream (hereinafter referred to as a "single-core bitstream") that does not apply parallel decoding (such as multi-slice, multi-tile, and wavefront) using a plurality of cores with low latency will be described.

[0194] Figure 18 is a schematic exemplary diagram showing a decoding process according to an embodiment of the present disclosure.

[0195] Figure 18 An example of

[0195] schematically shows a decoding process by which a single reconstructed image is generated from a bitstream. Here, entropy decoding is the step of parsing the bitstream performed by the entropy decoder 510, IQ represents inverse quantization performed by the inverse quantizer 515, IT represents inverse transformation performed by the inverse transformer 530, MC represents motion compensation, IP represents intra prediction, and loop filtering indicates the deblocking filter 562 / ALF 564 / SAO filter 566. Additionally, Et, Tt, Pt, and Lt shown for each step are the times required for each step and generally have the relationship Pt to Lt > Et > Tt.

[0196] Even when using multiple cores, the bitstream for which parallel decoding is not applied must be decoded sequentially as shown in the above figure. Therefore, the time taken to reconstruct one image is Et + Tt + Pt + Lt.

[0197] Hereinafter, reference will be made to Figure 19 Describe the process of applying multiple cores to the decoding of a single-core bitstream.

[0198] Figure 19 is an exemplary diagram showing the process of decoding a single-core bitstream using multiple cores according to an embodiment of the present disclosure.

[0199] In this embodiment, the video decoding device can apply multiple cores to the decoding of a single-core bitstream (loop filtering, MC / IP, entropy decoding, and IT / IQ) based on the time required for each decoding module. That is, considering the decoding time and the dependencies between decoding modules, the decoding modules can be assigned to each core. For example, entropy decoding / IT / IQ / IP have mutual dependencies, and in-loop filtering needs to be terminated before performing inter prediction on the next image, and thus in-loop filtering can be regarded as having relatively low dependencies. Additionally, when entropy decoding is completed, MC can be performed at any time. As described above, the decoding time of each decoding module is in the order of Pt to Lt > Et > Tt. Considering all these factors, the entropy decoding, IT / IQ, and IP of the Nth frame are assigned to core 0. Additionally, while performing the corresponding processing, the in-loop filtering of the (N - 1)th frame and the MC of the Nth frame are assigned to core 1. Since the decoding modules are assigned to each core in this way, the decoding time of the single-core bitstream can be significantly reduced to approximately Mt, as Figure 19 shown.

[0200] Figure 20 is an exemplary diagram showing the operation of a hardware video decoding device according to an embodiment of the present disclosure.

[0201] The video decoding device can perform parallel or serial decoding according to whether parallel processing is applied or whether multiple cores are supported.

[0202] The video decoding device determines whether to apply parallel processing to the bitstream (S2000).

[0203] If parallel processing is applied, the video decoding device determines whether multiple cores are supported (S2002). If multiple cores are supported, the video decoding device performs normal parallel decoding processing (S2004). If multiple cores are not supported, the video decoding device performs normal serial decoding processing (S2006).

[0204] Even when parallel processing is not applied, the video decoding device determines whether multiple cores are supported (S2008). If multiple cores are not supported, the video decoding device performs normal serial decoding processing (S2006).

[0205] If multiple cores are supported, the video decoding device allocates decoding modules to each core taking into account the decoding time and the dependencies between the decoding modules (S2010), and then performs the proposed parallel decoding processing as shown in Figure 19 and performs the proposed parallel decoding processing (S2012).

[0206] In this embodiment, the method for reducing memory consumption and latency performed by the video decoding device as described above can be equally applied to a video encoding device.

[0207] As described above, according to this embodiment, by providing a video encoding / decoding method for reducing memory consumption and latency during the encoding / decoding process, the bitrate of various contents such as game broadcasts, 360-degree video streams, VR / AR videos, and online lectures can be reduced, the burden on the network and energy consumption of a reproduction device that performs video decoding can be alleviated, and fast decoding can be achieved.

[0208] In each flowchart according to the embodiment, it is described that each process is executed sequentially, but the present disclosure is not limited thereto. In other words, since the processes described in the flowchart can be changed and executed or one or more processes can be executed in parallel, the flowchart is not limited to the time series order.

[0209] Meanwhile, various functions or methods described in the present disclosure can also be implemented by instructions stored in a non-transitory recording medium that can be read and executed by one or more processors. For example, non-volatile recording media include all types of recording devices that store data in a form readable by a computer system. For example, non-transitory recording media include storage media such as erasable programmable read-only memories (EPROMs), flash drives, optical drives, magnetic hard disk drives, and solid state drives (SSDs).

[0210] Although the exemplary embodiments of the present disclosure have been described for illustrative purposes, those skilled in the art will recognize that various modifications, additions, and substitutions are possible without departing from the spirit and scope of the claimed invention. Therefore, for the sake of brevity and clarity, the exemplary embodiments of the present disclosure have been described. The scope of the technical concept of this embodiment is not limited by the illustrations. Thus, those of ordinary skill in the art will understand that the scope of the claimed invention is not limited by the embodiments explicitly described above, but rather by the claims and their equivalents.

[0211] (Reference numerals)

[0212] 105: Downsampler 120: Predictor

[0213] 140: Transformer 145: Quantizer

[0214] 195: Resampler

[0215] 520: Inverse quantizer 530: Inverse transformer

[0216] 535: Inverse ACT unit 540: Predictor

[0217] 575: Resampler.

Claims

1. A method for decoding a video sequence of pictures to allow the pictures in the video sequence to have different chrominance formats, the method comprises: decoding chrominance format information of the video sequence and a flag indicating whether the spatial resolution of the pictures in the video sequence can be changed from a bitstream; decoding at least one first syntax element from the bitstream, the at least one first syntax element being used to specify the size of a luma picture including luma samples of a current picture belonging to the video sequence, thereby determining the size of the luma picture; and determining the size of a chroma picture including chroma samples of the current picture based on the flag, wherein when the flag indicates that the spatial resolution of the pictures in the video sequence can be changed, the size of the chroma picture is set based on at least one second syntax element included in the bitstream for specifying the size of the chroma picture, and wherein when the flag indicates that the spatial resolution of the pictures in the video sequence cannot be changed, the size of the chroma picture is set to be equal to the size defined by the size of the luma picture and the chrominance format information of the video sequence.

2. The method according to claim 1, wherein the at least one second syntax element includes additional chrominance format information to be applied to the current picture, and wherein the size of the chroma picture is set to be equal to the size defined by the size of the luma picture and the additional chrominance format information.

3. The method according to claim 1, wherein the at least one second syntax element includes: a syntax element for specifying the width of the chroma picture, and a syntax element for specifying the height of the chroma picture.

4. The method according to claim 1, wherein wherein when the flag indicates that the spatial resolution of the pictures in the video sequence can be changed, at least one of an Adaptive Color Transform (ACT), a Temporal Motion Vector Predictor (TMVP), a Decoder-side Motion Vector Refinement (DMVR), and a Prediction Refinement with Optical Flow (PROF) encoding tool is disabled.

5. A method for encoding a video sequence of pictures to allow the pictures in the video sequence to have different chrominance formats, the method comprises: encoding chrominance format information of the video sequence and a flag indicating whether the spatial resolution of the pictures in the video sequence can be changed into a bitstream; determining the size of a luma picture including luma samples of a current picture belonging to the video sequence, and encoding at least one first syntax element for specifying the dimensions of the luma picture into the bitstream; and based on the flag, determining the size of a chroma picture including chroma samples of the current picture, the chroma picture being encoded with at least one second syntax element for specifying the size of the chroma picture, wherein when the flag indicates that the spatial resolution of the pictures in the video sequence can be changed, the at least one second syntax element is encoded into the bitstream for specifying the size of the chroma picture, Among them, when the flag indicates that the spatial resolution of the picture in the video sequence cannot be changed, the at least one second syntax element is not encoded, and the size of the chrominance picture is set to be equal to the size defined by the size of the luma picture and the chrominance format information of the video sequence.

6. A method for providing encoded data of a video sequence of pictures to a video decoding device, the method comprising: generating a bitstream by encoding a video sequence of pictures to allow pictures in the video sequence to have different chrominance formats; and sending the bitstream to the video decoding device, wherein generating the bitstream comprises: encoding the chrominance format information of the video sequence and a flag indicating whether the spatial resolution of the pictures in the video sequence can be changed into the bitstream; determining the size of the luma picture including the luma samples of the current picture belonging to the video sequence, and encoding at least one first syntax element for specifying the size of the luma picture into the bitstream; and based on the flag, determining the size of the chrominance picture including the chrominance samples of the current picture, and encoding the chrominance picture with at least one second syntax element for specifying the size of the chrominance picture, wherein when the flag indicates that the spatial resolution of the pictures in the video sequence can be changed, encoding the at least one second syntax element into the bitstream for specifying the size of the chrominance picture, wherein when the flag indicates that the spatial resolution of the pictures in the video sequence cannot be changed, the at least one second syntax element is not encoded, and the size of the chrominance picture is set to be equal to the size defined by the size of the luma picture and the chrominance format information of the video sequence.

Citation Information

Patent Citations

  • Manufacturing method of back cover for mobile communication device

    KR1020200018864A

  • Method and apparatus for encoding / decoding multilayer video signal

    CN105850126A

  • Apparatus and method for coding and decoding of correlation between chrominance video and luminance video

    KR1020120008228A