Method and apparatus for processing video signal
By processing video signals through a partitioned transformation unit, the problems of insufficient coding efficiency and intra-frame prediction mode selection in existing technologies are solved, achieving more efficient video signal coding and prediction mode reception.
Patent Information
- Application Number
- CN202511242298.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-19
- Filing Date
- 2020-07-20
- Publication Date
- 2025-11-14
AI Technical Summary
Existing video signal processing methods have shortcomings in coding efficiency and intra-frame prediction mode selection, and it is necessary to improve coding efficiency and effectively receive the prediction mode selected by the encoder.
The video signal is processed by using partitioned transform units (blocks). The current transform block is divided into multiple transform blocks according to preset conditions, and the video signal is decoded through multiple transform blocks. The preset conditions include the comparison results of color components, width and height, to determine the partitioning direction, and the prediction mode selected by the encoder is received in the intra-frame prediction method.
It improves the coding efficiency of video signals and effectively receives the prediction mode selected by the encoder in the intra-frame prediction method, thereby enhancing the effect of video signal processing.
Smart Images

Figure CN120956894A_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 202080052141.1 (International Application No. PCT / KR2020 / 009563), filed on January 18, 2022, with an international application date of July 20, 2020, entitled "Method and Apparatus for Processing Video Signals". Technical Field
[0002] This disclosure relates to a video signal processing method and apparatus, and more specifically, to a video signal processing method and apparatus for encoding or decoding video signals. Background Technology
[0003] Compression coding refers to a series of signal processing techniques used to transmit digitized information over communication lines or to store information in a form suitable for storage media. The objects of compression coding include objects such as voice, video, and text, and in particular, techniques used to perform compression coding on images are called video compression. Compression coding of video signals is performed by removing excess information, taking into account spatial, temporal, and random correlations. However, with the recent development of various media and data transmission media, there is a need for more efficient video signal processing methods and devices. Summary of the Invention
[0004] Technical issues
[0005] The purpose of this invention is to increase the encoding efficiency of video signals.
[0006] The purpose of this invention is to improve the coding efficiency of video signals by using partition transformation units (blocks).
[0007] The purpose of this invention is to efficiently receive the prediction mode selected by the encoder in intra-frame prediction methods.
[0008] Technical solution
[0009] This specification provides a method for processing video signals using a quadratic transform.
[0010] Specifically, a video signal decoding device includes a processor, wherein the processor is configured to: determine a result value indicating the partitioning direction of a current transform block (TB) based on preset conditions, divide the current transform block into a plurality of transform blocks based on the result value, and decode a video signal by using the plurality of transform blocks, wherein the preset conditions include conditions related to the color components of the current transform block.
[0011] Additionally, in this specification, the preset conditions include conditions related to the result of comparing the width of the current transform block with the maximum transform block width, and determining the maximum transform block width based on the chroma format associated with the current transform block, the color components of the current transform block, and the maximum transform size.
[0012] Additionally, in this specification, the preset conditions further include conditions related to the result of comparing a first width value and a first height value, wherein the first width value is a value obtained by multiplying the width of the current transform block by a first value, and the first height value is a value obtained by multiplying the height of the current transform block by a second value, wherein the first value and the second value are respectively values related to the width and the height of the current transform block, and if the color component of the current transform block is luminance, then the first value and the second value are respectively set to 1, and if the color component of the current transform block is chroma, then the first value and the second value are respectively determined based on the chroma format associated with the current transform block.
[0013] Additionally, in this specification, when the width of the current transform block is greater than the width of the maximum transform block and the first width value is greater than the first height value: the result value is determined to be 1, which is a value indicating that the partitioning direction is vertical; the width of each of the plurality of transform blocks is a value obtained by dividing the width of the transform block by 2; and the height of each of the plurality of transform blocks is the same as the height of the transform block.
[0014] Additionally, in this specification, when the width of the current transform block is less than or equal to the width of the maximum transform block, or when the first width value is less than or equal to the first height value: the result value is determined to be 0, which is a value indicating that the partition direction is horizontal, the width of each of the plurality of transform blocks is the same as the width of the transform block, and the height of each of the plurality of transform blocks is a value obtained by dividing the height of the transform block by 2.
[0015] Additionally, in this specification, the maximum transform size is determined based on the size of the coding tree block (CTB) with a luminance component included in the coding tree unit (CTU) associated with the current transform block.
[0016] Additionally, in this specification, when the size of the coding tree block is 32, the maximum transformation size is 32.
[0017] Additionally, in this specification, if the color component of the current transform block is chroma, the processor parses a syntax element indicating whether the prediction method of the coding block associated with the current transform block is Block-Based Incremental Pulse Code Modulation (BDPCM). When, as a result of the parsing, the prediction method of the coding block is not BDPCM, the processor additionally parses the syntax element associated with the prediction method of the coding block, and determines the prediction method of the coding block based on the parsing result. Furthermore, the syntax element associated with the prediction method of the coding block is a syntax element indicating at least one of Cross-Component Linear Model (CCLM), Planar Mode, DC Mode, Vertical Mode, Horizontal Mode, Diagonal Mode, and DM Mode.
[0018] Additionally, a video signal encoding device includes a processor, wherein the processor determines a result value indicating the partitioning direction of a current transform block (TB) based on preset conditions, divides the current transform block into a plurality of transform blocks based on the result value, and generates a bit stream including information about the plurality of transform blocks, and the preset conditions include conditions related to the color components of the current transform block.
[0019] Additionally, in this specification, the preset conditions include conditions related to the result of comparing the width of the current transform block with the maximum transform block width, and determining the maximum transform block width based on the chroma format associated with the current transform block, the color components of the current transform block, and the maximum transform size.
[0020] Additionally, in this specification, the preset conditions further include conditions related to the result of comparing a first width value and a first height value, wherein the first width value is a value obtained by multiplying the width of the current transform block by a first value, and the first height value is a value obtained by multiplying the height of the current transform block by a second value, wherein the first value and the second value are respectively values related to the width and the height of the current transform block, and if the color component of the current transform block is luminance, then the first value and the second value are respectively set to 1, and if the color component of the current transform block is chroma, then the first value and the second value are respectively determined based on the chroma format associated with the current transform block.
[0021] Additionally, in this specification, when the width of the current transform block is greater than the width of the maximum transform block and the first width value is greater than the first height value: the result value is determined to be 1, which is a value indicating that the partitioning direction is vertical; the width of each of the plurality of transform blocks is a value obtained by dividing the width of the transform block by 2; and the height of each of the plurality of transform blocks is the same as the height of the transform block.
[0022] Additionally, in this specification, when the width of the current transform block is less than or equal to the width of the maximum transform block, or when the first width value is less than or equal to the first height value: the result value is determined to be 0, which is a value indicating that the partition direction is horizontal, the width of each of the plurality of transform blocks is the same as the width of the transform block, and the height of each of the plurality of transform blocks is a value obtained by dividing the height of the transform block by 2.
[0023] Additionally, in this specification, the maximum transform size is determined based on the size of the coding tree block (CTB) with a luminance component included in the coding tree unit (CTU) associated with the current transform block.
[0024] Additionally, in this specification, when the size of the coding tree block is 32, the maximum transformation size is 32.
[0025] Additionally, in this specification, if the color component of the current transform block is chroma, the processor parses a syntax element indicating whether the prediction method of the coding block associated with the current transform block is Block-Based Incremental Pulse Code Modulation (BDPCM). When, as a result of the parsing, the prediction method of the coding block is not BDPCM, the processor additionally parses the syntax element associated with the prediction method of the coding block, and determines the prediction method of the coding block based on the parsing result. The syntax element associated with the prediction method of the coding block is a syntax element indicating at least one of Cross-Component Linear Model (CCLM), Planar Mode, DC Mode, Vertical Mode, Horizontal Mode, Diagonal Mode, and DM Mode.
[0026] Additionally, a non-transitory computer-readable medium for storing a bitstream encoded by an encoding method comprising the steps of: determining a result value indicating the partitioning direction of a current transform block (TB) based on preset conditions; partitioning the current transform block into a plurality of transform blocks based on the result value; and encoding a bitstream including information about the plurality of transform blocks, wherein the preset conditions include conditions related to the color components of the current transform block.
[0027] Additionally, in this specification, the preset conditions include conditions related to the result of comparing the width of the current transform block with the maximum transform block width, and determining the maximum transform block width based on the chroma format associated with the current transform block, the color components of the current transform block, and the maximum transform size.
[0028] Additionally, in this specification, the preset conditions further include conditions related to the result of comparing a first width value and a first height value, wherein the first width value is a value obtained by multiplying the width of the current transform block by a first value, and the first height value is a value obtained by multiplying the height of the current transform block by a second value, wherein the first value and the second value are respectively values related to the width and the height of the current transform block, and if the color component of the current transform block is luminance, then the first value and the second value are respectively set to 1, and if the color component of the current transform block is chroma, then the first value and the second value are respectively determined based on the chroma format associated with the current transform block.
[0029] Additionally, in this specification, the maximum transform size is determined based on the size of the coding tree block (CTB) containing the luminance component in the coding tree unit (CTU) associated with the current transform block.
[0030] Beneficial effects
[0031] Embodiments of the present invention provide a method and apparatus for processing video signals using partitioning of a transformation unit.
[0032] Embodiments of the present invention provide a method and apparatus for processing video signals in an intra-frame prediction method to efficiently receive a prediction mode selected by an encoder. Attached Figure Description
[0033] Figure 1 This is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention.
[0034] Figure 2 This is a schematic block diagram of a video signal decoding apparatus according to an embodiment of the present invention.
[0035] Figure 3 An example is shown in which the coding tree unit is divided into coding units in the image.
[0036] Figure 4 An embodiment of a method for sending partitions of quadtrees and multi-type trees using signals is shown.
[0037] Figure 5 and Figure 6 The intra-frame prediction method is illustrated in more detail according to embodiments of the present disclosure.
[0038] Figure 7 This is a diagram illustrating an inter-frame prediction method according to an embodiment of the present invention.
[0039] Figure 8 This is a diagram illustrating a method for transmitting the motion vector of the current block using a signal according to an embodiment of the present invention.
[0040] Figure 9 This is a diagram illustrating a method for transmitting the motion vector difference of the current block using a signal according to an embodiment of the present invention.
[0041] Figure 10 This is a diagram illustrating the encoding unit and the transformation unit according to an embodiment of the present invention.
[0042] Figure 11 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0043] Figure 12 This is a diagram illustrating the decoding process in an intra-block according to an embodiment of the present invention.
[0044] Figure 13 This is a diagram illustrating the decoding process of a residual signal according to an embodiment of the present invention.
[0045] Figure 14 This is a diagram illustrating the relationship between color components according to an embodiment of the present invention.
[0046] Figure 15 This is a diagram illustrating the relationship between color components according to an embodiment of the present invention.
[0047] Figure 16 This is a diagram illustrating the maximum transformation size according to an embodiment of the present invention.
[0048] Figure 17 This is a diagram illustrating a higher level of syntax according to an embodiment of the present invention.
[0049] Figure 18 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0050] Figure 19 This is a diagram illustrating the TU partition according to an embodiment of the present invention.
[0051] Figure 20 This is a diagram illustrating the TU partition according to an embodiment of the present invention.
[0052] Figure 21 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0053] Figure 22 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0054] Figure 23 This is a diagram illustrating a higher-level syntax according to an embodiment of the present invention.
[0055] Figure 24 This is a diagram illustrating a higher-level syntax according to an embodiment of the present invention.
[0056] Figure 25 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0057] Figure 26 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0058] Figure 27 This is a diagram illustrating a method for performing BDPCM according to an embodiment of the present invention.
[0059] Figure 28 This is a diagram illustrating the syntax related to BDPCM according to an embodiment of the present invention.
[0060] Figure 29 This is a diagram illustrating the conditions under which BDPCM can be used according to an embodiment of the present invention.
[0061] Figure 30 This is a diagram illustrating CIIP and intra-frame prediction according to an embodiment of the present invention.
[0062] Figure 31 This is a diagram illustrating the merged data syntax according to an embodiment of the present invention.
[0063] Figure 32 This is a diagram illustrating the merged data syntax according to an embodiment of the present invention.
[0064] Figure 33 This is a diagram illustrating a method for executing CIIP mode according to an embodiment of the present invention.
[0065] Figure 34 This is a diagram illustrating the chroma BDPCM syntax structure according to an embodiment of the present invention.
[0066] Figure 35 This is a diagram illustrating the chroma BDPCM syntax structure according to an embodiment of the present invention.
[0067] Figure 36 This is a diagram illustrating a higher-level syntax related to BDPCM according to an embodiment of the present invention.
[0068] Figure 37 This is a diagram illustrating the syntax elements of BDPCM transmitted by signals at a higher level according to an embodiment of the present invention.
[0069] Figure 38 This is a diagram illustrating the syntax related to chroma BDPCM according to an embodiment of the present invention.
[0070] Figure 39 This is a diagram illustrating the syntax related to intra-frame prediction according to an embodiment of the present invention.
[0071] Figure 40 This is a diagram illustrating the syntax related to intra-frame prediction according to an embodiment of the present invention.
[0072] Figure 41 This is a diagram illustrating the sequence parameter set syntax according to an embodiment of the present invention.
[0073] Figure 42 This is a diagram illustrating the syntax elements related to sub-images according to an embodiment of the present invention.
[0074] Figure 43 This is a diagram illustrating operators according to an embodiment of the present invention.
[0075] Figure 44 These are pictures and sub-pictures illustrating embodiments of the present invention.
[0076] Figure 45 This is a diagram illustrating the syntax elements related to sub-images according to an embodiment of the present invention.
[0077] Figure 46 This is a diagram illustrating a method for partition transformation blocks according to an embodiment of the present invention. Detailed Implementation
[0078] Considering the functions of this disclosure, the terminology used in this specification may be general terms that are currently widely used, but may change according to the intent, customs, or emergence of new technologies of those skilled in the art. Additionally, in some cases, terms may be arbitrarily chosen by the applicant, and in such cases, their meanings are described in the corresponding descriptive sections of this disclosure. Therefore, the terminology used in this specification should be interpreted based on its substantive meaning throughout the specification.
[0079] In this specification, some terms may be interpreted as follows. In some cases, coding may be interpreted as encoding or decoding. In this specification, an apparatus for generating a video signal bitstream by performing encoding of a video signal is called an encoding apparatus or encoder, and an apparatus for reconstructing a video signal by performing decoding of the video signal bitstream is called a decoding apparatus or decoder. Additionally, in this specification, video signal processing apparatus is used as a term encompassing both the concepts of encoder and decoder. Information is a term that includes all values, parameters, coefficients, elements, etc. In some cases, the meaning is interpreted differently, therefore this disclosure is not limited thereto. "Unit" is used to refer to a basic unit of image processing or a specific location in an image, and refers to an image region that includes both luminance and chrominance components. Additionally, "block" refers to an image region that includes a specific component among the luminance and chrominance components (i.e., Cb and Cr). However, depending on the embodiment, terms such as "unit," "block," "partition," and "region" may be used interchangeably. Additionally, in this specification, "unit" can be used as a general term encompassing coding units, prediction units, and transformation units. Images refer to fields or frames, and these terms may be used interchangeably depending on the embodiment.
[0080] Figure 1 This is a schematic block diagram of a video signal encoding apparatus according to an embodiment of the present invention. (Reference) Figure 1 The encoding device 100 of the present invention includes a transformation unit 110, a quantization unit 115, an inverse quantization unit 120, an inverse transformation unit 125, a filtering unit 130, a prediction unit 150, and an entropy encoding unit 160.
[0081] Transform unit 110 obtains the values of transform coefficients by transforming the residual signal, which is the difference between the input video signal and the prediction signal generated by prediction unit 150. For example, discrete cosine transform (DCT), discrete sine transform (DST), or wavelet transform can be used. DCT and DST perform the transform by dividing the input image signal into multiple blocks. During the transform, the coding efficiency can vary depending on the distribution and characteristics of the values in the transform region. Quantization unit 115 quantizes the values of the transform coefficients output from transform unit 110.
[0082] To improve coding efficiency, instead of encoding the image signal as is, a method is used that predicts the image using the region already encoded by prediction unit 150, and obtains the reconstructed image by adding the residual value between the original image and the predicted image to the predicted image. To prevent mismatches in the encoder and decoder, information that can be used in the decoder should be used when performing prediction in the encoder. For this purpose, the encoder performs the processing of the current block of reconstruction encoding again. Inverse quantization unit 120 inverse quantizes the values of the transform coefficients, and inverse transform unit 125 uses the inverse quantized transform coefficient values to reconstruct the residual values. Simultaneously, filtering unit 130 performs filtering operations to improve the quality of the reconstructed image and improve coding efficiency. For example, it may include a deblocking filter, sample adaptive offset (SAO), and adaptive loop filter. The filtered image is output or stored in decoded image buffer (DPB) 156 for use as a reference image.
[0083] To improve coding efficiency, instead of encoding the image signal as is, a method is used to predict the image via prediction unit 150 by using the encoded region and adding the residual value between the original image and the predicted image to the predicted image, thereby obtaining a reconstructed image. Intra-frame prediction unit 152 performs intra-frame prediction within the current image, and inter-frame prediction unit 154 predicts the current image using a reference image stored in the decoded image buffer 156. Intra-frame prediction unit 152 performs intra-frame prediction from the reconstructed region in the current image and transmits the intra-frame coding information to entropy coding unit 160. Inter-frame prediction unit 154 may include motion estimation unit 154a and motion compensation unit 154b. Motion estimation unit 154a obtains the motion vector value of the current region by referencing a specific reconstructed region. Motion estimation unit 154a transmits the position information of the reference region (reference frame, motion vector, etc.) to entropy coding unit 160 so that the position information is included in the bitstream. Motion compensation unit 154b performs inter-frame motion compensation using the motion vector value transmitted from motion estimation unit 154a.
[0084] Prediction unit 150 includes intra-frame prediction unit 152 and inter-frame prediction unit 154. Intra-frame prediction unit 152 performs intra-frame prediction within the current image, and inter-frame prediction unit 154 performs inter-frame prediction to predict the current image using a reference image stored in DBP 156. Intra-frame prediction unit 152 performs intra-frame prediction from reconstructed samples in the current image and transmits intra-frame coding information to entropy coding unit 160. Intra-frame coding information may include at least one of intra-frame prediction mode, most probable mode (MPM) flag, and MPM index. Intra-frame coding information may include information about the reference samples. Inter-frame prediction unit 154 may include motion estimation unit 154a and motion compensation unit 154b. Motion estimation unit 154a obtains motion vector values for the current region by referencing a specific region of the reconstructed reference image. Motion estimation unit 154a transmits a set of motion information (reference image index, motion vector information, etc.) for the reference region to entropy coding unit 160. Motion compensation unit 154b performs motion compensation by using motion vector values passed from motion estimation unit 154a. Inter-frame prediction unit 154 passes inter-frame coding information, including motion information about the reference region, to entropy coding unit 160.
[0085] According to another embodiment, prediction unit 150 may include an intra-block copy (BC) prediction unit (not shown). The intra-BC prediction unit performs intra-BC prediction based on reconstructed samples in the current image and transmits intra-BC coding information to entropy coding unit 160. The intra-BC prediction unit obtains a specific region in the reference current image, indicating the block vector values of the reference region used to predict the current region. The intra-BC prediction unit can use the obtained block vector values to perform intra-BC prediction. The intra-BC prediction unit transmits intra-BC coding information to entropy coding unit 160. The intra-BC coding information may include block vector information.
[0086] When performing the image prediction described above, the transformation unit 110 transforms the residual value between the original image and the predicted image to obtain transformation coefficient values. In this case, the transformation can be performed on a unit of specific blocks within the image, and the size of the specific blocks can be changed within a preset range. The quantization unit 115 quantizes the transformation coefficient values generated in the transformation unit 110 and sends them to the entropy coding unit 160.
[0087] Entropy coding unit 160 performs entropy coding on information indicating quantized transform coefficients, intra-frame coding information, inter-frame coding information, etc., to generate a video signal bitstream. Variable-length coding (VLC) schemes, arithmetic coding schemes, etc., can be used in entropy coding unit 160. Variable-length coding (VLC) schemes involve transforming input symbols into continuous codewords, and the length of the codewords can be variable. For example, frequently occurring symbols are represented by short codewords, while rarely occurring symbols are represented by long codewords. Context-based adaptive variable-length coding (CAVLC) can be used as a variable-length coding scheme. Arithmetic coding can transform continuous data symbols into a single prime number, where arithmetic coding can obtain the optimal number of bits required to represent each symbol. Context-based adaptive binary arithmetic coding (CABAC) can be used as arithmetic coding. For example, entropy coding unit 160 can binarize information indicating quantized transform coefficients. Entropy coding unit 160 can generate a bitstream by arithmetic coding binary information.
[0088] The generated bitstream is encapsulated using Network Abstraction Layer (NAL) units as the basic unit. Each NAL unit comprises an integer number of encoded code tree units. To decode the bitstream in a video decoder, the bitstream must first be separated into NAL units, and then each separated NAL unit must be decoded. Simultaneously, the information required for decoding the video signal bitstream can be transmitted via Raw Byte Sequence Payload (RBSP), which consists of higher-level sets such as Picture Parameter Sets (PPS), Sequence Parameter Sets (SPS), and Video Parameter Sets (VPS).
[0089] at the same time, Figure 1 The block diagram illustrates an encoding device 100 according to an embodiment of the present invention, and the separately displayed blocks logically distinguish and illustrate the elements of the encoding device 100. Therefore, depending on the device design, the elements of the encoding device 100 can be mounted as one or more chips. According to an embodiment, the operation of each element of the encoding device 100 can be performed by a processor (not shown).
[0090] Figure 2 This is a schematic block diagram of a video signal decoding apparatus 200 according to an embodiment of the present invention. (See reference) Figure 2 The decoding device 200 of the present invention includes an entropy decoding unit 210, an inverse quantization unit 220, an inverse transform unit 225, a filtering unit 230, and a prediction unit 250.
[0091] Entropy decoding unit 210 performs entropy decoding on the video signal bitstream to extract transform coefficient information, intra-frame coding information, inter-frame coding information, etc., for each region. For example, entropy decoding unit 210 can obtain binary codes for transform coefficient information for a specific region from the video signal bitstream. Entropy decoding unit 210 obtains quantized transform coefficients by inverse binarizing the binary codes. Inverse quantization unit 220 inverse quantizes the quantized transform coefficients, and inverse transform unit 225 recovers the residual values using the inverse quantized transform coefficients. Video signal processing device 200 recovers the original pixel values by adding the residual values obtained by inverse transform unit 225 to the predicted values obtained by prediction unit 250.
[0092] Simultaneously, the filtering unit 230 performs filtering on the image to improve image quality. This may include a deblocking filter to reduce block distortion and / or an adaptive loop filter to remove distortion from the entire image. The filtered image is output or stored in the DPB 256 as a reference image for the next image.
[0093] Prediction unit 250 includes intra-frame prediction unit 252 and inter-frame prediction unit 254. Prediction unit 250 generates a prediction picture using the coding type decoded by the entropy decoding unit 210 described above, the transform coefficients of each region, and the intra / inter-frame coding information. To reconstruct the current block in which decoding is performed, the current picture or the decoded region of other pictures including the current block can be used. In reconstruction, only the current picture, i.e., the picture (or tile / slice) in which only intra-frame prediction or intra-frame BC prediction is performed, is referred to as an intra-frame picture or I-picture (or tile / slice), and the picture (or tile / slice) in which all intra-frame prediction, inter-frame prediction, and intra-frame BC prediction can be performed is referred to as an inter-frame picture (or tile / slice). To predict the sample value of each block within an inter-frame image (or tile / slice), an image (or tile / slice) using at most one motion vector and a reference image index is called a prediction image or P-image (or tile / slice), and an image (or tile / slice) using at most two motion vectors and a reference image index is called a bidirectional prediction image or B-image (or tile / slice). In other words, a P-image (or tile / slice) uses at most one set of motion information to predict each block, and a B-image (or tile / slice) uses at most two sets of motion information to predict each block. Here, the set of motion information includes one or more motion vectors and a reference image index.
[0094] Intra-prediction unit 252 generates prediction blocks using intra-coding information and reconstructed samples in the current image. As described above, the intra-coding information may include at least one of intra-prediction mode, most probable mode (MPM) flag, and MPM index. Intra-prediction unit 252 predicts sample values for the current block by using reconstructed samples located to the left and / or above the current block as reference samples. In this disclosure, the reconstructed samples, reference samples, and samples of the current block can represent pixels. Moreover, sample values can represent pixel values.
[0095] According to an embodiment, the reference sample may be a sample included in the adjacent blocks of the current block. For example, the reference sample may be a sample adjacent to the left boundary of the current block and / or a sample adjacent to the top boundary. Moreover, the reference sample may be a sample located on a line within a predetermined distance from the left boundary of the current block and / or a sample located on a line within a predetermined distance from the top boundary of the current block among the samples of the adjacent blocks of the current block. In this case, the adjacent blocks of the current block may include the left (L) block, the top (A) block, the bottom left (BL) block, the top right (AR) block, or the top left (AL) block.
[0096] Inter-frame prediction unit 254 generates prediction blocks using reference images and inter-frame coding information stored in DPB 256. The inter-frame coding information may include a set of motion information (reference image index, motion vector information, etc.) for the current block used as the reference block. Inter-frame prediction may include L0 prediction, L1 prediction, and bidirectional prediction. L0 prediction means making a prediction using a reference image included in the L0 image list, while L1 prediction means making a prediction using a reference image included in the L1 image list. For this purpose, a set of motion information (e.g., motion vectors and reference image index) may be required. In the bidirectional prediction method, up to two reference regions can be used, and the two reference regions may exist in the same reference image or in different images. That is, in the bidirectional prediction method, up to two sets of motion information (e.g., motion vectors and reference image indexes) can be used, and the two motion vectors may correspond to the same reference image index or different reference image indices. In this case, in terms of time, reference images can be displayed (or output) before and after the current image. According to an embodiment, the two reference regions used in the bidirectional prediction scheme can be regions selected from image list L0 and image list L1, respectively.
[0097] Inter-frame prediction unit 254 can obtain a reference block for the current block using motion vectors and a reference image index. The reference block is located in the reference image corresponding to the reference image index. Furthermore, the sample value of the block specified by the motion vector, or its interpolated value, can be used as a predictor for the current block. For motion prediction with sub-pellet unit pixel accuracy, for example, an 8-tap interpolation filter for the luma signal and a 4-tap interpolation filter for the chroma signal can be used. However, the interpolation filter for motion prediction at the sub-pixel level is not limited to this. In this way, inter-frame prediction unit 254 performs motion compensation to predict the texture of the current unit based on a motion image previously reconstructed using motion information. In such cases, the inter-frame prediction unit can use a set of motion information.
[0098] According to another embodiment, prediction unit 250 may include an intra-frame BC prediction unit (not shown). The intra-frame BC prediction unit can reconstruct a current region by referencing a specific region including reconstructed samples within the current image. The intra-frame BC prediction unit obtains intra-frame BC coding information of the current region from entropy decoding unit 210. The intra-frame BC prediction unit obtains block vector values indicating a specific region in the current image. The intra-frame BC prediction unit can perform intra-frame BC prediction using the obtained block vector values. The intra-frame BC coding information may include block vector information.
[0099] The reconstructed video image is generated by adding the predicted value output from the intra-frame prediction unit 252 or the inter-frame prediction unit 254 to the residual value output from the inverse transform unit 225. That is, the video signal decoding device 200 uses the predicted block generated by the prediction unit 250 and the residual obtained from the inverse transform unit 225 to reconstruct the current block.
[0100] at the same time, Figure 2 The block diagram illustrates a decoding device 200 according to an embodiment of the present invention, and the separately displayed blocks logically distinguish and illustrate the elements of the decoding device 200. Therefore, depending on the device design, the elements of the decoding device 200 can be mounted as one or more chips. According to an embodiment, the operation of each element of the decoding device 200 can be performed by a processor (not shown).
[0101] Figure 3The illustration shows an embodiment where a coding tree unit (CTU) in an image is segmented into coding units (CUs). During the encoding of a video signal, an image can be segmented into a series of coding tree units (CTUs). A coding tree unit consists of N×N blocks of luminance samples and two blocks of corresponding chrominance samples. A coding tree unit can be segmented into multiple coding units. A coding tree unit may not be segmented and may be a leaf node. In this case, the coding tree unit itself can be a coding unit. A coding unit refers to the basic unit used to process an image during the aforementioned video signal processing, i.e., intra / inter-frame prediction, transform, quantization, and / or entropy coding. The size and shape of a coding unit in an image may not be constant. A coding unit can have a square or rectangular shape. A rectangular coding unit (or rectangular block) includes vertical coding units (or vertical blocks) and horizontal coding units (or horizontal blocks). In this specification, a vertical block is a block whose height is greater than its width, and a horizontal block is a block whose width is greater than its height. Furthermore, in this specification, a non-square block may refer to a rectangular block, but this disclosure is not limited thereto.
[0102] refer to Figure 3 First, the coding tree unit is divided into a quadtree (QT) structure. That is, in a quadtree structure, a node of size 2×2N can be divided into four nodes of size N×N. In this specification, a quadtree can also be referred to as a quaternion tree. Quadtree partitioning can be performed recursively, and not all nodes need to be partitioned at the same depth.
[0103] Simultaneously, the leaf nodes of the aforementioned quadtree can be further divided into a multi-type tree (MTT) structure. According to embodiments of the present invention, in the MTT structure, a node can be divided into a horizontally or vertically partitioned binary or ternary tree structure. That is, in the MTT structure, there are four partitioning structures, such as vertical binary partitioning, horizontal binary partitioning, vertical ternary partitioning, and horizontal ternary partitioning. According to embodiments of the present invention, in each tree structure, the width and height of the node can both be powers of 2. For example, in a binary tree (BT) structure, a node of size 2×2N can be partitioned into two NX2N nodes through vertical binary partitioning, and further partitioned into two 2NXN nodes through horizontal binary partitioning. Additionally, in a ternary tree (TT) structure, a node of size 2×2N is partitioned into (N / 2)×2N, NX2N, and (N / 2)×2N nodes through vertical ternary partitioning, and further partitioned into 2NX(N / 2), 2NXN, and 2NX(N / 2) nodes through horizontal ternary partitioning. This multi-type tree split can be performed recursively.
[0104] Leaf nodes of multi-type trees can be coding units. When a coding unit is not greater than the maximum transform length, it can be used as a unit for prediction and / or transformation without further segmentation. As an example, when the width or height of the current coding unit is greater than the maximum transform length, the current coding unit can be segmented into multiple transform units without explicit signaling regarding segmentation. Alternatively, at least one of the following parameters in the quadtrees and multi-type trees described above can be predefined or transmitted through an RBSP of a higher-level set such as PPS, SPS, VPS, etc. 1) CTU size: Size of the root node of the quadtree; 2) Minimum QT size MinQtSize: Minimum allowed QT leaf node size; 3) Maximum BT size MaxBtSize: Maximum allowed BT root node size; 4) Maximum TT size MaxTtSize: Maximum allowed TT root node size; 5) Maximum MTT depth MaxMttDepth: Maximum allowed MTT depth derived from the leaf nodes of the QT; 6) Minimum BT size MinBtSize: Minimum allowed BT leaf node size; 7) Minimum TT size MinTtSize: Minimum allowed TT leaf node size.
[0105] Figure 4 The illustration shows an embodiment of a method for signaling the segmentation of quadtrees and multi-type trees. Preset flags can be used to signal the segmentation of the aforementioned quadtrees and multi-type trees. (Reference) Figure 4 At least one of the following flags can be used: "split_cu_flag" indicating whether to split a node, "split_qt_flag" indicating whether to split a quadtree node, "mtt_split_cu_vertical_flag" indicating the splitting direction of a multi-type tree node, or "mtt_split_cu_binary_flag" indicating the splitting shape of a multi-type tree node.
[0106] According to an embodiment of the present invention, a "split_cu_flag" signal can be sent first as an indication of whether the current node has been split. When the value of "split_cu_flag" is 0, it indicates that the current node has not been split, and the current node becomes a coding unit. When the current node is a coding tree unit, the coding tree unit includes an unsplit coding unit. When the current node is a quadtree node "QT node", the current node is a leaf node of the quadtree "QT leaf node" and becomes a coding unit. When the current node is a multi-class tree node "MTT node", the current node is a leaf node of the multi-class tree "MTT leaf node" and becomes a coding unit.
[0107] When the value of "split_cu_flag" is 1, the current node can be split into a quadtree or a multi-type tree node based on the value of "split_qt_flag". The encoding tree unit is the root node of the quadtree and can be initially split into a quadtree structure. In the quadtree structure, "split_qt_flag" is sent as a signal for each node, the "QT node". When the value of "split_qt_flag" is 1, the corresponding node is split into four square nodes, and when the value of "qt_split_flag" is 0, the corresponding node becomes a "QT leaf node" of the quadtree, and the corresponding node is split into multiple types of nodes. According to an embodiment of the present invention, the splitting of the quadtree can be restricted based on the type of the current node. Quadtree splitting is allowed when the current node is an encoding tree unit (the root node of the quadtree) or a quadtree node, and quadtree splitting is not allowed when the current node is a multi-type tree node. Each quadtree leaf node, the "QT leaf node", can be further split into a multi-type tree structure. As described above, when "split_qt_flag" is 0, the current node can be split into multiple types of nodes. To indicate the splitting direction and shape, "mtt_split_cu_vertical_flag" and "mtt_split_cu_binary_flag" can be sent using signals. When the value of "mtt_split_cu_vertical_flag" is 1, it indicates a vertical split of the "MTT node," and when the value of "mtt_split_cu_vertical_flag" is 0, it indicates a horizontal split of the "MTT node." Additionally, when the value of "mtt_split_cu_binary_flag" is 1, the "MTT node" is split into two rectangular nodes, and when the value of "mtt_split_cu_binary_flag" is 0, the "MTT node" is split into three rectangular nodes.
[0108] Image prediction (motion compensation) for encoding is performed on the coding units that are no longer divided (i.e., the leaf nodes of the coding unit tree). In the following text, the basic unit used to perform the prediction will be referred to as a "prediction unit" or "prediction block".
[0109] In the following text, the term "unit" as used herein may be used in place of a prediction unit, which is the basic unit used to perform prediction. However, this disclosure is not limited thereto, and "unit" may be understood to broadly encompass the concept of a coding unit.
[0110] Figure 5 and Figure 6The intra-frame prediction method according to an embodiment of the present invention is illustrated in more detail. As described above, the intra-frame prediction unit predicts the sample value of the current block by using reconstructed samples located to the left and / or above the current block as reference samples.
[0111] first, Figure 5 An embodiment of reference samples for prediction of the current block in intra-frame prediction mode is shown. According to the embodiment, the reference samples may be samples adjacent to the left boundary and / or the top boundary of the current block. Figure 5 As shown, when the size of the current block is WXH and a single reference line adjacent to the current block is used for intra-frame prediction, the reference samples can be configured using the maximum 2W+2H+1 neighboring samples located to the left and above the current block.
[0112] When at least some samples to be used as reference samples have not yet been recovered, the intra-prediction unit can obtain reference samples by performing a reference sample padding process. The intra-prediction unit can perform a reference sample filtering process to reduce errors in intra-prediction. That is, filtering can be performed on neighboring samples and / or reference samples obtained through the reference sample padding process to obtain filtered reference samples. The intra-prediction unit predicts samples for the current block using the reference samples obtained as described above. The intra-prediction unit predicts samples for the current block using either unfiltered or filtered reference samples. In this disclosure, neighboring samples can include samples on at least one reference line. For example, neighboring samples can include neighboring samples on lines adjacent to the boundary of the current block.
[0113] Next, Figure 6 An embodiment of a prediction mode for intra-frame prediction is illustrated. For intra-frame prediction, intra-frame prediction mode information indicating the direction of intra-frame prediction can be transmitted via signaling. The intra-frame prediction mode information indicates one of a plurality of intra-frame prediction modes included in the set of intra-frame prediction modes. When the current block is an intra-frame prediction block, the decoder receives the intra-frame prediction mode information of the current block from the bitstream. The decoder's intra-frame prediction unit performs intra-frame prediction on the current block based on the extracted intra-frame prediction mode information.
[0114] According to embodiments of the present invention, the intra-prediction mode set may include all intra-prediction modes used in intra-prediction (e.g., a total of 67 intra-prediction modes). More specifically, the intra-prediction mode set may include planar modes, DC modes, and multiple (e.g., 65) angle modes (i.e., orientation modes). Each intra-prediction mode can be indicated by a preset index (i.e., an intra-prediction mode index). For example, as... Figure 6As shown, intra-prediction mode index 0 indicates a planar mode, while intra-prediction mode index 1 indicates a DC mode. Furthermore, intra-prediction mode indices 2 through 66 can each indicate different angle modes. Each angle mode indicates an angle that differs from the others within a preset angle range. For example, an angle mode can indicate an angle within a clockwise angle range of 45 degrees and -135 degrees (i.e., a first angle range). An angle mode can be defined based on the 12 o'clock direction. In this case, intra-prediction mode index 2 indicates a horizontal diagonal (HDIA) mode, intra-prediction mode index 18 indicates a horizontal (horizontal, HOR) mode, intra-prediction mode index 34 indicates a diagonal (DIA) mode, intra-prediction mode index 50 indicates a vertical (VER) mode, and intra-prediction mode index 66 indicates a vertical diagonal (VDIA) mode.
[0115] Simultaneously, the preset angle range can be set differently depending on the shape of the current block. For example, if the current block is a rectangular block, a wide-angle mode indicating an angle exceeding 45 degrees or less than -135 degrees in the clockwise direction can be additionally used. When the current block is a horizontal block, the angle mode can indicate an angle within the clockwise direction between (45 + offset1) degrees and (-135 + offset1) degrees (i.e., the second angle range). In this case, angle modes 67 to 76, outside the first angle range, can be additionally used. Furthermore, if the current block is a vertical block, the angle mode can indicate an angle within the clockwise direction between (45 - offset2) degrees and (-135 - offset2) degrees (i.e., the third angle range). In this case, angle modes -10 to -1, outside the first angle range, can be additionally used. According to embodiments of this disclosure, the values of offset1 and offset2 can be determined differently depending on the ratio between the width and height of the rectangular block. Furthermore, offset1 and offset2 can be positive numbers.
[0116] According to another embodiment of the present invention, configuring multiple angle modes of the intra-frame prediction mode set may include a basic angle mode and an extended angle mode. In this case, the extended angle mode can be determined based on the basic angle mode.
[0117] According to an embodiment, the basic angle pattern is a pattern corresponding to the angle used in intra-frame prediction of the existing High Efficiency Video Coding (HEVC) standard, and the extended angle pattern can be a pattern corresponding to the angle newly added in intra-frame prediction of the next-generation video codec standard. More specifically, the basic angle pattern can be an angle pattern corresponding to any one of the intra-frame prediction patterns {2,4,6,...,66}, and the extended angle pattern can be an angle pattern corresponding to any one of the intra-frame prediction patterns {3,5,7,...,65}. That is, the extended angle pattern can be an angle pattern within a first angle range that falls between the basic angle patterns. Therefore, the angle indicated by the extended angle pattern can be determined based on the angle indicated by the basic angle pattern.
[0118] According to another embodiment, the basic angle mode can be a mode corresponding to an angle within a preset first angle range, and the extended angle mode can be a wide-angle mode outside the first angle range. That is, the basic angle mode can be an angle mode corresponding to any one of the intra-prediction modes {2,3,4,...,66}, and the extended angle mode can be an angle mode corresponding to any one of the intra-prediction modes {-10,-9,...,-1} and {67,68,...,76}. The angle indicated by the extended angle mode can be determined as the angle on the opposite side to the angle indicated by the corresponding basic angle mode. Therefore, the angle indicated by the extended angle mode can be determined based on the angle indicated by the basic angle mode. Meanwhile, the number of extended angle modes is not limited to this, and additional extended angles can be defined according to the size and / or shape of the current block. For example, the extended angle mode can be defined as an angle mode corresponding to any one of the intra-prediction modes {-14,-13,...,-1} and {67,68,...,80}. Furthermore, the total number of intra-prediction modes included in the intra-prediction mode set can vary depending on the configuration of the basic angle mode and the extended angle mode described above.
[0119] In the above embodiments, the interval between extended angle modes can be set based on the interval between corresponding basic angle modes. For example, the interval between extended angle modes {3,5,7,...,65} can be determined based on the interval between corresponding basic angle modes {2,4,6,...,66}. Additionally, the interval between extended angle modes {-10,-9,...,-1} can be determined based on the interval between corresponding basic angle modes {56,57,...,65} on the opposite side, and the interval between extended angle modes {67,68,...,76} can be determined based on the interval between corresponding basic angle modes {3,4,...,12} on the opposite side. The angular interval between extended angle modes can be set to be the same as the angular interval between corresponding basic angle modes. Furthermore, the number of extended angle modes in the intra-frame prediction mode set can be set to be less than or equal to the number of basic angle modes.
[0120] According to embodiments of the present invention, extended angle modes can be signaled based on a basic angle mode. For example, a wide-angle mode (i.e., an extended angle mode) can replace at least one angle mode (i.e., a basic angle mode) within a first angle range. The basic angle mode to be replaced can be a corresponding angle mode on the side opposite to the wide-angle mode. That is, the basic angle mode to be replaced is an angle mode corresponding to an angle in the opposite direction to the angle indicated by the wide-angle mode, or an angle mode corresponding to an angle differing from the angle in the opposite direction by a preset offset index. According to one embodiment of the present invention, the preset offset index is 1. The intra-frame prediction mode index corresponding to the basic angle mode to be replaced can be remapped to the wide-angle mode to signal the corresponding wide-angle mode. For example, wide-angle modes {-10, -9, ..., -1} can be signaled by intra-frame prediction mode indices {57, 58, ..., 66}, and wide-angle modes {67, 68, ..., 76} can be signaled by intra-frame prediction mode indices {2, 3, ..., 11}, respectively. In this way, the intra-prediction mode index used for the basic angle mode is used to signal the extended angle mode, and therefore, even if the configurations of the angle modes used for intra-prediction in each block are different, the same set of intra-prediction mode indexes can be used to signal the intra-prediction modes. Thus, signaling overhead caused by changes in the intra-prediction mode configuration can be minimized.
[0121] Simultaneously, the use of the extended angle mode can be determined based on at least one of the shape and size of the current block. According to an embodiment, when the size of the current block is greater than a preset size, the extended angle mode can be used for intra-frame prediction of the current block; otherwise, the basic angle mode can be used only for intra-frame prediction of the current block. According to another embodiment, when the current block is a block other than a square, the extended angle mode can be used for intra-frame prediction of the current block, and when the current block is a square block, only the basic angle mode can be used for intra-frame prediction of the current block.
[0122] In the following text, reference will be made to Figure 7 An inter-frame prediction method according to embodiments of the present invention is described. The inter-frame prediction method described herein may include a general inter-frame prediction method optimized for translational motion and an inter-frame prediction method based on an affine model. Furthermore, the motion vector may include at least one of a general motion vector for motion compensation according to the general inter-frame prediction method and a control point motion vector for affine motion compensation.
[0123] Figure 7 This diagram illustrates an inter-frame prediction method according to an embodiment of the present invention. As described above, the decoder can predict the current block by referring to a reconstructed sample from another decoded image. (Reference) Figure 7 The decoder obtains a reference block 702 in the reference image 720 based on the motion information set of the current block 701. In this case, the motion information set may include a reference image index and a motion vector 703. The reference image index indicates a reference image 720 that includes a reference block for inter-frame prediction of the current block in the reference image list. According to an embodiment, the reference image list may include at least one of the L0 image list and the L1 image list described above. The motion vector 703 represents the offset between the coordinate values of the current block 701 in the current image 710 and the coordinate values of the reference block 702 in the reference image 720. The decoder obtains a predictor for the current block 701 based on the sample values of the reference block 702 and uses the predictor to reconstruct the current block 701.
[0124] Specifically, the encoder can obtain the aforementioned reference block by searching for blocks similar to the current block in images with an earlier reconstruction order. For example, the encoder can search for a reference block within a preset search area where the sum of the differences between the current block and the sample values is minimized. In this case, to measure the similarity between the samples of the current block and the reference block, at least one of the sum of absolute differences (SAD) and the sum of Hadamard transform differences (SATD) can be used. Here, SAD can be a value obtained by adding all the absolute values of the individual differences between the sample values included in the two blocks. Alternatively, SATD can be a value obtained by adding all the absolute values of the Hadamard transform coefficients obtained by performing a Hadamard transform on the differences between the sample values included in the two blocks.
[0125] Simultaneously, one or more reference regions can be used to predict the current block. As described above, two or more reference regions can be used to perform inter-frame prediction of the current block using a dual prediction method. According to an embodiment, the decoder can obtain two reference blocks based on two sets of motion information of the current block. Furthermore, the decoder can obtain a first predictor and a second predictor for the current block based on the sample values of each of the two obtained reference blocks. Additionally, the decoder can use the first predictor and the second predictor to reconstruct the current block. For example, the decoder can reconstruct the current block based on the average of each sample from the first predictor and the second predictor.
[0126] As described above, for motion compensation of the current block, one or more sets of motion information can be signaled. In this case, the similarity between the motion information sets used for motion compensation of each of the multiple blocks can be used. For example, the motion information set for predicting the current block can be derived from the motion information set used to predict any of the other pre-reconstructed samples. In this way, the encoder and decoder can reduce signaling overhead. Various embodiments of signaling the motion information set of the current block will be described below.
[0127] Figure 8 This diagram illustrates a method for transmitting the motion vector of a current block using signals according to an embodiment of the present invention. According to an embodiment of the present invention, the motion vector of the current block can be derived from a motion vector predictor (MVP) of the current block. According to an embodiment, a candidate list of motion vector predictors (MVPs) can be used to obtain motion vector predictors referenced for deriving the motion vector of the current block. The MVP candidate list may include a preset number of MVP candidates (candidate 1, candidate 2, ..., candidate N).
[0128] According to an embodiment, the MVP candidate list may include at least one of spatial candidates and temporal candidates. A spatial candidate may be a set of motion information for predicting motion of neighboring blocks within a specific range from the current block in the current image. Spatial candidates can be constructed based on available neighboring blocks among the current block's neighboring blocks. Conversely, a temporal candidate may be a set of motion information for predicting motion of blocks in an image different from the current image. For example, a temporal candidate may be constructed based on a specific block in a specific reference image corresponding to the position of the current block. In this case, the position of the specific block represents the position of the upper-left sample of the specific block in the reference image. According to another embodiment, the MVP candidate list may include zero motion vectors. According to another embodiment, a rounding process may be performed on the MVP candidates included in the MVP candidate list for the current block. In this case, the resolution of the motion vector difference of the current block, which will be described later, can be used. For example, each of the MVP candidates for the current block may be rounded based on the resolution of the motion vector difference of the current block.
[0129] In this disclosure, the MVP candidate list may include an Advanced Temporal Motion Vector Prediction (ATMVP) list, a merge candidate list for merging inter-frame predictions, a control point motion vector candidate list for affine motion compensation, a sub-block-based Temporal Motion Vector Prediction (STMVP) list, and combinations thereof.
[0130] According to an embodiment, encoder 810 and decoder 820 can construct an MVP candidate list for motion compensation of the current block. For example, among the samples reconstructed before the current block, there may be candidates corresponding to samples that may have been predicted based on the same or similar motion information set as the current block. Encoder 810 and decoder 820 can construct an MVP candidate list for the current block based on multiple candidate blocks. In this case, encoder 810 and decoder 820 can construct the MVP candidate list according to predefined rules between encoder 810 and decoder 820. That is, the MVP candidate lists constructed in encoder 810 and decoder 820 respectively can be identical to each other.
[0131] Furthermore, the predefined rules can vary depending on the prediction mode of the current block. For example, when the prediction mode of the current block is an affine prediction mode based on an affine model, the encoder and decoder can construct the MVP candidate list for the current block using a first method based on the affine model. The first method could be a method for obtaining a candidate list of control point motion vectors. On the other hand, when the prediction mode of the current block is a general inter-frame prediction mode not based on an affine model, the encoder and decoder can use a second method not based on an affine model to construct the MVP candidate list for the current block. In this case, the first method and the second method can be different methods.
[0132] Decoder 820 can derive the motion vector of the current block based on any one of at least one MVP candidate included in the MVP candidate list for the current block. For example, encoder 810 can signal an MVP index of a motion vector predictor to be referenced for deriving the motion vector of the current block. Decoder 820 can obtain the motion vector predictor of the current block based on the signaled MVP index. Decoder 820 can use the motion vector predictor to derive the motion vector of the current block. According to an embodiment, decoder 820 can use the motion vector predictor obtained from the MVP candidate list as the motion vector of the current block without separate motion vector interpolation. Decoder 820 can reconstruct the current block based on the motion vector of the current block. The inter-frame prediction mode in which the motion vector predictor obtained from the MVP candidate list is used as the motion vector of the current block without separate motion vector interpolation can be referred to as a merging mode.
[0133] According to another embodiment, decoder 820 can obtain a separate motion vector difference for the motion vector of the current block. Decoder 820 can obtain the motion vector of the current block by adding the motion vector difference of the current block to the motion vector predictor obtained from the MVP candidate list. In this case, encoder 810 can signal a motion vector (MV) difference value, MV difference, indicating the difference between the motion vector of the current block and the motion vector predictor. (Refer to...) Figure 9 This describes in detail a method for transmitting motion vector differences using signals. Decoder 820 can obtain the motion vector of the current block based on the motion vector difference MV. Decoder 820 can reconstruct the current block based on the motion vector of the current block.
[0134] Additionally, a reference image index for motion compensation of the current block can be signaled. In the prediction of the current block's pattern, encoder 810 can signal a reference image index indicating a reference image of the reference block. Decoder 820 can obtain the POC of the reference image to be referenced to reconstruct the current block based on the signaled reference image index. In this case, the POC of the reference image may differ from the POC of the reference image corresponding to the MVP used to derive the motion vectors of the current block. In this case, decoder 820 can perform motion vector scaling. That is, decoder 820 can obtain MVP' by scaling MVP. In this case, motion vector scaling can be performed based on the POC of the current image, the POC of the signaled reference image of the current block, and the POC of the reference image corresponding to the MVP. Furthermore, decoder 820 can use MVP' as a motion vector predictor for the current block.
[0135] As described above, the motion vector of the current block can be obtained by adding the motion vector predictor of the current block to the motion vector difference. In this case, the motion vector difference can be transmitted as a signal from the encoder. The encoder can encode the motion vector difference to generate and transmit information representing the motion vector difference as a signal. Hereinafter, a method for transmitting motion vector differences as a signal according to an embodiment of the present invention will be described.
[0136] Figure 9 This diagram illustrates a method for transmitting motion vector differences of the current block using signals according to an embodiment of the present invention. According to the embodiment, the information indicating the motion vector difference may include at least one of absolute value information and sign information of the motion vector difference. The absolute value and sign of the motion vector difference may be encoded separately.
[0137] According to an embodiment, the absolute value of the motion vector difference may not be transmitted as the value itself in the signal. The encoder can reduce the magnitude of the value to be transmitted in the signal by using at least one flag indicating the characteristics of the absolute value of the motion vector difference. The decoder can derive the absolute value of the motion vector difference from the value transmitted in the signal by using at least one flag.
[0138] For example, at least one flag may include a first flag indicating whether the absolute value of the motion vector difference is greater than N. In this case, N may be an integer. When the magnitude of the absolute value of the motion vector difference is greater than N, the value of (absolute value of motion vector difference - N) may be signaled along with the activated first flag. In this case, the activated flag may indicate that the magnitude of the absolute value of the motion vector difference is greater than N. The decoder may obtain the absolute value of the motion vector difference based on the activated first flag and the value signaled.
[0139] refer to Figure 9 A second flag, abs_mvd_greater0_flag, indicating whether the absolute value of the motion vector difference is greater than "0" can be sent via a signal. When the second flag abs_mvd_greater0_flag[] indicates that the absolute value of the motion vector difference is not greater than "0", the absolute value of the motion vector difference can be "0". Alternatively, when the second flag abs_mvd_greater0_flag indicates that the absolute value of the motion vector difference is greater than "0", the decoder can use other information about the motion vector difference to obtain its absolute value.
[0140] According to an embodiment, a third flag, abs_mvd_greater1_flag, indicating whether the absolute value of the motion vector difference is greater than "1" can be sent by signaling. When the third flag abs_mvd_greater1_flag indicates that the absolute value of the motion vector difference is not greater than '1', the decoder can determine that the absolute value of the motion vector difference is '1'.
[0141] Conversely, when the third flag `abs_mvd_greater1_flag` indicates that the absolute value of the motion vector difference is greater than '1', the decoder can use additional information about the motion vector difference to obtain its absolute value. For example, the value of `abs_mvd_minus2` (the absolute value of the motion vector difference - 2) can be sent as a signal. This is because, when the absolute value of the motion vector difference is greater than '1', the absolute value of the motion vector difference can be 2 or a larger value.
[0142] As described above, the absolute value of the motion vector difference of the current block can be transformed into at least one flag. For example, the transformed absolute value of the motion vector difference can be represented according to the magnitude of the motion vector difference (absolute value of motion vector difference - N). According to an embodiment, the transformed absolute value of the motion vector difference can be signaled using at least one bit. In this case, the number of bits signaled to indicate the transformed absolute value of the motion vector difference can be variable. The encoder can encode the transformed absolute value of the motion vector difference using a variable-length binarization method. For example, the encoder can use at least one of truncated unary binarization, unary binarization, truncated Ricean binarization, or exponential Columbus binarization as the variable-length binarization method.
[0143] Additionally, the sign of the motion vector difference can be signaled using the sign flag `mvd_sign_flag`. Furthermore, the sign of the motion vector difference can be implicitly signaled using sign bit hiding.
[0144] Simultaneously, the motion vector difference of the current block can be transmitted by signal in units of a specific resolution. In this disclosure, the resolution of the motion vector difference can indicate the unit in which the motion vector difference is transmitted by signal. That is, in this disclosure, a resolution other than the resolution of the image can represent the precision or granularity of transmitting the motion vector difference by signal. The resolution of the motion vector difference can be represented in units of samples or pixels. For example, the resolution of the motion vector difference can be represented using sample units (such as one-quarter, one-half, one, two, or four sample units). Furthermore, as the resolution of the motion vector difference of the current block decreases, the precision of the motion vector difference of the current block can increase.
[0145] According to embodiments of the present invention, motion vector differences can be transmitted as signals based on various resolutions. According to embodiments, the absolute value or deformed absolute value of the motion vector difference can be transmitted as a value in units of integer samples. Alternatively, the absolute value of the motion vector difference can be transmitted as a value in units of 1 / 2 subpixels. That is, the resolution of the motion vector difference can be set differently depending on the circumstances. The encoder and decoder according to embodiments of the present invention can efficiently transmit the motion vector difference of the current block as signals by appropriately utilizing various resolutions of the motion vector difference.
[0146] According to an embodiment, for each unit of at least one of a block, coding unit, slice, or tile, the resolution of the motion vector difference can be set to a different value. For example, the first resolution of the motion vector difference for the first block can be a unit of 1 / 4 sample. In this case, "64" can be signaled, which is the value obtained by dividing the absolute value of the motion vector difference "16" by the first resolution. Additionally, the second resolution of the motion vector difference for the second block can be an integer sample unit. In this case, "16" can be signaled, which is the value obtained by dividing the absolute value of the second motion vector difference "16" by the second resolution. In this way, even when the absolute values of the motion vector differences are the same, different values can be signaled according to the resolution. In this case, when the value obtained by dividing the absolute value of the motion vector difference by the resolution includes decimal places, a rounding function can be applied to the corresponding value.
[0147] The encoder can transmit information indicating the motion vector difference using a signal based on the resolution of the motion vector difference. The decoder can obtain a modified motion vector difference from the transmitted motion vector difference. The decoder can modify the motion vector difference based on the resolution of the resolution difference. The relationship between the transmitted motion vector difference valuePerResolution and the modified motion vector difference valueDetermined for the current block is represented by Equation 1 below. In the following, in this disclosure, unless otherwise stated, the motion vector difference indicates the modified motion vector difference valueDetermined. Furthermore, the transmitted motion vector difference represents the value before resolution modification.
[0148] [Equation 1]
[0149] valueDetermined=resolution*valuePerResolution
[0150] In Equation 1, resolution represents the resolution of the motion vector difference of the current block. That is, the decoder can obtain the modified motion vector difference by multiplying the motion vector difference transmitted by the signal for the current block by the resolution. Next, the decoder can obtain the motion vector of the current block based on the modified motion vector difference. Furthermore, the decoder can reconstruct the current block based on its motion vector.
[0151] When a relatively small value is used as the resolution for the motion vector difference of the current block (i.e., when the precision is high), it may be advantageous to represent the motion vector difference of the current block more accurately. However, in this case, the signaling overhead of the motion vector difference of the current block may increase because the value being signaled itself becomes larger. Conversely, when a relatively large value is used as the resolution for the motion vector difference of the current block (i.e., when the precision is low), the signaling overhead of the motion vector difference can be reduced by decreasing the amount of the value being signaled. That is, when the resolution of the motion vector difference is large, the motion vector difference of the current block can be signaled with fewer bits than when the resolution of the motion vector difference of the current block is small. However, in this case, it may be difficult to represent the motion vector difference of the current block accurately.
[0152] Therefore, the encoder and decoder can select a favorable resolution from multiple resolutions for transmitting motion vector differences using signals, depending on the circumstances. For example, the encoder can transmit the selected resolution using signals based on the circumstances. Additionally, the decoder can obtain the motion vector differences of the current block based on the resolution used for signal transmission. Hereinafter, a method for transmitting the resolution of motion vector differences of the current block using signals according to an embodiment of the present invention will be described. According to an embodiment of the present invention, the resolution of the motion vector differences of the current block can be any one of multiple available resolutions included in a resolution set. Here, multiple available resolutions can indicate the resolutions available under specific circumstances. Furthermore, the type and number of available resolutions included in the resolution set can vary depending on the circumstances.
[0153] Figure 10 This is a diagram illustrating the encoding unit and the transformation unit according to an embodiment of the present invention.
[0154] Figure 10 (a) is a diagram illustrating an encoding unit according to an embodiment of the present invention, and Figure 10 (b) is a diagram illustrating a transformation unit according to an embodiment of the present invention.
[0155] According to embodiments of the present invention, there may be block units that perform transformations on it. For example, the block units that perform transformations on it may be smaller than or equal to the coding unit (CU), or smaller than or equal to the block units that perform predictions on it. Additionally, the size of the block units that perform transformations on it may be limited by a maximum transformation size. For example, the width or height of the block units that perform transformations on it may be limited by the maximum transformation size. Specifically, the width or height of the block units that perform transformations on it may be smaller than or equal to the maximum transformation size. Furthermore, the maximum transformation size may vary depending on the color difference components (i.e., the luminance component and the chrominance component) of the transform block. In this case, the block units that perform transformations on it can be represented as transform units (TUs). If the coding unit or prediction unit (PU) is larger than the maximum transformation size, the coding unit or prediction unit can be partitioned to generate multiple transform units. Among the multiple transform units, the transform units larger than the maximum transformation size can be partitioned, and multiple transform units can be generated. The sizes of the multiple transform units generated by this process can all be smaller than or equal to the maximum transformation size. In the present invention, the fact that the size of a unit (block) is larger than the maximum transformation size may mean that the width or height of the unit (block) is larger than the maximum transformation size. Simultaneously, the fact that the size of a cell (block) is less than or equal to the maximum transform size may mean that both the width and height of the cell (block) are less than or equal to the maximum transform size. According to embodiments of the invention, operations based on partitioning transform units according to the maximum transform size can be performed, such as when performing intra-frame prediction, when processing residual signals, etc. Operations based on partitioning transform units according to the maximum transform size can be performed recursively, and can be performed until the size of the transform unit becomes less than or equal to the maximum transform size.
[0156] refer to Figure 10 (a) The width and height of the encoding unit can be represented as cbWidth and cbHeight, respectively. (See reference) Figure 10 (b) The width and height of the transform unit can be represented as tbWidth and tbHeight, respectively. The maximum transform size can be represented as MaxTbSizeY. Specifically, MaxTbSizeY can be the maximum transform size used for the luminance component. When the size of the transform unit is larger than the maximum transform size and therefore the transform unit needs to be segmented, the tbWidth and tbHeight of the transform unit before segmentation can be cbWidth and cbHeight, respectively. That is, the transform unit can be segmented based on whether its width or height is greater than the maximum transform size, and tbWidth or tbHeight can be updated. For example, refer to... Figure 10(a) where cbWidth is greater than MaxTbSizeY. In this case, the transform unit can be segmented. Before segmenting the transform unit, the tbWidth and tbHeight of the transform unit can be cbWidth and cbHeight, respectively. (See reference...) Figure 10 (b) Since tbWidth is greater than MaxTbSizeY, TU can be partitioned into transform unit 1 and transform unit 2. In this case, newTbWidth, which is the width of the new transform unit (the width of the transform unit after partitioning), can be tbWidth / 2.
[0157] The transformation unit (TU) partition described in this specification may refer to a partition included in a transformation block (TB) within the TU.
[0158] In this specification, the terms "unit" and "block" may be used interchangeably. Furthermore, a unit may be a concept comprising one or more blocks based on color difference components. For example, a unit may be a concept comprising blocks for luminance and one or more blocks for chromaticity.
[0159] Figure 11 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0160] Figure 11 The syntax can be used to execute references Figure 10 The syntax for describing the partitioning of the TU.
[0161] It can be called from the encoding unit syntax or the prediction unit syntax. Figure 11 The transformation tree syntax is shown. In this case, tbWidth and tbHeight, which are the input values for the transformation tree syntax, can be referenced. Figure 10 The described values are cbWidth and cbHeight. Additionally, the decoder can check if tbWidth is greater than the maximum transform size MaxTbSizeY or if tbHeight is greater than MaxTbSizeY. As a result of this check, if tbWidth or tbHeight is greater than MaxTbSizeY, TU partitioning can be performed, and the transform tree syntax can be called again. Otherwise, the transform unit syntax can be called. In this case, when the transform tree syntax is called again, the input values can be updated. For example, if tbWidth or tbHeight is greater than MaxTbSizeY, the input values can be updated to tbWidth / 2 and tbHeight / 2 respectively. Figure 11The updated tbWidth and tbHeight are represented as trafoWidth and trafoHeight. Additionally, in the syntax element transform_tree(x0,y0,tbWidth,tbHeight,treeType), x0 and y0 can be the horizontal and vertical coordinates indicating the position of the block to which the transform tree syntax is applied. The decoder can use the values of trafoWidth and trafoHeight to parse the syntax element transform_tree(x0,y0,trafoWidth,trafoHeight,treeType). When tbWidth is greater than MaxTbSizeY, the decoder can parse the syntax element transform_tree(x0+trafoWidth,y0,trafoWidth,trafoHeight,treeType). When tbHeight is greater than MaxTbSizeY, the decoder can parse the syntax element transform_tree(x0,y0+trafoHeight,trafoWidth,trafoHeight,treeType). When tbWidth is greater than MaxTbSizeY and tbHeight is greater than MaxTbSizeY, the decoder can parse the syntax element transform_tree(x0+trafoWidth,y0+trafoHeight,trafoWidth,trafoHeight,treeType). In this case, the order in which the transform tree syntax "transform_tree()" is called multiple times may be important in terms of encoder-decoder matching or encoding performance. In this case, the order can follow a preset order, and the preset order can be based on... Figure 11 The order of the publicly available grammar.
[0162] Additionally, when first called in the coding unit or prediction unit Figure 11 When parsing the transform tree syntax, tbWidth, tbHeight, x0, and y0 can be values set based on the luma component. For example, when the luma block size is 16×16, the chroma block size according to the chroma format can be 8×8. In this case, even if the decoder parses the transform tree syntax used for the chroma block, the tbWidth and tbHeight values can be 16 and 16 respectively.
[0163] Additionally, it can be analyzed Figure 11The transform tree syntax is determined regardless of the `treeType` value. `treeType` indicates whether the block structure of the luma component is the same as or different from the block structure of the chroma component. For example, when `treeType` is `SINGLE_TREE`, the block structures of the luma and chroma components can be the same. Conversely, when `treeType` is not `SINGLE_TREE`, the block structures of the luma and chroma components can be different. When `treeType` is `DUAL_TREE_LUMA`, the block structures of the luma and chroma components can be different. In this case, `treeType` can indicate the tree used for the luma component or indicate the transform tree syntax used to parse the luma component. On the other hand, when `treeType` is `DUAL_TREE_CHROMA`, the block structures of the luma and chroma components can be different. In this case, `treeType` can indicate the tree used for the chroma component or indicate the transform tree syntax used to parse the chroma component.
[0164] Figure 12 This is a diagram illustrating the decoding process in an intra-block according to an embodiment of the present invention.
[0165] Figure 12 This is a diagram illustrating the transformation unit partitioning and intra-frame prediction process.
[0166] In the embodiments described with reference to the accompanying drawings, the terminology used is defined.
[0167] (xTb0, yTb0): The coordinates of the top left sample position of the current transform block.
[0168] nTbW: Width of the current transform block.
[0169] nTbH: The height of the current transform block.
[0170] Figure 12 Clause 8.4.1 disclosed herein illustrates the decoding process for the luma block and the chroma block of the coding unit. The decoding process for the chroma block may include the decoding process for the Cb block and the Cr block. (See reference...) Figure 12 Clause 8.4.1 states that different inputs can be used in the decoding process of luma blocks and chroma blocks, and the decoder can use different inputs to perform [the decoding process]. Figure 12 The operation disclosed in Clause 8.4.5.1.
[0171] Figure 12The `cbWidth` and `cbHeight` disclosed in Clause 8.4.1 can refer to the width and height of the current coded block, and the values of `cbWidth` and `cbHeight` can be based on the values of the luma samples. For example, if the luma block size is 16×16, then the chroma block size can be 8×8 depending on the chroma format. In this case, the values of `cbWidth` and `cbHeight` used for the chroma coded block can be 16 and 16, respectively. Additionally, Figure 12 The (xCb, yCb) disclosed can indicate the coordinates of the current coding block, or it can be the coordinate value of the top-left sample of the current coding block. In this case, the coordinate value can be set based on the coordinates of the top-left sample of the current image. Alternatively, (xCb, yCb) can be a value based on the luminance sample representation.
[0172] refer to Figure 12 Clause 8.4.1 states that when `treeType` is `SINGLE_TREE` or `DUAL_TREE_LUMA`, the decoder can perform the decoding process on the luma block. In this case, it can call... Figure 12 According to clause 8.4.5.1 disclosed in the document, the sample position input can be (xCb, yCb), and the width and height of the block can be cbWidth and cbHeight, respectively. Additionally, the value of cIdx, which indicates the color component, can be 0.
[0173] When `treeType` is `SINGLE_TREE` or `DUAL_TREE_CHROMA`, the decoder can perform the decoding process on the chroma blocks. In this case, it can call... Figure 12 Clause 8.4.5.1 disclosed herein can be invoked for each of the Cb and Cr blocks. In this case, the sample position input can be (xCb / SubWidthC, yCb / SubHeightC), and the width and height of the block can be cbWidth / SubWidthC and cbHeight / SubHeightC, respectively. In this case, the value of cIdx can be non-0, and the value of cIdx can be set to 1 for the Cb block and 2 for the Cr block. SubWidthC and SubHeightC can be preset values based on the chroma format. The values of SubWidthC and SubHeightC can be 1 or 2, and can be values used to represent the relationship between the luminance and chroma components. See below for further details. Figures 14 to 15 Describe SubWidthC and SubHeightC.
[0174] When the decoder executes Figure 12During the decoding process disclosed in Clause 8.4.5.1, the input can be in a transition state to correspond to luma and chroma blocks. For example, Figure 12 The block width nTbW, block height nTbH, and sample position coordinates (xTb0, yTb0) disclosed in Clause 8.4.5.1 can be determined based on the number of chroma samples during the decoding process of the chroma block, and can also be determined based on the number of luminance samples during the decoding process of the luma block.
[0175] The above TU partition can be accessed via Figure 12 To be executed in accordance with clause 8.4.5.1. Figure 12 The disclosed values `maxTbWidth` and `maxTbHeight` indicate the maximum transformation size corresponding to the color component, and can be values corresponding to width and height, respectively. As mentioned above, `MaxTbSizeY` can be the maximum transformation size of the luma block. Therefore, for the chroma block (i.e., when `cIdx` is not 0), `maxTbWidth` can be `MaxTbSizeY / SubWidth`, and `maxTbHeight` can be `MaxTbSizeY / SubHeightC`. For the luma block, both `maxTbWidth` and `maxTbHeight` can be `MaxTbSizeY`.
[0176] Additionally, refer to Figure 12 8-43, the coordinates (xTbY, yTbY) can be transformed based on brightness, and in this case, the coordinates can be calculated based on cIdx, SubWidthC, and SubHeightC.
[0177] Figure 12 The decoding process described herein should match the syntax structure. That is to say, Figure 12 The decoding process should match Figure 11 The grammatical structure.
[0178] When in Figure 12 In clause 8.4.5.1, if nTbW is greater than maxTbWidth or nTbH is greater than maxTbHeight, execution is permitted. Figure 12 Steps 1 to 5 are disclosed in Clause 8.4.5.1. See also: Figure 12 Step 1 involves updating the block's width (nTbW) and height (nTbH) using `newTbW` and `newTbH`. Specifically, if `nTbW` is greater than `maxTbWidth`, the block's width is updated to `nTbW / 2`; otherwise, the width remains `nTbW`. If `nTbH` is greater than `maxTbHeight`, the block's height is updated to `nTbH / 2`; otherwise, the height remains `nTbH`. (See reference...) Figure 12Step 2, by using coordinates (xTb0, yTb0), newTbW, and newTbH as input, can be called again. Figure 12 Clause 8.4.5.1. (See reference.) Figure 12 Step 3, if nTbW is greater than maxTbWidth, can be called again by using the coordinates (xTb0+newTbW,yTb0), newTbW, and newTbH as input. Figure 12 Clause 8.4.5.1. (See reference.) Figure 12 In process 4, if nTbH is greater than maxTbHeight, it can be called again by using the coordinates (xTb0, yTb0+newTbH), newTbW, and newTbH as input. Figure 12 Section 8.4.5.1. (See reference.) Figure 12 Step 5, if nTbW is greater than maxTbWidth and nTbH is greater than maxTbHeight, can be called again by using the coordinates (xTb0+newTbW, yTb0+newTbH), newTbW, and newTbH as input. Figure 12 Clause 8.4.5. Figure 12 Steps 2 to 5 can be combined with Figure 11 The process of calling the transformation tree syntax again is the same.
[0179] Additionally, if nTbW is less than or equal to maxTbWidth and nTbH is less than or equal to maxTbHeight, then the following can be executed: Figure 12 The process other than steps 1 to 5. In this case, besides Figure 12 The processes beyond steps 1 to 5 can be processes related to actual intra-frame prediction, residual signal decoding, transformation, and reconstruction.
[0180] Figure 13 This is a diagram illustrating the decoding process of a residual signal according to an embodiment of the present invention.
[0181] exist Figure 13 The content disclosed herein will omit any content that is redundant with the above.
[0182] When applying inter-frame prediction, intra-block copy (IBC) prediction, etc., it can be executed. Figure 13 The decoding process is publicly disclosed. Additionally... Figure 13 The TU partition mentioned above can be made public.
[0183] refer to Figure 13 S1301 can perform calls for each color component. Figure 13 The steps in section 8.5.8. In this case, refer to... Figure 12The above can be performed using inputs corresponding to each color component. Figure 13 The decoding process of clause 8.5.8. As... Figure 13 The input sample position coordinates (xTb0, yTb0), block width nTbW, and block height nTbH for the decoding process in Clause 8.5.8 can be (xCb, yCb), CbWidth, and CbHeight for luma blocks, and (xCb / SubWidthC, yCb / SubHeightC), cbWidth / SubWidthC, and cbHeight / SubHeightC for chroma blocks. Additionally, the cIdx value indicating the color component can be set to 0 when the block is a luma component and to a value other than 0 when the block is a chroma component. For example, when the block is a Cb component, the cIdx value can be set to 1, and when the block is a Cr component, the cIdx value can be set to 2.
[0184] When execution Figure 13 During the decoding process of Clause 8.5.8, (xTb0, yTb0), nTbW, and nTbH can be values corresponding to each color component. That is, (xTb0, yTb0), nTbW, and nTbH can be values corresponding to the number of samples for each color component.
[0185] refer to Figure 13 8-849 and 8-850 can adaptively calculate the maximum transformation size for each color component. Additionally, refer to... Figure 13 8-852 and 8-853 update the width or height of the block based on whether nTbW or nTbH is greater than the maximum transform size, and can be described as newTbW and newTbH. Figure 13 Steps 2 through 5 of Clause 8.5.8 disclose the procedure for re-invoking Clause 8.5.8, which can be referenced above. Figure 12 The content described is the same.
[0186] Additionally, if nTbW is less than or equal to maxTbWidth and nTbH is less than or equal to maxTbHeight, then the following can be executed: Figure 13 The process other than steps 1 to 5. In this case, besides Figure 13 The processes other than steps 1 to 5 can be processes related to actual residual signal decoding, transformation, etc.
[0187] Figure 14 This is a diagram illustrating the relationship between color components according to an embodiment of the present invention.
[0188] refer to Figure 14Elements related to color components may include chroma_format_idc, chroma format, separate_colour_plane_flag, etc.
[0189] For example, when the chroma format is monochrome, there may be only one sample array, and SubWidthC and SubHeightC can both be 1. When the chroma format is 4:2:0 sampling, there may be two chroma arrays. In this case, the chroma array can have half the width and half the height of the luma array, and both SubWidthC and SubHeightC can be 2. When the chroma format is 4:2:2 sampling, there may be two chroma arrays. In this case, the chroma array can have half the width of the luma array and the same height as the luma array, and SubWidthC and SubHeightC can be 2 and 1 respectively. When the chroma format is 4:4:4 sampling, there may be two chroma arrays. In this case, the chroma array can have the same width and the same height as the luma array, and both SubWidthC and SubHeightC can be 1.
[0190] Furthermore, when the chroma format is 4:4:4 sampling, the process executed based on `separate_colour_plane_flag` can differ. When `separate_colour_plane_flag` is 0, the chroma array can have the same width and height as the luma array. When `separate_colour_plane_flag` is 1, separate processes can be performed for the three color planes (i.e., luma, Cb, and Cr). When `separate_colour_plane_flag` is 1, only one color component can exist in a slice. Conversely, when `separate_colour_plane_flag` is 0, multiple color components can exist in a slice. For example... Figure 14 As shown, when the chroma format is 4:4:4 sampling, both SubWidthC and SubHeightC can be 1, regardless of separate_colour_plane_flag.
[0191] SubWidthC and SubHeightC can indicate how large the chroma array is compared to the luma array. When the width or height of the chroma array is half that of the luma array, SubWidthC or SubHeightC can be 2, and when the width or height of the chroma array is the same as that of the luma array, SubWidthC or SubHeightC can be 1.
[0192] refer to Figure 14SubWidthC and SubHeightC can have different values only when the chroma format is 4:2:2 sampling. Therefore, when the chroma format is 4:2:2 sampling, the relationship between width and height based on the luma component can be different from the relationship between width and height based on the chroma component.
[0193] Figure 15 This is a diagram illustrating the relationship between color components according to an embodiment of the present invention.
[0194] Figure 15 (a) The illustration shows the case where the chroma format is 4:2:0 sampling. Figure 15 (b) The illustration shows the case where the chroma format is 4:2:2 sampling, and Figure 15 (c) The illustration shows the case where the chroma format is 4:4:4 sampling.
[0195] refer to Figure 15 (a) When the chroma format is 4:2:0 sampling, one chroma sample (one Cb, one Cr) can be located for every two luminance samples in the horizontal direction. In addition, one chroma sample (one Cb, one Cr) can be located for every two luminance samples in the vertical direction.
[0196] refer to Figure 15 (b) When the chroma format is 4:2:2 sampling, one chroma sample (one Cb, one Cr) can be located for every two luminance samples in the horizontal direction. In addition, one chroma sample (one Cb, one Cr) can be located for every luminance sample in the vertical direction.
[0197] refer to Figure 15 (c) When the chroma format is 4:4:4 sampling, one chroma sample (one Cb, one Cr) can be located for each luminance sample in the horizontal direction. Additionally, one chroma sample (one Cb, one Cr) can be located for each luminance sample in the vertical direction.
[0198] According to Figure 15 The relationship between the luminance samples and chrominance samples shown is used to determine the aforementioned SubWidthC and SubHeightC, and luminance sample-based transformations and chrominance sample-based transformations can be performed based on SubWidthC and SubHeightC.
[0199] Figure 16 This is a diagram illustrating the maximum transformation size according to an embodiment of the present invention.
[0200] The maximum transform size can be variable, and the complexity of the encoder or decoder can be adjusted by changing the maximum transform size. For example, a smaller maximum transform size can reduce the complexity of the encoder or decoder.
[0201] In an embodiment of the present invention, a value that can be the maximum transform size can be limited. For example, the maximum transform size can be limited to either of two values. In this case, the two values can be 32 or 64, and can be based on the size of luminance.
[0202] The maximum transform size can be signaled at a higher level. In this case, the higher level can be the level that includes the current block. For example, units such as a sequence, sequence parameter, slice, tile, tile group, picture, and coding tree unit (CTU) can be the higher level.
[0203] Reference Figure 16 , sps_max_luma_transform_size_64_flag can be a flag indicating the maximum transform size. For example, when the value of sps_max_luma_transform_size_64_flag is 1, the maximum transform size can be 64, and when the value of sps_max_luma_transform_size_64_flag is 0, the maximum transform size can be 32. In this case, the maximum transform size can be based on the size of luminance samples.
[0204] In addition, MaxTbLog2SizeY can be a value obtained by taking the log2 of the maximum transform size. Therefore, MaxTbLog2SizeY can be (sps_max_luma_transform_size_64_flag? 6:5). That is, when the value of sps_max_luma_transform_size_64_flag is 1, the value of MaxTbLog2SizeY is 6, and when the value of sps_max_luma_transform_size_64_flag is 0, the value of MaxTbLog2SizeY is 5.
[0205] In addition, MaxTbSizeY indicating the maximum transform size can be (1 << MaxTbLog2SizeY). That is, if the value of MaxTbLog2SizeY is 6, the bits are shifted left by 6 spaces, and the value of MaxTbSizeY becomes 64. If the value of MaxTbLog2SizeY is 5, the bits are shifted left by 5 spaces, and the value of MaxTbSizeY becomes 32.
[0206] Figure 16 MinTbSizeY disclosed in can be a value indicating the minimum transform size.
[0207] The size of the luminance coding tree block, CtbSizeY, or the size of the coding tree unit, can be variable. For example, when CtbSizeY is less than 64, the minimum transform size can be less than 64. Therefore, when CtbSizeY is less than 64, the value of sps_max_luma_transform_size_64_flag can be 0.
[0208] The CtbSizeY described below in this invention refers to the size of the luminance coding tree block, specifically indicating the width and height of the luminance coding tree block.
[0209] Figure 17 This is a diagram illustrating a higher level of syntax according to an embodiment of the present invention.
[0210] refer to Figure 17 , Figure 17 The syntax structure shown in the diagram can be a higher-level syntax and can include the sps_max_luma_transform_size_64_flag.
[0211] in addition, Figure 17 The syntax structure can include log2_ctu_size_minus5. Based on log2_ctu_size_minus5, the CTU size and size CtbSizeY of the luminance coding tree block can be determined. For example, log2_ctu_size_minus5+5 is CtbLog2SizeY, and CtbLog2SizeY represents log2(CtbSizeY). Additionally, CtbSizeY can be determined as (1 < ... <CtbLog2SizeY)。
[0212] in addition, Figure 17 The syntax structure can include `sps_sbt_enabled_flag` and `sps_sbt_max_size_64_flag`. `sps_sbt_enabled_flag` can be a flag indicating whether Subblock Transformation (SBT) can be used. SBT can transform only some samples of the Cubs or Pus. `sps_sbt_max_size_64_flag` can be a flag indicating the maximum size of the SBT that can be used. (See reference...) Figure 17When `sps_sbt_enabled_flag` indicates that SBT can be used (e.g., when `sps_sbt_enabled_flag` is 1), `sps_sbt_max_size_64_flag` can be signaled. Conversely, when `sps_sbt_enabled_flag` indicates that SBT can be disabled (e.g., when `sps_sbt_enabled_flag` is 0), `sps_sbt_max_size_64_flag` does not need to be signaled. Additionally, SBT can be used when both the width and height of the block are less than the maximum size that allows SBT to be used.
[0213] The maximum size of the SBT that can be used, indicated by `sps_sbt_max_size_64_flag`, can include 32. Alternatively, the maximum size of the SBT that can be used, indicated by `sps_sbt_max_size_64_flag`, can be either 32 or 64. When the maximum transform size is less than the maximum size of the SBT that can be used, indicated by `sps_sbt_max_size_64_flag`, the maximum size of the SBT that can be used can be set to the maximum transform size. (See reference) Figure 17 7-31, the smaller of the maximum transform size MaxTbSizeY and the maximum size of the SBT that can be used, indicated by sps_sbt_max_size_64_flag, can be set as the maximum size of the SBT that can be used, MaxSbtSize.
[0214] Figure 17 The syntax structure can include `sps_transform_skip_enabled_flag`. `sps_transform_skip_enabled_flag` can be a flag indicating whether transform skipping is enabled. In this case, transform skipping can indicate that no transform is performed.
[0215] Figure 18 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0216] Figure 18 The transformation tree syntax can be used to support references Figures 16-17 The syntax for describing the variable maximum transform size.
[0217] When the maximum size of CU or PU is twice the maximum transform size and a fixed maximum transform size is used, it can be used Figures 11 to 13 The syntax and decoding process described in [the document]. Specifically, when the maximum size of the CU or PU is 128 and the maximum transform size is 64, [the following can be used]. Figures 11 to 13An example of this implementation. In this case, if the CU or PU is larger than the maximum transform size and therefore the TU is segmented, the TU can be partitioned into two or four. Within a partition, the TU can be divided into up to two TUs in each of the horizontal and vertical directions.
[0218] However, when using a variable maximum transform size, or when supporting a maximum transform size smaller than the normal maximum transform size, or when the maximum size of the CU or PU is greater than twice the maximum transform size, the TU should be partitioned into two or more sections in both the horizontal and vertical directions. However, there are references... Figures 11 to 13 The described syntax and procedures do not support this partitioning issue.
[0219] Therefore, in order to solve this problem, reference will be made to Figure 18 This describes the syntax for recursively performing the operation of dividing a TU into two parts only in the horizontal or vertical direction. (Reference) Figure 18 When tbWdith is greater than MaxTbSizeY or tbHeight is greater than MaxTbSizeY, it can be called again. Figure 18 The transformation tree syntax is "transform_tree()". In this case, "transform_tree()" can be called twice. Figure 18 (6) and (8) or Figure 18 (6) and (10)). The two calls can be determined based on the first partition direction of TU. Figure 18 (6) and (8) are still Figure 18 (6) and (10). The first partition direction of TU can be determined based on verSplitFirst. The input to "transform_tree()" called based on verSplitFirst can be different. In addition, it can be determined based on verSplitFirst that the two calls are Figure 18 (6) and (8) are still Figure 18 (6) and (10). tbWidth and tbHeight are the width and height of the block, and MaxTbSizeY is the maximum transform size.
[0220] refer to Figure 18 (3) can determine the verSplitFirst value based on tbWidth, tbHeight and MaxTbSizeY.
[0221] In an embodiment of the present invention, when tbWidth is greater than MaxTbSizeY and tbWidth is greater than tbHeight, the verSplitFirst value can be set to 1; otherwise, the verSplitFirst value can be set to 0.
[0222] As Figure 18 The width and height of the input block for "transform_tree()" can be... Figure 18 The trafoWidth and trafoHeight in (6), (8), and (10) can be determined based on verSplitFirst. For example, when verSplitFirst is 1, trafoWidth can be set to tbWidth / 2, and when verSplitFirst is 0, trafoWidth can be set to tbWidth. Additionally, when verSplitFirst is 0, trafoHeight can be set to tbHeight / 2, and when verSplitFirst is 1, trafoHeight can be set to tbHeight. That is, as... Figure 18 The inputs trafoWidth and trafoHeight of (6), (8), and (10) can be the same as either of the existing block width and height, while the other can be half of the existing value. In other words, trafoWidth and trafoHeight can be tbWidth / 2, tbHeight, or tbWidth, tbHeight / 2.
[0223] In an embodiment of the present invention, when tbWidth is greater than or equal to MaxTbSizeY and tbWidth is greater than or equal to tbHeight, verSplitFirst can be set to 1; otherwise, verSplitFirst can be set to 0.
[0224] The width and height of the block can be determined by... Figure 18 Update using (4) and (5). In this case, the width and height of the updated block can be trafoWidth and trafoHeight. Figure 18 The input of (6), i.e., the "transform_tree()" call again, can use trafoWidth and trafoHeight. Figure 18 (6) can be a process that calls "transform_tree()" using coordinates (x0, y0), trafoWidth, and trafoHeight as input. That is to say, Figure 18(6) can be a process that calls "transform_tree()" with the same coordinates as the input of the existing "transform_tree()", and the width (i.e., trafoWidth) or height (i.e., trafoHeight) of the update block as input.
[0225] Based on verSplitFirst, it can be executed Figure 18 (8) or (10). When executing Figure 18 When (8) is executed, "transform_tree()" can be called using coordinates (x0+trafoWidth,y0), trafoWidth, and trafoHeight as input. Figure 18 When (10) is called, "transform_tree()" can be invoked using coordinates (x0, y0 + trafoHeight), trafoWidth, and trafoHeight as input. That is, the call to the existing "transform_tree()" function, except for the one with the coordinates (x0, y0 + trafoHeight), can be used. Figure 18 The remaining blocks outside the block corresponding to (6) (because their coordinates are different) can have their corresponding "transform_tree()" calls... Figure 18 (8) and (10).
[0226] The steps of calling "transform_tree()" can be executed recursively.
[0227] When tbWidth is less than or equal to MaxTbSizeY and tbHeight is less than or equal to MaxTbSizeY, "transform_unit()" can be called.
[0228] Figure 18 In this context, x0, y0, tbWidth, and tbHeight can be values based on luminance samples and can be compared with... Figure 11 The values described are the same. Therefore, a size comparison for determining verSplitFirst can be performed using values based on luminance samples. This is because, when the size comparison is based on chrominance samples, even when the width and height of the transform block (unit) are the same, it is possible that the tbWidth of the chrominance block is greater than tbHeight. For example, when the chrominance format is 4:2:2 sampling, it is possible that tbWidth is greater than tbHeight and tbWidth / SubWidthC is equal to tbHeight / SubHeightC.
[0229] Figure 19This is a diagram illustrating the TU partition according to an embodiment of the present invention.
[0230] Figure 19 The partitions of TU shown in the diagram can indicate according to Figure 18 The syntax for partitioning. For example, when the block width tbWidth is greater than the maximum transform size MaxTbSizeY and the block width tbWidth is greater than the block height tbHeight, such as... Figure 19 As shown, a vertical partitioning method can be applied to divide TU into left (1) TU and right (2) TU. If the value of verSplitFirst is 1, as shown... Figure 18 As described above, TU can be divided into two TUs, left TU and right TU, such as Figure 19 As shown. In this case, the TU can be partitioned again by calling "transform_tree()" again. Figure 19 The diagram illustrates the case where tbHeight is also greater than the maximum transform size MaxTbSizeY. (This can be done...) Figure 19 The left TU ((1)) is segmented, and can be divided within it. Figure 19 The right TU ((2)) is partitioned internally. Figure 18 The syntax, due to the recursive call to "transform_tree()", allows all operations related to partitioning in the left TU((2)) to be parsed or executed after all operations related to partitioning in the left TU((1)) have been parsed or executed. Meanwhile, the value of verSplit can be set to 0 when the width of the block is less than the maximum transform size or the width of the block is less than the height of the block. In this case, the TU can be partitioned into an upper TU and a lower TU by applying horizontal partitioning.
[0231] Figure 20 This is a diagram illustrating the TU partition according to an embodiment of the present invention.
[0232] Figure 20 The TU partition can be based on the reference Figures 11 to 13 The TU partitioning of the described embodiment. For example, when the block width tbWidth is greater than the maximum transform size MaxTbSizeY and the block height tbHeight is greater than MaxTbSizeY, one TU can be partitioned into four TUs at a time, such as... Figure 20 As shown. Additionally, see reference... Figures 11 to 13 In some embodiments, there may be a preset order for performing the decoding process on the TU of a partition. As an example of a preset order, it can be according to... Figure 20 The decoding process is executed in the order of (a), (b), (c), and (d).
[0233] However, due to the corresponding Figure 19 The region of (1) is Figure 20 (a) and (c), and corresponding to Figure 19 The region of (2) is Figure 20 Therefore, there may exist (b) and (d) for... Figure 19 and 20 There is a problem with the TUs of the partitions performing the decoding process in different orders. Furthermore, there is an increase in implementation complexity when the syntax order and the decoding process order are different, and when the block order in the syntax and the block order in the decoding process are different. This is because, during the decoding process of a certain block, the syntax corresponding to the decoding process applied to each block must be applied.
[0234] As in Figure 19 and Figure 20 In the embodiments, the case where the decoding order of blocks differs from one another can be that tbWidth may be greater than MaxTbSizeY and tbWidth may be greater than tbHeight. Specifically, it can be that tbWidth is greater than MaxTbSizeY, tbWidth is greater than tbHeight, and tbHeight is greater than MaxTbSizeY. For example, when MaxTbSize is 32, tbWidth is 128, and tbHeight is 64, the decoding order of blocks can be different.
[0235] Figure 21 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0236] Figure 21 The content disclosed herein relates to intra-frame blocks, but is not limited thereto, and may be applied to other embodiments of TU partitioning, such as residual signal decoding processes.
[0237] Figure 21 The publicly disclosed decoding process is related to Figure 18 The decoding process corresponding to the syntax.
[0238] In order to solve Figure 20 The aforementioned problems described in the embodiments, etc., necessitate changes to the decoding process associated with the TU partition. As stated above, the TU partition described herein can have the same meaning as the TB partition. Figure 21 The publicly disclosed decoding process can be defined as partitioning a TB into two and operating recursively. For example, a TB can be partitioned into two based on verSplitFirst. (See reference...) Figure 21 8-41 and 8-42 can be compared with Figure 12 In the same manner as in the embodiments, maxTbWidth and maxTbHeight are determined. Additionally, Figure 21nTbW and nTbH can be values set based on each color component. maxTbWidth and maxTbHeight indicate the width and height of the maximum transform block, while nTbW and nTbH indicate the width and height of the block.
[0239] For example, when nTbW is greater than maxTbWidth and nTbW is greater than nTbH, verSplitFirst can be set to 1; otherwise, verSplitFirst can be set to 0. Additionally, when verSplitFirst is 1, newTbW can be set to nTbW / 2; otherwise, it can be set to nTbW( Figure 21 (Refer to section 8-44). Additionally, when verSplitFirst is 0, newTbH can be set to nTbH / 2; otherwise, it can be set to nTbH (…). Figure 21 (8-45 in the middle).
[0240] At the same time, if in Figure 18 Different definitions of verSplitFirst in the publicly available syntax allow for different definitions. Figure 21 verSplitFirst. For example, in Figure 18 In the above, when tbWidth is greater than MaxTbSizeY and tbWidth is greater than or equal to tbHeight, verSplitFirst is set to 1; otherwise, if verSplitFirst is set to 0, then... Figure 21 verSplitFirst can be set to 1; otherwise, verSplitFirst can be set to 0 when nTbW is greater than maxTbWidth and nTbW is greater than or equal to nTbH.
[0241] exist Figure 21 In step 2, it can be called again. Figure 21 Clause 8.4.5.1 disclosed in the document. In this case, the coordinates of the sample location (xTb0, yTb0), newTbH, and newTbH can be used as input. When as Figure 19 When partitioning the left TU (1) TU, these inputs can be used for the first (left or top) TU.
[0242] In addition, Figure 21 In step 3 or 4, it can be called again based on verSplitFirst. Figure 21 Chapter 8.4.5.1. When... Figure 19 When partitioning the right TU (2) part of the TU, these inputs can be used for the second (right or down) TU. Figure 21 In step 3, when calling Figure 21In clause 8.4.5.1, the coordinates (xTb0+newTbW,yTb0), newTbH, and newTbH can be used as input. Figure 21 In step 4, when calling Figure 21 According to Clause 8.4.5.1, the coordinates (xTb0, yTb0+newTbH), newTbH, and newTbH can be used as inputs.
[0243] As mentioned above, in Figure 18 In the text, tbWidth and tbHeight can be values based on brightness, and... Figure 21 In this context, nTbW and nTbH can be values based on each color component. Therefore, when performing a block on the chromaticity components... Figure 21 When the decoding process is disclosed in the document, Figure 18 grammar and Figure 21 Mismatches may occur between the decoding processes. For example, the relationship between tbWidth and nTbW can be as follows: In the case of a block with a luma component, tbWidth can be equal to nTbW, and tbHeight can be equal to nTbH. In the case of a block with a chroma component, tbWidth / SubWidthC can be equal to nTbW, and tbHeight / SubHeightC can be equal to nTbH. Therefore, Figure 18 The comparison blocks shown are width and height to determine the result of the verSplit portion and Figure 21 The results of comparing the width and height of the blocks shown in step 1 to determine the parts of the verSplit can differ from each other. For example, there are... Figure 18 The result of tbWidth>tbHeight and Figure 21 The results for nTbW > nTbH may differ, therefore, Figure 18 and 21 The verSplitFirst values are different.
[0244] Figure 22 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0245] Figure 22 The content disclosed herein relates to intra-frame blocks, but is not limited thereto, and can be applied to other embodiments of TB partitioning, such as residual signal decoding processes. Figure 22 The content that is the same as the above will be omitted from the publicly available information.
[0246] Figure 22 The content disclosed in the middle can be related to Figure 18 The decoding process corresponding to the publicly disclosed syntax. Additionally, reference will be made to... Figure 22 Description for reference Figures 20 to 21The decoding process described, particularly the decoding process related to TB segmentation.
[0247] refer to Figure 22 As input to the intra-block decoding process or the residual signal decoding process, it can include the coordinates (xTb0, yTb0) indicating the block's position, nTbW indicating the block's width, nTbH indicating the block's height, cIdx indicating the color components, treeType indicating the tree structure, etc. In this case, nTbW and nTbH can be the width and height of the transform block. That is, nTbW and nTbH can be the width and height based on the color components of the current block. For example, when the chroma subsampling is 4:2:0, if the width and height of the luma block are 16 and 16 respectively, then for the chroma block (when cIdx is not 0), Figure 22 The nTbW and nTbH of the decoding process disclosed herein can be 8 and 8, respectively.
[0248] refer to Figure 22 Section 8-41 states that `maxTbWidth` can be determined based on `cIdx` and `SubWidthC`. For example, when `cIdx` is 0, `maxTbWidth` is `MaxTbSizeY`, while when `cIdx` is not 0, `maxTbWidth` can be determined based on `MaxTbSizeY` and `SubWidthC`. When `cIdx` is not 0, `maxTbWidth` can be set to `MaxTbSizeY / SubWidthC`.
[0249] refer to Figure 22 As shown in section 8-42, maxTbHeight can be determined based on cIdx and subHeightC. For example, when cIdx is 0, maxTbHeight is MaxTbSizeY, while when cIdx is not 0, maxTbHeight can be determined based on MaxTbSizeY and SubHeightC. When cIdx is not 0, maxTbHeight can be set to MaxTbSizeY / SubHeightC.
[0250] If nTbW is greater than maxTbWidth or nTbH is greater than maxTbHeight, other inputs can be used to call the function again. Figure 22 The decoding process is disclosed in the documentation. For example, the decoding process can be called twice. Additionally, recursive calls may occur. When called again... Figure 22 During the publicly disclosed decoding process, TB can be segmented.
[0251] According to embodiments of the present invention, when a TB is partitioned, the TB can be partitioned based on nTbW, nTbH, MaxTbSizeY, cIdx, SubWidthC, and SubHeightC. Alternatively, when a TB is partitioned, the TB can be partitioned based on nTbW, nTbH, maxTbWidth, cIdx, SubWidthC, and SubHeightC. Alternatively, the TB can be partitioned based on the verSplitFirst value.
[0252] The verSplitFirst value can be determined based on cIdx. For example, when cIdx is 0, the verSplitFirst value can be determined based on nTbW, nTbH, and maxTbWidth. When cIdx is 0, the verSplitFirst value can be determined without considering SubWidthC and SubHeightC. Meanwhile, when cIdx is not 0, that is, when cIdx is 1 or 2, the verSplitFirst value can be determined based on nTbW, nTbH, maxTbWidth, SubWidthC, and SubHeightC.
[0253] Additionally, the verSplitFirst value can be determined based on one or more conditions. For example, verSplitFirst can be determined based on whether the block width is greater than the maximum transform size. It can also be determined based on whether the block height is greater than the maximum transform size. In this case, whether the block width and height are greater than the maximum transform size is irrelevant whether the result is based on a luma block or a chroma block. In other words, it does not affect the result regardless of the color component used. Therefore, refer to... Figure 22 The above conditions are checked using input based on each color component. This can be to avoid calculations that are transformed into values based on brightness.
[0254] Additionally, one or more conditions used to determine the verSplitFirst value can be based on the width and height of the block. For example, the verSplitFirst value can be determined based on whether the block width is greater than the block height. In this case, the block width and height can be values for blocks with luma components. In other words, the decoder can determine the verSplitFirst value by comparing the block width and block height based on the values of luma pixels and luma components. That is, when cIdx is 0 (in the case of luma blocks), nTbW and nTbH can be used for one or more of the above conditions, while when cIdx is not 0 (in the case of chroma blocks), nTbW*SubWidthC and nTbH*SubHeightC can be used. (See reference) Figure 22When cIdx is 0, we can check whether nTbW is greater than nTbH. When cIdx is not 0, we can check whether nTbW*SubWidthC is greater than nTbH*SubHeightC.
[0255] In this case, the verSplitFirst value can be set to 1 when one or more of the above conditions are met; otherwise (when at least one of the one or more conditions is not met), the verSplitFirst value can be set to 0.
[0256] In other words, the decoder can use values that take into account color components and chroma format when comparing the width of a block with the maximum transform size, and can use values based on the luminance component when comparing the width of a block with the height of a block.
[0257] If we express the above reference through an equation Figure 22 The conditions described for determining the value of verSplitFirst can be expressed as Equation 2 below.
[0258] [Equation 2]
[0259] verSplitFirst=(nTbW>maxTbWidth&&(cldx==0?nTbW:nTbW*SubWidthC)>(cldx==0?nTbH:nTbH*SubHeightC))? 1:0
[0260] Referring again to Equation 2, when cIdx is 0, verSplitFirst can be determined in Equation 3 below.
[0261] [Equation 3]
[0262] verSplitFirst=(nTbW>maxTbWidth&&nTbW>nTbH)? 1:0
[0263] Referring again to Equation 2, when cIdx is not 0 (for example, when cIdx is 1 or 2), verSplitFirst can be determined as in Equation 4 below.
[0264] [Equation 4]
[0265] verSplitFirst=(nTbW>maxTbWidth&&nTbW*SubWidthC>nTbH*SubHeightC)? 1:0
[0266] When comparing the sizes in equations 2 through 4, the equal sign (=) can be included. In other words, it is possible to check if nTbW is greater than or equal to maxTbWidth. When cIdx is 0, it is possible to check if nTbW is greater than or equal to nTbH, and when cIdx is not 0, it is possible to check if nTbW*SubWidthC is greater than nTbH*SubHeightC. In this case, the equal sign (=) can also be included when comparing sizes. Figure 18 In the syntax.
[0267] refer to Figure 22 In sections 8-44 and 8-45, newTbW and newTbH can be determined based on the verSplitFirst value. For example, when verSplitFirst is 1, newTbW can be determined as nTbW / 2, and newTbH can be determined as nTbH. When verSplitFirst is 0, newTbW can be determined as nTbW, and newTbH can be determined as nTbH / 2.
[0268] refer to Figure 22 Based on verSplitFirst, it is possible to... Figure 22 The decoding process can be called again in steps 2 and 3, or it can be called again in steps 2 and 4.
[0269] exist Figure 22 In step 2, the decoding process can be invoked again using the block position (xTb0, yTb0), newTbW, and newTbH. This can be for applications such as... Figure 19 The first (left or top) TU of the partition of the left TU (1) part.
[0270] exist Figure 22 In step 3, the decoding process can be invoked again using the block positions (xTb0+newTbW, yTb0), newTbW, and newTbH. This can be used for... Figure 19 The second (right) TU of the partition of the right TU (2) part.
[0271] exist Figure 22 In step 4, the decoding process can be invoked again using the block positions (xTb0, yTb0+newTbH), newTbW, and newTbH. This can be used for... Figure 19 The right TU (2) part of the partition of the second (lower) TU.
[0272] Figure 23 This is a diagram illustrating a higher-level syntax according to an embodiment of the present invention.
[0273] The syntax element indicating the maximum transform size can be based on the CTU size (CtbSizeY) or a syntax element indicating CtbSizeY. For example, a maximum transform size (CtbSizeY) greater than the CTU size and a size indicated by a syntax element indicating CtbSizeY can be excluded from the candidates. When the CTU size, CtbSizeY, and the size indicated by the syntax element indicating CtbSizeY are all less than 64, values greater than or equal to 64 can be excluded from the maximum transform size candidates. When two possible values exist for the maximum transform size, the maximum transform size can be determined based on the CTU size, CtbSizeY, and the size indicated by the syntax element indicating CtbSizeY, without the need for signaling indicating the maximum transform size, and the value of the maximum transform size can be determined. The CTU size refers to the width and height of the coding tree unit, and CtbSizeY indicates the size of the luma coding tree block and the width and height of the chroma coding tree block.
[0274] refer to Figure 23 The syntax element indicating the maximum transform size can be `sps_max_luma_transform_size_64_flag`. This syntax element can represent the maximum transform size as 64 or a value less than 64. If `CtbSizeY` is less than 64, `sps_max_luma_transform_size_64_flag` can be inferred without signaling. For example, if the maximum transform size indicates 64 when `sps_max_luma_transform_size_64_flag` is 1, then when `CtbSizeY` is less than 64, the value of `sps_max_luma_transform_size_64_flag` can be inferred to be 0. Alternatively, the value of `MaxTbLog2SizeY` can be determined to be 5, or the value of `MaxTbSizeY` can be determined to be 32.
[0275] If the minimum CtbSizeY is 32, then the case where CtbSizeY is less than 64 can be handled in the same way as the case where CtbSizeY is 32. For example, if CtbSizeY is 32, the value of sps_max_luma_transform_size_64_flag can be inferred to be equal to 0, and therefore the maximum transform size (MaxTbSizeY) value can be determined to be 32.
[0276] Figure 24 This is a diagram illustrating a higher-level syntax according to an embodiment of the present invention.
[0277] If the maximum transform size is limited, the maximum size of the subblock transform (SBT) that can be used is also limited accordingly. For example, the maximum size of the SBT that can be used can be less than or equal to the maximum transform size. Therefore, the syntax element indicating the maximum size of the SBT that can be used can vary. For example, a syntax element indicating the maximum size of the SBT that can be used can be signaled based on either the syntax element indicating the maximum transform size or the maximum transform size.
[0278] refer to Figure 24 When the maximum transform size is less than 64, no signal is needed to indicate whether the maximum size of the SBT is 64, and the maximum size of the SBT can be inferred. The syntax element indicating that the maximum size of the SBT is 64 can be sps_sbt_max_size_64_flag. (See reference) Figure 24 When both `sps_sbt_enabled_flag` and `sps_max_luma_tranform_size_64_flag` are 1 (i.e., `sps_sbt_enabled_flag && sps_max_luma_tranform_size_64_flag`), `sps_sbt_max_size_64_flag` is signaled; otherwise, it is not signaled. Additionally, when `sps_sbt_max_size_64_flag` does not exist, its value can be inferred to be 0. Alternatively, when `sps_sbt_max_size_64_flag` does not exist, its value can be inferred to be the value of `sps_max_luma_tranform_size_64_flag`. In this case, as in... Figure 24 As disclosed in 7-31, when sps_sbt_max_size_64_flag is 1, MaxSbtSize can be set to 64, and when sps_sbt_max_size_64_flag is 0, MaxSbtSize can be set to 32.
[0279] Figure 25 This is a diagram illustrating the transformation tree syntax according to an embodiment of the present invention.
[0280] exist Figure 25 In the middle, the method used to determine the reference was changed. Figure 18 The syntax describes the condition for the verSplitFirst value, and the remaining identical description will be omitted.
[0281] The TU partitioning method can vary based on the tree type (treeType). For example, when the tree type is DUAL_TREE_CHROMA, TU partitioning can be performed on a width and height basis, based on blocks with each color component (i.e., chroma component). Figure 18 In the embodiments, using Figure 18 The conditions disclosed in (3) determine verSplitFirst, but in the case of DUAL_TREE_CHROMA, the following can be used: Figure 25 The conditions disclosed in (2) are used to determine verSplitFirst.
[0282] When determining the width of a block and the maximum transform size, it is irrelevant whether this determination is based on luminance or on each color component; therefore, luminance-based values can be used as input to the transform tree syntax "transform_tree()". However, when comparing the width and height of a block, values based on the chrominance components can be used. That is, tbWidth / SubWidthC and tbHeight / SubHeightC can be compared. Specifically, it can be checked whether tbWidth / SubWidthC is greater than tbHeight / SubHeightC.
[0283] Therefore, when the tree type is dual-tree, TU partitioning for the chroma component can be performed on the same basis as when the luminance component is used for TU partitioning.
[0284] refer to Figure 25 When the tree type is SINGLE_TREE or DUAL_TREE_LUMA, the decoder can execute... Figure 25 Step (1) determines verSplitFirst. That is, partitions can be determined using values based on the luminance components. When the tree type is DUAL_TREE_CHROMA, the decoder can perform... Figure 25 Step (2) determines verSplitFirst. That is, the partition can be determined using values based on the chroma components.
[0285] Figure 26 This is a diagram illustrating the decoding process according to an embodiment of the present invention.
[0286] Figure 26 The embodiments of the decoding process disclosed herein can be related to... Figure 25 The decoding process corresponds to the syntax disclosed in the document. As mentioned above, the TU partition described in this document can have the same meaning as the TB partition.
[0287] The method for determining the TB partition or verSplitFirst can be based on cIdx, treeType, etc. For example, when cIdx is 0, verSplitFirst can be determined by using the width and height of the block based on each color component. When the tree type is DUAL_TREE_CHROMA, verSplitFirst can be determined by using the width and height of the block based on each color component. When cIdx is not 0 and the tree type is not DUAL_TREE_CHROMA (i.e., cIdx is not 0 and the tree type is SINGLE_TREE), verSplitFirst can be determined by using the width and height of the block based on the luminance component.
[0288] That is, when cIdx is 0 or the tree type is DUAL_TREE_CHROMA (cIdx == 0 || treeType == DUAL_TREE_CHROMA), verSplitFirst can be determined by comparing whether nTbW is greater than nTbH. In other words, verSplitFirst can be determined without using SubWidthC and SubHeightC. When cIdx is not 0 and the tree type is not DUAL_TREE_CHROMA (when cIdx != 0 && treeType != DUAL_TREE_CHROMA), verSplitFirst can be determined by comparing whether nTbW / SubWidthC is greater than nTbH / SubHeightC.
[0289] The maximum transition skip size can be smaller than the maximum transition size. For example, if the maximum transition size is variable, the maximum transition skip size can change accordingly. Therefore, when the maximum transition size is variable, the method of sending the maximum transition skip size using a signal can change accordingly.
[0290] The maximum transform size and the maximum transform skip size can be the same. In this case, when the size of the CU or PU is greater than the maximum transform size, the signaling indicating whether transform skip is used in the partitioned TU can be shared. In other words, when the size of the CU or PU is greater than the maximum transform size, the signaling of the syntax element indicating whether transform skip is used in the partitioned TU can be partially omitted. This is because when the size of the CU or PU is greater than the maximum transform size, the residual signal shape of the partitioned TU can be similar. Sharing the signaling of whether transform skip is used in the partitioned TU or omitting the relevant syntax element can be limited to the case where the maximum transform size and the maximum transform skip size are the same when using variable maximum transform size and variable maximum transform skip size. Sharing whether transform skip is used in the TU of the additional partition can be limited to the case where transform skip is used in the TU used as a reference (e.g., the first TU). That is, when transform skip is not used in the TU used as a reference, the signal indicating whether transform skip is used can be sent to the TU of other partitions. In addition, if the signal indicating whether transform skip is used is sent to the TUs partitioned in a preset order and the signal indicating that transform skip is used is sent to a specific TU, then the signal indicating whether transform skip is used is not sent to subsequent TUs, and it can be determined that transform skip is used.
[0291] When signaling regarding whether to use the syntax element for transformation skipping at the CU or PU level is used for shared transform skipping in the TU of partitioning or when omitting related syntax elements, it is possible to signal whether to use the syntax element for transform skipping at the CU or PU level.
[0292] Figure 27 This is a diagram illustrating a method for performing BDPCM according to an embodiment of the present invention.
[0293] Block-based incremental pulse code modulation (BDPCM) can be either an intra-frame prediction method or a coding method. Furthermore, BDPCM can possess unique characteristics compared to prediction methods and residual signal generation methods. That is, BDPCM can be a method of encoding and transmitting values based on the difference between specific values (e.g., residual signals or prediction signals, image samples or reconstructed samples). For example, an encoder can transmit values based on the difference between specific values, and a decoder can reconstruct the specific values based on the signaling. In this case, during the decoder's reconstruction process, since the process of calculating the difference between specific values should be performed in reverse, the specific values can be reconstructed by performing an addition process based on the values transmitted with the signal. BDPCM can be described as quantized residual differential pulse code modulation (RDPCM).
[0294] A syntax element indicating whether BDCPM is used may exist. For example, the syntax element could be `BdpcmFlag` or `intra_bdpcm_flag`. Signaling indicating whether BDCPM is used can be executed at the CU level. Additionally, a syntax element at a higher level indicating whether BDCPM can be used may exist. For example, a higher-level syntax element indicating whether BDCPM is used could be `sps_bdpcm_enabled_flag`. Furthermore, the higher level at which the signaled syntax element is sent can be the SPS level, tile level, patch level, or patch group level. When signaling indicating whether BDCPM is used at a higher level is possible, the signaling of the syntax element indicating whether BDCPM is used may not be executed. Conversely, when signaling indicating whether BDCPM is used at a higher level is not possible, the signaling of the syntax element indicating whether BDCPM is used may not be executed.
[0295] The prediction mode of BDPCM can be restricted. The syntax element indicating the prediction mode of BDPCM can be either `BdpcmDir` or `intra_bdpcm_dir_flag`. When using BDPCM, the prediction mode can be an angular mode. Specifically, when using BDPCM, the prediction mode can be a horizontal mode (mode 18; `INTRA_ANGULAR18`) or a vertical mode (mode 50; `INTRA_ANGULAR50`). Therefore, when using BDPCM, the value of a predicted sample at a certain location can be the same as the value of the reference sample to the left of that location in the horizontal mode. Additionally, when using BDPCM in the horizontal mode, predicted samples corresponding to the same row (i.e., predicted samples with the same y-coordinate) can have the same value. Furthermore, when using BDPCM in the vertical mode, the value of a predicted sample at a certain location can be the same as the value of the reference sample above that location in the vertical mode. Additionally, when using BDPCM in the vertical mode, predicted samples corresponding to the same column (i.e., predicted samples with the same x-coordinate) can have the same value.
[0296] Figure 27 (a) is a diagram illustrating the derivation of the prediction model of BDPCM. (Reference) Figure 27 (a) The prediction mode derivation of BDPCM can be based on BdpcmDir. When in Figure 27In (a), a value of 1 for BdpcmFlag likely indicates the use of BDPCM. IntraPredModeY can be a value indicating the prediction mode. Specifically, IntraPredModeY can be a value indicating the prediction mode used for the luminance component. For example, IntraPredModeY can be set to INTRA_ANGULAR50 or INTRA_ANGULAR18 based on BdpcmDir. When BdpcmDir is 1, IntraPredModeY is set to INTRA_ANGULAR50 and can be in vertical mode. Conversely, when BdpcmDir is 0, IntraPredModeY is set to INTRA_ANGULAR18 and can be in horizontal mode. Prediction operations can be performed based on IntraPredmodeY.
[0297] There may be residual signal generation methods or scaling and transformation methods for BDPCM. As described above, when using BDPCM, values based on the difference between specific values can be encoded and transmitted as signals. For example, when using BDPCM, values based on the difference between the residual signal and the value based on the residual signal can be encoded and transmitted as signals. The value transmitted as a signal for a sample corresponding to a first position can be the difference between the value at a second position and the value at the first position. Specifically, the value transmitted as a signal for a sample corresponding to a first position can be the difference between the value at the first position and the value at a second position, where the second position is adjacent to the first position. The value transmitted as a signal for a sample corresponding to the first position (x, y) can be the difference between the value at the second position (x-1, y) and the value at the first position in the horizontal mode. Additionally, the value transmitted as a signal for a sample corresponding to the first position (x, y) can be the difference between the value at the second position (x, y-1) and the value at the first position in the vertical mode. In this case, the value at the second position and the value at the first position can be the residual signal or a value based on the residual signal.
[0298] Figure 27 (b) is a diagram illustrating the operations in the decoder when using BDPCM. Reference Figure 27 (b) dz can be an array created based on values sent via signals using syntax elements. Additionally, in horizontal mode, the value of dz[x][y] can be set based on dz[x][y] and dz[x-1][y]. For example, in horizontal mode, the value of dz[x][y] can be set based on (dz[x-1][y]+dz[x][y]) (see reference). Figure 27(8 - 961) in (b). This setting can be executed when x is greater than 0 (not tangent to the left boundary of the block). When x is 0, dz[x][y] can remain as dz[x][y]. In vertical mode, the value of dz[x][y] can be set based on dz[x][y] and dz[x][y - 1]. In vertical mode, the value of dz[x][y] can be set based on (dz[x][y - 1]+dz[x][y]) ( Figure 27 (8 - 962) in (b). This setting can be executed when y is greater than 0 (not tangent to the upper boundary of the block). When y is 0, dz[x][y] can remain as dz[x][y]. The reason for calculating dz based on the summation (+) operation is because the signal transmitted by signaling is based on the difference (-) operation. For example, the signal dz[x][y] transmitted by signaling can be the value set to (dz[x][y]-dz[x - 1][y]) in horizontal mode and can be the value set to (dz[x][y]-dz[x][y - 1]) in vertical mode. Therefore, in order to reconstruct this, the decoder can perform the calculation based on the summation operation. Additionally, based on the value of dz[x][y], dnc[x][y] can be derived, and d[x][y] can be derived ( Figure 27 (8 - 963 and 8 - 964) in (b).
[0299] In an embodiment of the present invention, a transform can be performed based on d[x][y]. In this case, performing the transform can include performing a transform skip mode. That is, performing the transform can include not performing an operation based on a transform matrix. When using BDPCM, the transform skip mode can always be used. When using BDPCM, the value of transform_skip_flag, which is a syntax element indicating whether the transform skip mode is used, can always be 1. When using the transform skip mode, the residual signal can be determined without the step of multiplying d[x][y] by a vector or matrix. When using the transform skip mode, the residual signal can be determined without performing an operation on d[x][y] other than shifting or rounding. For example, when using the transform skip mode, the value of (d[x][y]<<tsShift) can be r[x][y], which is the value of the residual sample array. Here, tsShift can be a value determined based on the width and height of the transform block. Additionally, the intermediate residual sample and the residual sample can be determined based on r[x][y]. Additionally, BDPCM can only be used when the transform skip mode is applied. This is because when using BDPCM, the transform skip mode can be used. For example, only when the syntax element signaled at a higher level indicates that the transform skip mode can be used, can the signaling of the syntax element at a higher level indicating that BDPCM can be used be performed.
[0300] Figure 28This is a diagram illustrating the syntax related to BDPCM according to an embodiment of the present invention.
[0301] BDPCM can be used when a syntax element indicating its availability is sent via a signal at a higher level. (See reference.) Figure 28 Whether to parse intra_bdpcm_flag can be determined based on sps_bdpcm_enabled_flag.
[0302] Additionally, the size of the block that can be skipped using transforms may be limited. For example, the maximum value of the width or height of the block that can be skipped using transforms may be MaxTsSize. Therefore, transform skipping can be used when the width of the block is less than or equal to MaxTsSize and the height of the block is less than or equal to MaxTsSize. On the other hand, if the width of the block is greater than MaxTsSize or the height of the block is greater than MaxTsSize, transform skipping can be avoided. Since transforms are performed at the TU level, the width and height of the block can be the width and height of the transformed block. In this case, the width of the transformed block can be described as tbWidth, and the height of the transformed block can be described as tbHeight. Since BDPCM uses transform skipping, whether BDPCM can be used can be determined based on MaxTsSize. Alternatively, whether BDPCM can be used can be signaled at the CU level. Therefore, whether BDPCM can be used can be determined based on the width and height of the coded block (CB) (which can be described as cbWidth and cbHeight) and MaxTsSize. For example, BDPCM can be used when cbWidth is less than or equal to MaxTsSize and cbHeight is less than or equal to MaxTsSize. BDPCM may not be usable when cbWidth is greater than MaxTsSize or cbHeight is greater than MaxTsSize. (See reference) Figure 28 When cbWidth <= MaxTsSize and cbHeight <= MaxTsSize, intra_bdpcm_flag can be parsed; when cbWidth > MaxTsSize or cbHeight > MaxTsSize, intra_bdpcm_flag does not need to be parsed. In this case, if intra_bdpcm_flag does not exist, its value can be inferred to be 0.
[0303] According to embodiments of the invention, MaxTsSize can be selected and used as any of a plurality of values. For example, MaxTsSize can be a value of 32 or less. Alternatively, MaxTsSize can be a value less than or equal to MaxTbSizeY. Alternatively, MaxTsSize can be one of 32, 16, 8, or 4. Additionally, syntax elements for determining MaxTsSize can exist. For example, a value based on log2(MaxTsSize) can be signaled. log2_transform_skip_max_size_minusN can be signaled, and MaxTsSize can be (1<<(log2_transform_skip_max_size_minusN+N)). Specifically, N can be 2, in which case log2_transform_skip_max_size_minus2 is signaled, and MaxTsSize can be (1<<(log2_transform_skip_max_size_minus2+2)). In this case, log2_transform_skip_max_size_minus2 can be a value ranging from 0 to 3, and MaxTsSize can be one of 4, 8, 16, and 32. The syntax element used to determine MaxTsSize can be signaled at a higher level. For example, the syntax element can be signaled at the image parameter set or slice level.
[0304] Figure 29 This is a diagram illustrating the conditions under which BDPCM can be used according to an embodiment of the present invention.
[0305] refer to Figure 28BDPCM can be used when the width cbWidth and height cbHeight of the coded block are less than or equal to MaxTsSize. If MaxTbSizeY is greater than MaxTsSize and the width and height of the coded block are less than or equal to MaxTsSize, then the width and height of the coded block can be the same as the width and height of the transform block. This is because TU partitioning does not occur in this case. Specifically, when MaxTbSizeY is 64 and MaxTsSize is 32 or less, if cbWidth is MaxTsSize or less and cbHeight is MaxTsSize or less, then cbWidth can be equal to tbWidth and cbHeight can be equal to tbHeight. Therefore, the fact that BDPCM can be used when cbWidth and cbHeight are MaxTsSize or less can mean that BDPCM can be used when tbWidth and tbHeight are MaxTsSize or less.
[0306] However, according to embodiments of the invention, MaxTbSizeY can be less than or equal to MaxTsSize. In this case, if cbWidth is greater than MaxTsSize or cbHeight is greater than MaxTsSize, BDPCM can be used. This is because, in this case, TU partitioning may occur, and the width and height of the transform block after the TU partitioning are less than or equal to MaxTsSize. When MaxTbSizeY is equal to MaxTsSize, if cbWidth is greater than MaxTsSize or cbHeight is greater than MaxTsSize, TU partitioning may occur, and the width and height of the transform block after the TU partitioning can be equal to MaxTbSizeY. Therefore, a transform skip mode, which is the mode used in the case of BDPCM, can be executed. This is because the transform skip mode is executed at the TU level. When BDPCM is available, the intra_bdpcm_flag can be resolved.
[0307] Therefore, BDPCM can be used even when MaxTbSizeY and MaxTsSize are the same. Alternatively, BDPCM can be used when a syntax element indicating whether BDPCM is usable, signaled at a higher level, indicates that BDPCM is usable and MaxTbSizeY and MaxTsSize are the same. Alternatively, BDPCM can be used even when MaxTbSizeY and MaxTsSize are different, when cbWidth is less than or equal to MaxTsSize and cbHeight is less than or equal to MaxTsSize. Alternatively, BDPCM can be used when cbWidth is less than or equal to MaxTsSize and cbHeight is less than or equal to MaxTsSize, signaled at a higher level, indicates that BDPCM is usable and MaxTbSizeY and MaxTsSize are different. On the other hand, when MaxTbSizeY and MaxTsSize are different, when cbWidth is greater than MaxTsSize or cbHeight is greater than MaxTsSize, BDPCM can be omitted.
[0308] When BDPCM is available, the intra_bdpcm_flag can be parsed as a syntax element indicating whether BDPCM is used.
[0309] As mentioned above, MaxTbSizeY can have a value of 32 or larger. Specifically, MaxTbSizeY can have a value of 32 or 64. Additionally, MaxTsSize can have a value of 32 or smaller. Specifically, MaxTsSize can have a value of 4, 8, 16, or 32. Therefore, when MaxTbSizeY and MaxTsSize are the same, MaxTbSizeY can be 32 and MaxTsSize can be 32.
[0310] Figure 29 (a) and (b) are diagrams illustrating the syntax structures related to the above embodiments.
[0311] refer to Figure 29 (a) The intra_bdpcm_flag can be parsed when the value of sps_bdpcm_enabled_flag is 1 and at least one of the following two conditions is met.
[0312] Condition a-1) cbWidth<=MaxTsSize&&cbHeight<=MaxTsSize
[0313] Condition a-2)MaxTbSizeY==32&&MaxTsSize==32
[0314] On the other hand, if conditions a-1 and a-2 are not met, then intra_bdpcm_flag does not need to be parsed. Also, when intra_bdpcm_flag does not exist, its value can be inferred to be 0.
[0315] refer to Figure 29 (b) The intra_bdpcm_flag can be parsed when the value of sps_bdpcm_enabled_flag is 1 and at least one of the following two conditions is met.
[0316] Condition b-1)cbWidth<=MaxTsSize&&cbHeight<=MaxTsSize
[0317] Condition b-2) MaxTbSizeY==MaxTsSize
[0318] Furthermore, if conditions b-1 and b-2 are not met, then intra_bdpcm_flag does not need to be parsed. Also, if intra_bdpcm_flag does not exist, its value can be inferred to be 0.
[0319] Figure 30 This is a diagram illustrating CIIP and intra-frame prediction according to an embodiment of the present invention.
[0320] CIIP is an abbreviation for Combined Inter-Frame and Intra-Frame Prediction, or CIIP for Combined Inter-Picture Merging and Intra-Picture Prediction. CIIP refers to a method of combining intra-frame and inter-frame prediction signals when generating prediction signals. When using CIIP, the intra-frame or inter-frame prediction method may be limited. For example, when using CIIP, only the planar mode (MODE_PLANAR) or only merging mode can be used for intra-frame prediction. When CIIP is used for a coding block (CB) or coding unit (CU), intra-frame prediction can be performed on the entire coding block. That is, when CIIP is used for a CB or CU, intra-frame prediction can be performed on a block with a size of (cbWidth x cbHeight). In other words, when CIIP is used for a CB or CU, intra-frame prediction can be performed without TU partitioning. Even if the CB is larger than the maximum transform size, intra-frame prediction can be performed without TU partitioning when using CIIP.
[0321] The block size that can use CIIP can be limited. In this case, the limited block size can be a fixed value. The maximum block size that can use CIIP can be predetermined. For example, CIIP can be avoided when the block width cbWidth or height cbHeight is 128 or greater. CIIP can be used when cbWidth is less than 128, cbHeight is less than 128, and cbWidth*cbHeight is 64 or greater.
[0322] Figure 30 The diagram illustrates CIIP mode and intra-frame prediction mode when the CB or CU is 64×64. When using CIIP mode, inter-frame prediction and intra-frame prediction for 64×64 blocks can be performed. Additionally, 64×64 prediction blocks can be generated based on inter-frame and intra-frame predictions. Figure 30 "Combined inter-frame and intra-frame prediction blocks").
[0323] When intra-prediction mode is not used (i.e., when only the current image is used, or when no reference image is used, or when CuPredMode is MODE_INTRA), the TB to be partitioned from the 64×64 CB can be determined based on MaxTbSizeY. For example, when MaxTbSizeY is less than 64 (e.g., 32), the 64×64 CB can be partitioned into 32×32 TUs, and intra-prediction can be performed on each of the partitioned TUs. Thus, four prediction blocks, each 32×32, can be generated.
[0324] In this scenario, there may be a misalignment between intra-prediction in CIIP mode and intra-prediction in intra-prediction mode. In other words, when using CIIP mode, intra-prediction can be performed on blocks larger than 32×32, and intra-prediction in intra-prediction mode can be performed on blocks smaller than or equal to 32×32. Therefore, in intra-prediction mode, hardware and software capable of handling intra-prediction for blocks smaller than or equal to 32×32 are required; however, in CIIP mode, hardware and software capable of handling intra-prediction for blocks larger than 32×32 are required. This can be a significant implementation burden. For example, some encoders aim to limit the maximum transform size to 32 instead of 64 to perform intra-prediction on small blocks and reduce hardware and software burden, but may still need to prepare intra-prediction for blocks larger than 32 to use CIIP mode.
[0325] Figure 31 This is a diagram illustrating the merged data syntax according to an embodiment of the present invention.
[0326] According to embodiments of the present invention, the grouping method can be used as a mode merging signaling method. For example, a group_1_flag can be sent using a signal, and the decoder can determine whether the selected mode belongs to group 1 based on the group_1_flag. When group_1_flag indicates that it does not belong to group 1, a group_2_flag can be sent using a signal. Additionally, the decoder can determine whether the selected mode belongs to group 2 based on the group_2_flag. This operation can be performed even when multiple groups exist. Furthermore, signaling indicating modes within a group can exist. In the grouping method, the signaling depth can be reduced compared to the sequential signaling method. Additionally, the maximum length of the signaling (e.g., the maximum length of the codeword) can be reduced.
[0327] The grouping method will be described in detail below.
[0328] First, assume there are three groups. A particular group can contain more than one mode. For example, there can be a mode included in group 1. Group 1 can include subblock merging modes, group 2 can include regular merging modes as well as merging with motion vector difference (MMVD), and group 3 can include CIIP and triangular merging modes. `group_1_flag` can be `merge_subblock_flag`, and `group_2_flag` can be `regular_merge_flag`. Additionally, `ciip_flag` and `mmvd_merge_flag` can exist as syntax elements indicating modes within a group. In an embodiment, `merge_subblock_flag` is signaled, and it can be used to determine whether the current mode is a subblock merging mode. In this case, if the current mode is not a subblock merging mode, `regular_merge_flag` can be signaled. The decoder can use `regular_merge_flag` to determine whether the mode is included in group 2 (regular merging mode or MMVD) or group 3 (CIIP or triangular merging mode). In this scenario, when `regular_merge_flag` indicates group 2, the current mode (whether it's regular merge mode or MMVD) can be determined based on `mmvd_merge_flag`. Conversely, when `regular_merge_flag` indicates group 3, the current mode (whether it's CIIP or triangular merge mode) can be determined based on `ciip_flag`.
[0329] refer to Figure 31When using merge mode, the `merge_subblock_flag` can be signaled. Implementations using merge mode can be the same as described above, and can be cases where `general_merge_flag` is 1. Furthermore, this invention can be applied to cases where `CuPredMode` is not `MODE_IBC` or `CuPredMode` is `MODE_INTER`. The decoder can determine whether to parse `merge_subblock_flag` based on `MaxNumSubblockMergeCand` or the block size. When `merge_subblock_flag` is 1, the decoder can determine to use subblock merge mode and can additionally determine candidate indices based on `merge_subblock_idx`. When `merge_subblock_flag` is 0, the decoder can parse `regular_merge_flag`. In this case, there may be conditions that allow parsing `regular_merge_flag`. For example, there may be conditions based on the block size. Additionally, there may be conditions based on syntax elements that indicate whether a mode is available, signaled at a higher level. In this context, syntax elements indicating whether a mode is available, signaled at a higher level, can include `sps_ciip_enabled_flag` and `sps_triangle_enabled_flag`. In this case, `sps_triangle_enabled_flag` can be a syntax element signaled at a higher level indicating whether the decoder can use the triangular mode, and can be signaled from the sequence parameter set. Conditions based on slice type may also exist. Additionally, conditions based on `cu_skip_flag` may exist. In this case, `cu_skip_flag` can be a syntax element indicating whether a skip mode can be used. When using skip mode, residual signals can be sent without signaling. Therefore, when using skip mode, the decoder can reconstruct blocks from the predicted signals without residual signals (or transform coefficients).
[0330] Conditions related to the block size that can use CIIP include a block width multiplied by a block height of 64 or greater, a block width less than 128, and a block height less than 128. Additionally, the block size condition for using the triangle merge mode is a block width multiplied by a block height of 64 or greater. (See reference) Figure 31 The decoder may not parse `regular_merge_flag` if the condition that `block width * block height` is 64 or greater is not met. Conversely, the decoder may parse `regular_merge_flag` if the block width is equal to (or greater than or equal to) 128 or the block height is equal to (or greater than or equal to) 128.
[0331] Specifically, i) when the value of `sps_ciip_enabled_flag` is 1, the value of `cu_skip_flag` is 0, the block width is less than 128, and the block height is less than 128, the decoder can parse `regular_merge_flag` if the condition that the block width * block height is 64 or greater is met. Or ii) when the value of `sps_triangle_enabled_flag` is 1, the value of `MaxNumTriangleMergeCand` is greater than 1, and `slice_type` is B, the decoder can parse `regular_merge_flag` if the condition that the block width * block height is 64 or greater is met.
[0332] Meanwhile, when the conditions for allowing the parsing of `regular_merge_flag` are not met—namely, i) when `sps_ciip_enabled_flag` is 1, `cu_skip_flag` is 0, the block width is less than 128, and the block height is less than 128; and ii) when `sps_triangle_enabled_flag` is 1, `MaxNumTriangleMergeCand` is greater than 1, and `slice_type` is B—the decoder may not parse `regular_merge_flag`. In this case, `MaxNumTriangleMergeCand` can be the maximum number of candidates that can be used in the triangular (merging) mode. Additionally, `slice_type` can be a signaling indication of the slice type. A `slice_type` of B indicates that double prediction can be used. A `slice_type` of P indicates that double prediction can be disabled and single prediction can be used. A `slice_type` of P or B indicates that inter-frame prediction can be used. A `slice_type` of I indicates that inter-frame prediction can be disabled.
[0333] Additionally, refer to Figure 31 When the decoder determines whether to parse the `ciip_flag`, it can use a condition based on the block size. For example, if the block width is less than 128 and the block height is less than 128, the decoder can parse the `ciip_flag`. On the other hand, if the block width is 128 (or 128 or greater) or the block height is 128 (or 128 or greater), the decoder can choose not to parse the `ciip_flag`. This is because when the block width or block height is 128 (or 128 or greater), only one of the `CIIP` and triangle merge modes can be used. For example, `CIIP` can be omitted, and only the triangle merge mode can be used.
[0334] In this invention, the fact that ciip_flag can be parsed can mean that CIIP can be used. Alternatively, the fact that ciip_flag can be parsed can mean that CIIP or the triangular merge mode can be used.
[0335] Conditions for using CIIP can include a value of 1 for `sps_ciip_enabled_flag` and a value of 0 for `cu_skip_flag`. Additionally, conditions related to the block size for which CIIP can be used can include a block width multiplied by a block height of 64 or greater, a block width less than 128, and a block height less than 128.
[0336] Conditions for enabling the triangle merge mode can include a value of 1 for `sps_triangle_enabled_flag`, a value greater than 1 for `MaxNumTriangleMergeCand`, and a `slice_type` of B. Additionally, conditions related to the block size for which the triangle merge mode can be used can include a block width multiplied by a block height of 64 or greater.
[0337] The decoder can parse `regular_merge_flag` when either of the conditions for using CIIP or the conditions for using the triangular merge mode are met. Conversely, if neither the conditions for using CIIP nor the conditions for using the triangular merge mode are met, the decoder may choose not to parse `regular_merge_flag`.
[0338] When `regular_merge_flag` does not exist, its value can be inferred to be 1. In this invention, when `regular_merge_flag` is 1, either the regular merge mode or MMVD can be used. Therefore, when neither the conditions related to the block size for using CIIP nor the conditions related to the block size for using the triangular merge mode are met, the usable mode can be either the regular merge mode or MMVD. In this case, `regular_merge_flag` can be set to 1 without resolution.
[0339] Additionally, when neither the conditions for using CIIP nor the conditions for using the triangular merge mode are met, the mode that can be used is the regular merge mode or MMVD. Therefore, regular_merge_flag can be inferred to be equal to 1 and no parsing is performed.
[0340] refer to Figure 31When `regular_merge_flag` is 1, syntax elements can be resolved based on the value of `sps_mmvd_enabled_flag`. `sps_mmvd_enabled_flag` can be a syntax element signaled at a higher level indicating whether MMVD is usable. When `sps_mmvd_enabled_flag` is 0, MMVD is not used. (See reference) Figure 31 When the value of `sps_mmvd_enabled_flag` is 0, `mmvd_merge_flag`, `mmvd_cand_flag`, `mmvd_distance_idx`, and `mmvd_direction_idx` do not need to be parsed. When `mmvd_merge_flag` does not exist, its value can be inferred to be equal to 0.
[0341] Additionally, refer to Figure 31 If the value of `regular_merge_flag` is 0, the decoder can parse `ciip_flag` when both the conditions for using CIIP and the conditions for using the triangle merge mode are met. In this case, a value of `ciip_flag` of 1 indicates that CIIP can be used, and a value of 0 indicates that the triangle merge mode can be used. A value of 0 for `ciip_flag` may mean that CIIP is not used. On the other hand, if neither the conditions for using CIIP nor the conditions for using the triangle merge mode are met, the decoder may not parse `ciip_flag`.
[0342] When ciip_flag does not exist, its value can be inferred to be 1 if all of the following conditions are met. Conversely, if none of the following conditions are met, its value can be inferred to be 0.
[0343] Condition c-1) sps_ciip_enabled_flag==1
[0344] Condition c-2) general_merge_flag==1
[0345] Condition c-3) merge_subblock_flag == 0
[0346] Condition c-4)regular_merge_flag==0
[0347] Condition c-5)cbWidth<128
[0348] (Condition c-6)cbHeight<128
[0349] Condition c-7)cbWidth*cbHeight>=64
[0350] Condition c-8)cu_skip_flag == 0
[0351] `sps_ciip_enabled_flag` can be a syntax element signaled at a higher level to indicate whether CIIP is used. `sps_ciip_enabled_flag` can be signaled from the sequence parameter set. `general_merge_flag` can be a syntax element indicating whether a merge mode is used. `merge_subblock_flag` can be a syntax element indicating whether a subblock merge mode is used. In this case, the subblock merge mode can be an affine merge mode or a subblock-based temporal motion vector prediction (SbTMVP). `regular_merge_flag` can be a syntax element indicating whether an existing merge mode (e.g., regular merge mode) or MMVD is used.
[0352] Figure 32 This is a diagram illustrating the merged data syntax according to an embodiment of the present invention.
[0353] As mentioned above, the block size that can be used with CIIP is limited, and in this case, the block size can be a fixed value. Furthermore, when the maximum transform size is variable, there is a problem of misalignment between intra-prediction in CIIP mode and intra-prediction in intra-prediction mode.
[0354] In the following text, reference will be made to Figure 32 Describe the methods used to solve this problem. References Figure 32The block size that can be used in CIIP mode can be variable. Specifically, the block size that can be used in CIIP mode can be based on MaxTbSizeY. For example, CIIP mode can be used when the block width is less than or equal to MaxTbSizeY and the block height is less than or equal to MaxTbSizeY. Conversely, CIIP can be omitted when the block width is greater than MaxTbSizeY or the block height is greater than MaxTbSizeY. In this case, the block width and height can be the width and height of the coded block (unit). Because the block width and height are determined based on MaxTbSizeY, the intra-prediction size in CIIP mode is limited to less than or equal to MaxTbSizeY, and existing intra-prediction sizes are also limited to less than or equal to MaxTbSizeY. Therefore, a unified approach, hardware, and software can be used to perform intra-prediction. This has the effect of reducing the resources required for hardware and software. Even when MaxTbSizeY is 32, both intra-prediction in CIIP mode and regular intra-prediction can be performed on blocks of size 32×32 or smaller.
[0355] refer to Figure 32 The conditions for allowing the use of CIIP mode can include a value of 1 for `sps_ciip_enabled_flag`, `cbWidth` being less than or equal to `MaxTbSizeY`, `cbHeight` being less than or equal to `MaxTbSizeY`, `cbWidth*cbHeight` being 64 or greater, and a value of 0 for `cu_skip_flag`. Therefore, CIIP mode is not used when `sps_ciip_enabled_flag` is 0, `cbWidth` is greater than `MaxTbSizeY`, `cbHeight` is greater than `MaxTbSizeY`, `cbWidth*cbHeight` is less than 64, or `cu_skip_flag` is not 0. When the conditions for using CIIP mode are met, the decoder can parse `ciip_flag`. Additionally, when the conditions for using CIIP mode are met, the decoder can parse `regular_merge_flag`. Furthermore, the decoder can consider additional conditions when parsing `ciip_flag` or `regular_merge_flag`. Additionally, when CIIP mode can be omitted, the decoder may not need to parse ciip_flag.
[0356] Conditions that allow the use of the triangle (merge) mode may include a value of 1 for sps_triangle_enabled_flag, a value greater than 1 for MaxNumTriangleMergeCand, a slice_type of B, and a cbWidth*cbHeight of 64 or greater.
[0357] In this scenario, the decoder can parse `ciip_flag` when both the conditions for allowing CIIP mode and the conditions for allowing triangle (merge) mode are met. If neither condition for allowing CIIP mode nor the conditions for allowing triangle (merge) mode are met, the decoder may not parse `ciip_flag`. Furthermore, if `ciip_flag` does not exist and the conditions for allowing CIIP mode are not met, the value of `ciip_flag` can be inferred to be 0. Conversely, if `ciip_flag` does not exist, the conditions for allowing CIIP mode are met, `general_merge_flag` is 1, `merge_subblock_flag` is 0, and `regular_merge_flag` is 0, the value of `ciip_flag` can be inferred to be 1.
[0358] The decoder can parse the `regular_merge_flag` when either the conditions for allowing CIIP mode or the conditions for allowing triangle (merge) mode are met.
[0359] If neither the conditions for CIIP mode nor the conditions for allowing the use of triangle (merge) mode are met, the decoder may not parse regular_merge_flag. In this case, if regular_merge_flag does not exist, its value can be inferred as (general_merge_flag && !merge_subblock_flag).
[0360] refer to Figure 32 When cbWidth is less than or equal to MaxTbSizeY and cbHeight is less than or equal to MaxTbSizeY, the decoder can parse ciip_flag. Conversely, when cbWidth is greater than MaxTbSizeY or cbHeight is greater than MaxTbSizeY, the decoder may not parse ciip_flag.
[0361] When MaxTbSizeY is used, the value of ciip_flag can be inferred to be equal to 0.
[0362] When cbWidth is less than or equal to MaxTbSizeY and cbHeight is less than or equal to
[0363] When MaxTbSizeY is met, the decoder can parse regular_merge_flag. On the other hand, when i) cbWidth is greater than MaxTbSizeY or cbHeight is greater than MaxTbSizeY, and ii) the conditions for allowing the use of the triangle (merge) mode are not met, regular_merge_flag may not be parsed.
[0364] The decoder can parse ciip_flag when the conditions of Equation 5 below are met.
[0365] [Equation 5]
[0366] (sps_ciip_enabled_flag&&sps_triangle_enabled_flag&&
[0367] MaxNumTriangleMergeCand>1&&slice_type=B&&
[0368] cu_skip_flag[x0][y0]==0&&
[0369] (cbWidth*cbHeight)>=64&&cbWidth<=MaxTbSizeY&&cbHeight<=MaxTbSizeY)
[0370] In the absence of ciip_flag, the value of ciip_flag can be inferred to be equal to 1 if all of the following conditions are met. On the other hand, the value of ciip_flag can be inferred to be equal to 0 if any of the following conditions are not met.
[0371] Condition d-1) sps_ciip_enabled_flag==1
[0372] Condition d-2) general_merge_flag==1
[0373] Condition d-3)merge_subblock_flag == 0
[0374] Condition d-4)regular_merge_flag==0
[0375] Condition d-5)cbWidth<=MaxTbSizeY
[0376] Condition d-6)cbHeight<=MaxTbSizeY
[0377] Condition d-7)cbWidth*cbHeight>=64
[0378] Condition d-8)cu_skip_flag == 0
[0379] The decoder can parse regular_merge_flag when the conditions of Equation 6 below are met.
[0380] [Equation 6]
[0381] ((cbWidh*cbHeight)>=64&&((sps_clip_enabled_flag&&
[0382] cu_skip_flag[x0][y0]===0&&cbWidth<=MaxTbSizeY&&cbHeight<=MaxTbSizeY)||
[0383] (sps_triabnle_enabled_flag&&MaxNumTriangleMergeCand>1&&
[0384] slice_type==B)))
[0385] When `regular_merge_flag` does not exist, its value can be inferred as (`general_merge_flag && !merge_subblock_flag`). In other words, if `general_merge_flag` is 1 and `merge_subblock_flag` is 0, `regular_merge_flag` can be inferred to be 1. Conversely, if `general_merge_flag` is 0 or `merge_subblock_flag` is 1, `regular_merge_flag` can be inferred to be 0.
[0386] Figure 33 This is a diagram illustrating a method for executing CIIP mode according to an embodiment of the present invention.
[0387] Figure 33 Diagrams can be used to solve this problem. Figure 30 A diagram illustrating an embodiment of the problem described herein.
[0388] In embodiments of the present invention, intra-frame prediction in CIIP mode can be performed on a transform block (unit) basis. That is, TU partitioning can be performed when using CIIP mode. For example, when cbWidth is greater than MaxTbSizeY or cbHeight is greater than MaxTbSizeY, CIIP mode can be used, in which case TU partitioning is performed, and intra-frame prediction can be performed on blocks of (tbWidth x tbHeight) on a (MaxTbSizeY x MaxTbSizeY) basis. Therefore, intra-frame prediction in CIIP mode can be performed using the same resources as existing intra-frame prediction.
[0389] Figure 33 The diagram illustrates the prediction method when CIIP mode is used for a CU of size 64×64 with MaxTbSizeY of 32. Because CIIP mode is used, prediction signals can be generated based on prediction signals using inter-frame prediction and intra-frame prediction. In this case, inter-frame prediction can be performed on a 64×64 CU. Furthermore, since a 64×64 block exceeds MaxTbSizeY, TU partitioning can be performed before intra-frame prediction. The block after TU partitioning can be a block of size MaxTbSizeY × MaxTbSizeY. That is, after TU partitioning, multiple blocks of size 32×32 can be generated. Intra-frame prediction can then be performed separately on the blocks generated after TU partitioning.
[0390] The following section describes an intra-frame prediction method for the case of using CIIP mode and performing TU partitioning.
[0391] For intra-frame prediction after TU partitioning, reference samples located adjacent to each TU can be used after partitioning. Using reference samples from nearby blocks offers advantages in both prediction performance and coding efficiency. For example, the reference sample could be a reconstructed sample. However, in this case, to perform intra-frame prediction on a particular TU, it's necessary to wait for adjacent TUs (i.e., the TU corresponding to the reference sample's location) to be reconstructed, potentially causing unnecessary latency. Alternatively, the reference sample could be a prediction sample. In this case, prediction performance might be lower than using reconstructed samples, but prediction for the current TU can begin once prediction is complete, even if adjacent TUs are not fully reconstructed. Therefore, latency can be reduced compared to using reconstructed samples as reference samples. However, even when using prediction samples as reference samples, latency issues may still exist.
[0392] For intra-frame prediction after TU partitioning, reference samples from locations adjacent to CU before TU partitioning can be used. This solves the aforementioned latency problem. However, prediction performance may degrade due to the use of non-adjacent samples. Furthermore, since intra-frame prediction is performed on some partitions of the TU using reference samples not adjacent to the TU, reference samples from different locations than in conventional intra-frame prediction are required, along with a different process.
[0393] When performing TU partitioning in CIIP mode, predictions in CIIP mode can be performed even if the CU size is large. This is because intra-frame predictions included in CIIP mode can be performed on CUs with large sizes, and intra-frame predictions included in CIIP mode can also be performed on smaller blocks of partitions following the TU partition. Therefore, the conditions related to the block size that can be used in CIIP mode may not have an upper limit. That is, there may not be an upper limit. Figures 31 to 32 The condition described in [the document] is that cbWidth is less than 128 and cbHeight is less than 128. Therefore, when the condition that the block size (cbWidth * cbHeight) is 64 or greater is met, the decoder can resolve either ciip_flag or regular_merge_flag. In other words, even if cbWidth is greater than 64 and cbHeight is greater than 64 (see [the document]...), the decoder can resolve either ciip_flag or regular_merge_flag. Figure 31 ), cbWidth is greater than MaxTbSizeY, or cbHeight is greater than MaxTbSizeY (see Figure 32 The decoder can also parse ciip_flag or regular_merge_flag.
[0394] Specifically, ciip_flag can be parsed when the following conditions are met. On the other hand, ciip_flag can be left unparsed when none of the following conditions are met.
[0395] Condition e-1) sps_ciip_enabled_flag==1
[0396] Condition e-2) sps_triangle_enabled_flag==1
[0397] Condition e-3)MaxNumTriangleMergeCand>1
[0398] Condition e-4) slice_type==B
[0399] Condition e-5)cu_skip_flag == 0
[0400] Condition e-6)cbWidth*cbHeight>=64
[0401] In the absence of ciip_flag, ciip_flag is inferred to be equal to 1 if all of the following conditions are met, and ciip_flag is inferred to be equal to 0 if any of the following conditions are not met.
[0402] Condition f-1) sps_ciip_enabled_flag==1
[0403] Condition f-2) general_merge_flag==1
[0404] Condition f-3) merge_subblock_flag==0
[0405] Condition f-4)regular_merge_flag==0
[0406] Condition f-5)cbWidth*cbHeight>=64
[0407] Condition f-6)cu_skip_flag == 0
[0408] The `regular_merge_flag` can be parsed if all of the following conditions are met. Conversely, the `regular_merge_flag` may not be parsed if none of the following conditions are met.
[0409] Condition g-1)cbWidth*cbHeight>=64
[0410] Condition g-2)
[0411] (sps_ciip_enabled_flag&&cu_skip_flag==0)||(sps_triangle_enabled_flag&&MaxNumTriangleMergeCand>1&&slice_type==B)
[0412] When regular_merge_flag does not exist, its value can be inferred, which is consistent with the reference. Figure 31 and 32 Since the descriptions are identical, their descriptions will be omitted.
[0413] Figure 34 This is a diagram illustrating the chroma BDPCM syntax structure according to an embodiment of the present invention.
[0414] It can perform operations on chromaticity components.Figures 27 to 29 The BDPCM described herein. By performing BDPCM on the chroma components, the compression performance of specific video content can be improved. In this invention, applying BDPCM to chroma blocks will be described as chroma BDPCM.
[0415] To perform BDPCM on the chromaticity components, a separate signaling method is required, and a reference... Figure 34 To describe. Additionally. Figure 34 The publicly disclosed syntax can exist in the encoding unit syntax.
[0416] When the treeType value is SINGLE_TREE or DUAL_TREE_CHROMA, it can be parsed. Figure 34 The syntax elements disclosed in [the document]. Additionally, when ChromaArrayType is not 0, it can be parsed. Figure 34 The syntax elements disclosed in [the document] are as follows. When the value of separate_colour_plane_flag is 0, ChromaArrayType can be set to the value of chroma_format_idc. When the value of separate_colour_plane_flag is 1, ChromaArrayType can be set to 0. In this case, separate_colour_plane_flag can indicate whether the color components of the 4:4:4 chroma format are encoded separately. For example, when the value of separate_colour_plane_flag is 0, it can indicate that the color components are not encoded separately. When the value of separate_colour_plane_flag is 0 and it is not monochrome, ChromaArrayType can be non-zero.
[0417] When `pred_mode_plt_flag` is 1 and `treeType` is `DUAL_TREE_CHROMA`, `palette_coding()` can be executed. `pred_mode_plt_flag` is a syntax element indicating whether palette mode is used, and `palette_coding()` can be part of parsing syntax elements related to palette mode. When `pred_mode_plt_flag` is 0 or `treeType` is not `DUAL_TREE_CHROMA`, the decoder can execute... Figure 34 Step (1) and subsequent steps.
[0418] Additionally, based on cu_act_enabled_flag, it is possible to execute... Figure 34Step (1) and subsequent steps are described in the text. `cu_act_enabled_flag` can be a syntax element indicating whether adaptive color transformation is applied. When `cu_act_enabled_flag` is 1, the residual signal can be encoded in a different color space, and when `cu_act_enabled_flag` is 0, the residual signal can be encoded in the original color space. The original color space can be the YUV color space or the YCbCr color space. The other color space can be the YCgCo color space or the RGB color space. (See reference...) Figure 34 When the value of cu_act_enabled_flag is 0, execution is possible. Figure 34 Step (1) and subsequent steps. Also, when the value of cu_act_enabled_flag is 1, it can be omitted. Figure 34 Step (1) and subsequent steps. That is, when the value of cu_act_enabled_flag is 0, the decoder can parse the syntax elements related to chroma intra-frame prediction, and when the value of cu_act_enabled_flag is not 0, the decoder can not parse the syntax elements related to chroma intra-frame prediction.
[0419] refer to Figure 34 The decoder can parse intra_bdpcm_chroma_flag when the conditions of Equation 7 below are met.
[0420] [Equation 7]
[0421] (cbWidth<=MaxTsSize&&cbHeight<=MaxTsSize&&sps_bdpcm_chroma_enabled_flag)
[0422] `intra_bdpcm_chroma_flag` can be a syntax element indicating whether BDPCM is applied to the current chroma code block. For example, a value of 1 for `intra_bdpcm_chroma_flag` indicates that BDPCM is applied to the current chroma code block, and a value of 0 indicates that BDPCM is not applied to the current chroma code block. (See reference) Figure 34There may be conditions for resolving `intra_bdpcm_chroma_flag`. Alternatively, there may be conditions for using chroma BDPCM. The conditions for using chroma BDPCM can be the same as the conditions for resolving `intra_bdpcm_chroma_flag`. The conditions that allow resolving `intra_bdpcm_chroma_flag` are the same as those in Equation 7 above, which will be described in detail below.
[0423] Condition h-1)cbWidth<=MaxTsSize
[0424] Condition h-2)cbheight<=MaxTsSize
[0425] Condition h-3) sps_bdpcm_chroma_enable_flag==1
[0426] The decoder may not parse `intra_bdpcm_chroma_flag` if any of the above conditions are not met. When `intra_bdpcm_chroma_flag` does not exist, its value can be inferred to be 0. When the value of `intra_bdpcm_chroma_flag` is 1, the decoder can parse the syntax related to chroma BDPCM. In this case, the syntax related to chroma BDPCM may include `intra_bdpcm_chroma_dir_flag`. `intra_bdpcm_chroma_dir_flag` can be a flag indicating the prediction direction of chroma BDPCM. For example, `intra_bdpcm_chroma_dir_flag` can indicate whether the prediction direction of chroma BDPCM is horizontal or vertical. `sps_bdpcm_chroma_enable_flag` can be a syntax element sent at a higher level indicating whether chroma BDPCM is usable. For example, the `sps_bdpcm_chroma_enable_flag` can be sent using a signal at a level that includes the current coding unit (e.g., sequence, image, slice, etc.). When the value of `sps_bdpcm_chroma_enable_flag` is 1, chroma BDPCM can be used, and there can be additional syntax elements indicating whether chroma BDPCM is used. When the value of `sps_bdpcm_chroma_enable_flag` is 0, chroma BDPCM is not used.
[0427] refer to Figure 34When the decoder parses `intra_bdpcm_chroma_flag` and, as a result, `intra_bdpcm_chroma_flag` has a value of 0, the decoder does not parse syntax related to chroma intra-prediction. In this case, syntax related to chroma intra-prediction may include `cclm_mode_flag`, `cclm_mode_idx`, `intra_chroma_pred_mode`, etc. In this case, `cclm_mode_flag` can be a syntax element indicating whether the Cross-Component Linear Model (CCLM) is used as the chroma intra-prediction mode. `cclm_mode_flag` can be a syntax element indicating whether the chroma intra-prediction mode is one of `INTRA_LT_CCLM`, `INTRA_L_CCLM`, and `INTRA_T_CCLM`. `CcLMEnabled` can be a value indicating whether CCLM is usable. Alternatively, `CclmEnabled` can be a value indicating whether to parse a syntax element indicating whether CCLM is usable. CCLM is a chroma prediction method based on luminance samples. When using CCLM, `cclm_mode_idx` can be parsed. `cclm_mode_idx` can be a syntax element indicating which of several CCLM methods is used. When CCLM is not used, the decoder can resolve `intra_chroma_pred_mode`. In this case, `intra_chroma_pred_mode` can be a syntax element indicating the intra-prediction mode for chroma samples. Specifically, `intra_chroma_pred_mode` can be a syntax element indicating which of the following modes—planar mode (mode index 0), DC mode (mode index 1), vertical mode (mode index 50), horizontal mode (mode index 18), diagonal mode (mode index 66), and DM mode (the same mode as luma mode)—is used as the prediction mode for the chroma components.
[0428] refer to Figure 34 Even if the conditions for allowing the parsing of intra_bdpcm_chroma_flag are met or the conditions for using chroma BDPCM are met, when a syntax element that does not use chroma BDPCM is sent by signal, the problem of uncertain chroma prediction mode may occur because the syntax related to the above chroma intra-frame prediction is not parsed.
[0429] Figure 35 This is a diagram illustrating the chroma BDPCM syntax structure according to an embodiment of the present invention.
[0430] Figure 35 The embodiments are for solving the reference Figure 34 The embodiments of the problem described will be presented, and redundant content will be omitted.
[0431] refer to Figure 35 The decoder can parse `intra_bdpcm_chroma_flag` when the conditions for parsing `intra_bdpcm_chroma_flag` or the conditions for using chroma BDPCM are met. The conditions for parsing `intra_bdpcm_chroma_flag` are... Figure 35 (1) is the same as Equation 7 above. That is, chroma BDPCM can be used when all of the following conditions are met, and chroma BDPCM can be not used when any of the following conditions are not met. In addition, intra_bdpcm_chroma_flag may not exist when any of the following conditions are not met. In this case, when intra_bdpcm_chroma_flag does not exist, the value of intra_bdpcm_chroma_flag can be inferred to be equal to 0.
[0432] Condition i-1)cbWidth<=MaxTSSize
[0433] Condition i-2)cbHeight<=MaxTsSize
[0434] Condition i-3) sps_bdpcm_chroma_enable_flag==1
[0435] Figure 29 The BDPCM application method described in [the document] can also be applied to chroma BDPCM. That is, chroma BDPCM can be used even if the block size cbWidth or cbHeight is greater than MaxTsSize, when the TU or TB size of the partition is less than or equal to MaxTsSize. Therefore, chroma BDPCM can be used when all of the following conditions are met, and chroma BDPCM can be omitted when none of the following conditions are met.
[0436] Condition j-1)
[0437] (cbWidth<=MaxTsSize&&cbHeight<=MaxTsSize)||(MaxTbSizeY==MaxTsSize)
[0438] Condition j-2) sps_bdpcm_chroma_enable_flag==1
[0439] When the value of `intra_bdpcm_chroma_flag` is 1, the decoder can parse syntax related to chroma BDPCM. In this case, the syntax related to chroma BDPCM may include `intra_bdpcm_chroma_dir_flag`, `sps_bdpcm_chroma_enable_flag`, etc. `intra_bdpcm_chroma_dir_flag`, `sps_bdpcm_chroma_enable_flag`, etc., can be compared with the reference... Figure 34 The syntax for the description is the same.
[0440] refer to Figure 35 When the decoder parses intra_bdpcm_chroma_flag and uses it as a result, if the value of intra_bdpcm_chroma_flag is 0 (when chroma BDPCM is not used), the decoder can execute... Figure 35 Step (2) and subsequent steps. That is, the decoder can parse the syntax related to other chroma intra-prediction. The syntax related to other chroma intra-prediction can be compared with the reference. Figure 34 The syntax described is the same. Syntax related to other chroma intra-prediction functions may include cclm_mode_flag, cclm_mode_idx, intra_chroma_pred_mode, etc.
[0441] The decoder can parse cclm_mode_flag when the value of intra_bdpcm_chroma_flag is 0 and the value of CclmEnabled is 1. Additionally, the decoder can parse cclm_mode_idx when the value of cclm_mode_flag is 1, and can parse intra_chroma_pred_mode when the value of cclm_mode_flag is not 1.
[0442] refer to Figure 35Even when the conditions for parsing `intra_bdpcm_chroma_flag` are met, and a syntax element indicating that chroma BDPCM is not used is sent as a result of parsing `intra_bdpcm_chroma_flag`, the decoder can determine the chroma prediction mode because syntax elements used to determine other chroma prediction modes can be parsed. Specifically, if the block width and height are less than or equal to the maximum transform skip size and a syntax element indicating whether chroma BDPCM can be used is sent at a higher level, the decoder can parse the syntax element indicating whether chroma BDPCM can be used. In this case, when the parsing result indicates that chroma BDPCM is not used, the decoder can determine the chroma intra prediction mode based on syntax elements indicating which chroma intra prediction mode will be used (e.g., `cclm_mode_flag`, `intra_chroma_pred_mode`).
[0443] Figure 36 This is a diagram illustrating a higher-level syntax related to BDPCM according to an embodiment of the present invention.
[0444] As described above, higher-level syntax can exist, including syntax elements indicating whether BDPCM is usable. For example, syntax elements indicating whether BDPCM is usable, signaled at a higher level, can include sps_bdpcm_enabled_flag, sps_bdpcm_chroma_enabled_flag, etc., as mentioned above. Specifically, sps_bdpcm_enabled_flag is a syntax element signaled at a higher level indicating whether BDPCM for the luminance component is usable. Additionally, sps_bdpcm_chroma_enabled_flag is a syntax element signaled at a higher level indicating whether BDPCM for the chroma component is usable. In this case, the higher level at which syntax elements indicating whether BDPCM is usable can be signaled can include sequence level, sequence parameter set level (SPS level), slice level, picture level, and picture parameter set level (PPS level). When a syntax element indicating whether BDPCM is usable, signaled at a higher level, indicates that BDPCM is usable (e.g., when its value is 1), additional syntax elements regarding whether BDPCM is used can exist. For example, syntax elements regarding whether BDPCM is used can exist at both the block level and the coding unit (block) level. However, when a syntax element indicating whether BDPCM is usable, signaled at a higher level, indicates that BDPCM is not usable (e.g., when its value is 0), additional syntax elements regarding whether BDPCM is used may not exist. Furthermore, when a syntax element indicating whether BDPCM is usable, signaled at a higher level, does not exist, the value of that syntax element can be inferred to be equal to 0. Additionally, syntax elements related to whether `sps_bdpcm_enabled_flag` is used can be referenced. Figure 28 and 29 The description refers to `intra_bdpcm_flag`. `intra_bdpcm_flag` can be the same as `intra_bdpcm_luma_flag`. Additionally, syntax elements related to whether or not `sps_bdpcm_chroma_enabled_flag` are available in the reference section. Figure 34 and 35 The description is for intra_bdpcm_chroma_flag.
[0445] A higher-level indicator signal can be used to skip over whether a syntax element is usable. (See reference.) Figure 36At a higher level, the syntax element indicating whether transform skipping is usable can be `sps_transform_skip_enalbed_flag`. In this case, the higher level for signaling the syntax element indicating whether transform skipping is usable can include sequence level, sequence parameter set level (SPS level), slice level, picture level, and picture parameter set level (PPS level). When the syntax element indicating whether transform skipping is usable at a higher level indicates that transform skipping is usable (e.g., when its value is 1), additional syntax elements regarding whether transform skipping is used can exist. For example, syntax elements regarding whether transform skipping is used can exist at the block level and the coding unit (block) level. Conversely, when the syntax element indicating whether transform skipping is usable at a higher level indicates that transform skipping is not usable (e.g., when its value is 0), no additional syntax elements regarding whether transform skipping is used. When the syntax element indicating whether transform skipping is usable at a higher level does not exist, the value of the syntax element signaled at the higher level can be inferred to be equal to 0. Syntax elements related to whether or not to use the `sps_transform_skip_enabled_flag` can be found in the reference. Figure 27 and 28 The described transform_skip_flag.
[0446] As mentioned above, in order to use BDPCM, it should be possible to skip transformations. (See reference) Figure 36 When a syntax element indicating whether a transform skip is usable is signaled at a higher level, the decoder can parse the syntax element indicating whether BDPCM is usable. Conversely, when a syntax element indicating whether a transform skip is unusable is signaled at a higher level, the decoder may not parse the syntax element indicating whether BDPCM is usable. In this case, the syntax element indicating whether BDPCM is usable can include information about the luminance component and information about the chrominance component.
[0447] When the value of `sps_bdpcm_enabled_flag` is 1, the decoder can parse `sps_bdpcm_chroma_enabled_flag`. Conversely, when the value of `sps_bdpcm_enabled_flag` is 0, the decoder can choose not to parse it. In other words, the BDPCM for the chroma component can only be used if the BDPCM for the luma component is available.
[0448] refer to Figure 36 When the decoder parses `sps_transform_skip_enabled_flag` and uses it as a result, if the value of `sps_transform_skip_enabled_flag` is 1, the decoder can parse `sps_bdpcm_enabled_flag`. If the value of `sps_bdpcm_enabled_flag` is 1, the decoder can parse `sps_bdpcm_chroma_enabled_flag`.
[0449] Figure 37 This is a diagram illustrating the syntax elements of BDPCM transmitted by signals at a higher level according to an embodiment of the present invention.
[0450] exist Figure 37 In the embodiments, the references to Figure 36 The content is redundant.
[0451] refer to Figure 37 The usability of chroma BDPCM can be determined based on the chroma format. Information related to the color format may include `chroma_format_idc`, `separate_colour_plance_flag`, chroma format, `SubWidthC`, `SubHeightC`, etc. Specifically, chroma BDPCM can be used when the color format is 4:4:4. Conversely, chroma BDPCM may not be usable when the color format is not 4:4:4. Therefore, whether chroma BDPCM can be used can be based on `chroma_format_idc`. For example, when the value of `chroma_format_idc` is 3, the decoder can parse the syntax element `sps_bdpcm_chroma_enabled_flag` indicating whether chroma BDPCM is usable; otherwise, the decoder may not parse the syntax element `sps_bdpcm_chroma_enabled_flag` indicating whether chroma BDPCM is usable. In this invention, the color format can be used with the same meaning as the chroma format.
[0452] Figure 38 This is a diagram illustrating the syntax related to chroma BDPCM according to an embodiment of the present invention.
[0453] Syntax elements indicating whether a BDPCM for the luma component and a BDPCM for the chroma component can be used, both transmitted at a higher level, may not exist independently. That is, a single syntax element can indicate whether a BDPCM for both the luma and chroma components can be used. For example, when a syntax element indicating whether a BDPCM for the luma component can be used indicates that a BDPCM is usable, either the one for the luma component or the one for the chroma component can be used. Alternatively, in this case, additional syntax elements indicating whether a BDPCM is used can exist. For example, these additional syntax elements could be the `intra_bdpcm_flag` and `intra_bdpcm_chroma_flag` mentioned above. Conversely, when a syntax element indicating whether a BDPCM for the luma component can be used indicates that a BDPCM is not usable, the BDPCM may not be used for either the luma or chroma components.
[0454] The syntax elements indicating whether to use a BDPCM for the luma component and whether to use a BDPCM for the chroma component can be parsed based on the syntax elements sent at a higher level, which indicate whether the same BDPCM can be used for both the luma and chroma components. For example, when the syntax element indicating whether the same BDPCM can be used is `sps_intra_bdpcm_flag`, as shown in the reference... Figures 28 to 29 The decoder can parse syntax elements indicating whether BDPCM for the luma component is used. Specifically, when the value of `sps_bdpcm_enabled_flag` is 1, the decoder can parse `intra_bdpcm_flag` (in which case additional conditions can be considered), and when the value of `sps_bdpcm_enabled_flag` is not 1, the decoder can choose not to parse `intra_bdpcm_flag`.
[0455] refer to Figure 38 The value of `sps_bdpcm_enabled_flag` can be used to indicate whether to use the syntax elements of BDPCM for the chroma components. For example, when the value of `sps_bdpcm_enabled_flag` is 1, the decoder can parse `intra_bdpcm_chroma_flag` (in which case additional conditions can be considered), and when the value of `sps_bdpcm_enabled_flag` is not 1, the decoder can choose not to parse `intra_bdpcm_chroma_flag`.
[0456] For the luma and chroma components, the syntax elements indicating whether to use BDPCM at the coding unit (block) level can be omitted. That is, when the syntax element indicating whether to use BDPCM specifies its use, BDPCM can be used for both the luma and chroma components. Conversely, when the syntax element indicating whether to use BDPCM specifies its non-use, BDPCM may not be used for either the luma or chroma components.
[0457] Figure 39 This is a diagram illustrating the syntax related to intra-frame prediction according to an embodiment of the present invention.
[0458] refer to Figure 39 It can send syntax elements related to multiple intra-frame predictions in a preset order using signals.
[0459] Syntax elements can be grouped and sent as signals for each color component. For example, syntax elements related to the luma component can be grouped and sent as signals. Syntax elements related to the chroma component can also be grouped and sent as signals. Furthermore, groups of syntax elements related to the luma component and groups of syntax elements related to the chroma component can be sent as signals sequentially. Conversely, groups of syntax elements related to the chroma component and groups of syntax elements related to the luma component can be sent as signals sequentially.
[0460] refer to Figure 39 This will describe the structure of sending the syntax element groups related to the luminance component and the syntax element groups related to the chrominance component in sequence using signals.
[0461] The syntax element group related to the luminance component can include the following syntax elements.
[0462] i) Syntax elements related to BDPCM used for the luminance component:
[0463] Syntax elements can include syntax elements indicating whether BDPCM can be used and syntax elements indicating the prediction direction. For example, Figure 39 The intra_bdpcm_luma_flag and intra_bdpcm_luma_dir_flag disclosed in the document can correspond to syntax elements.
[0464] ii) Matrix-based intra-frame prediction (MIP) related syntax elements:
[0465] Syntax elements refer to syntax elements that indicate whether MIP prediction is used, syntax elements that indicate whether the MIP input vector is transposed, and syntax elements that indicate the MIP mode. For example, Figure 39The intra_mip_flag, intra_mip_transposed, and intra_mip_mode disclosed in the document can correspond to syntax elements.
[0466] iii) Syntax elements indicating reference lines used for prediction:
[0467] Syntax elements can be syntax elements that indicate the position of reference lines used for intra-frame prediction. For example, Figure 39 The intra_luma_ref_idx disclosed in the document can correspond to a syntax element.
[0468] iv) Intra-Frame Sub-Partition (ISP) related syntax elements:
[0469] Syntax elements can include syntax elements indicating whether to use ISP and syntax elements indicating the direction of ISP partitioning. For example, Figure 39 The intra_subpartitions_mode_flag and intra_subpartitions_split_flag disclosed in the document can correspond to syntax elements. In this case, ISP is either an intra-subpartition mode or a method of prediction by partitioning a block into subpartitions.
[0470] v) Syntax elements indicating intra-prediction mode:
[0471] Syntax elements can include those indicating whether to use the MPM list (candidate list), those indicating whether to use a flat pattern, those indicating an MPM index, and those indicating a pattern not included in the MPM list. For example, Figure 39 The intra_luma_mpm_flag, intra_luma_not_planar_flag, intra_luma_mpm_idx, and intra_luma_mpm_remainder disclosed in the document can correspond to syntax elements.
[0472] The syntax element group related to chroma components can include the following syntax elements.
[0473] i) Chroma component syntax elements related to BDPCM:
[0474] The chromaticity component syntax elements can include syntax elements indicating whether BDPCM is used and syntax elements indicating the prediction direction. For example, Figure 39 The intra_bdpcm_chroma_flag and intra_bdpcm_chroma_dir_flag disclosed in the document can correspond to chroma component syntax elements.
[0475] ii) Syntax elements related to Cross-Component Linear Model (CCLM):
[0476] CCLM-related syntax elements can include syntax elements indicating whether CCLM is used and syntax elements indicating the CCLM pattern used for prediction. For example, Figure 39 The cclm_mode_flag and cclm_mode_idx disclosed herein can correspond to CCLM-related syntax elements. CCLM can be a method for prediction based on another color component value. Specifically, CCLM can be a method for predicting the chromaticity component based on the luminance component value.
[0477] iii) Syntax elements related to chroma intra-frame prediction mode:
[0478] Syntax elements can be syntax elements used to determine the intra-prediction mode index of the chroma components. For example, Figure 39 The intra_chroma_pred_mode disclosed in the document can correspond to a syntax element.
[0479] refer to Figure 39 When the treeType is SINGLE_TREE or DUAL_TREE_LUMA, the decoder can parse syntax elements related to the luma component, and when the treeType is SINGLE_TREE or DUAL_TREE_CHROMA, the decoder can parse syntax elements related to the chroma component. Therefore, the decoder can continuously parse groups of syntax elements related to the luma component, and it can also continuously parse groups of syntax elements related to the chroma component.
[0480] refer to Figure 39 The syntax elements intra_bdpcm_chroma_flag and intra_bdpcm_chroma_dir_flag related to the chroma component can be located after the syntax element group related to the luma component. Therefore, when the tree type is SINGLE_TREE, intra_bdpcm_chroma_flag and intra_bdpcm_chroma_dir_flag located after the syntax element group related to the luma component can be resolved.
[0481] By resolving in this way, locating (parsing) syntax elements related to the luminance component and syntax elements related to the chrominance component sequentially, there are advantages in memory management. In other words, it can be efficient in terms of implementation.
[0482] Figure 40 This is a diagram illustrating the syntax related to intra-frame prediction according to an embodiment of the present invention.
[0483] refer to Figure 39 It can parse syntax element groups related to the luminance component and syntax element groups related to the chrominance component out of order. For example, it can reverse and parse the reference. Figure 39 The description specifies the order of some syntax elements in the groups of syntax elements related to the luma component and the groups of syntax elements related to the chroma component. Specifically, syntax elements related to BDPCM for the luma component and syntax elements related to BDPCM for the chroma component can be parsed consecutively. Syntax element groups that exclude BDPCM-related elements from the syntax elements included in the luma component-related syntax element group and syntax element groups that exclude BDPCM-related elements from the syntax elements included in the chroma component-related syntax element group can also be parsed consecutively. Additionally, BDPCM-related syntax elements can be parsed before other syntax element groups.
[0484] An example of parsing the order of syntax elements is shown below.
[0485] First, syntax elements related to the BDPCM used for the luma component can be parsed, and syntax elements related to the BDPCM used for the chroma component can also be parsed. Then, among the syntax elements included in the group of syntax elements related to the luma component, a group of syntax elements excluding those related to the BDPCM used for the luma component can be parsed. Then, among the syntax elements included in the group of syntax elements related to the chroma component, a group of syntax elements excluding those related to the BDPCM used for the chroma component can also be parsed.
[0486] refer to Figure 40 , intra_bdpcm_luma_flag, intra_bdpcm_luma_dir_flag, intra_bdpcm_chroma_flag and intra_bdpcm_chroma_dir_flag can be located at intra_mip_flag, intra_mip_transposed, intra_mip_mode, intra_luma_ref_idx, intra_subpartiti before ons_mode_flag, intra_subpartitions_split_flag, intra_luma_mpm_flag, intra_luma_not_planar_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, cclm_mode_flag, cclm_mode_idx and intra_chroma_pred_mode.
[0487] Alternatively, intra_bdpcm_luma_flag, intra_bdpcm_luma_dir_flag, intra_bdpcm_chroma_flag, and intra_bdpcm_chroma_dir_flag may precede at least one of intra_mip_flag, intra_mip_transposed, intra_mip_mode, intra_luma_ref_idx, intra_subpartitions_mode_flag, intra_subpartitions_split_flag, intra_luma_mpm_flag, intra_luma_not_planar_flag, intra_luma_mpm_idx, intra_luma_mpm_remainder, cclm_mode_flag, cclm_mode_idx, and intra_chroma_pred_mode.
[0488] refer to Figure 40 Syntax elements related to BDPCM for the chroma component can be placed before those related to the luma component. By positioning them in this way, there is a reduction in the actions / operations used to determine whether to parse syntax elements when chroma BDPCM is used extensively.
[0489] Luminance BDPCM and chromaticity BDPCM can share the same prediction direction. Therefore, as... Figure 40 As shown, `intra_bdpcm_luma_dir_flag` and `intra_bdpcm_chroma_dir_flag` do not exist independently, and can both indicate the prediction direction of luma BDPCM and chroma BDPCM through a single syntax element. In this case, whether luma BDPCM is used and whether chroma BDPCM is used may not be independent. For example, luma BDPCM may be needed to use chroma BDPCM. Therefore, the syntax element indicating whether luma BDPCM is used and the syntax element indicating whether chroma BDPCM is used can be related. The syntax element indicating whether chroma BDPCM is used can only be parsed if the syntax element indicating whether luma BDPCM is used indicates that luma BDPCM is used.
[0490] Figure 41 This is a diagram illustrating the sequence parameter set syntax according to an embodiment of the present invention.
[0491] Sequence Parameter Set (SPS) syntax can be applied to encoded video sequences (CVS). Whether SPS syntax is applied to CVS can be determined by the syntax element values of the Picture Parameter Set (PPS), which are referenced by the syntax elements of the slice header.
[0492] Figure 41 The syntax elements disclosed herein may be some of the SPS syntax elements. (See reference) Figure 41 The SPS syntax can include syntax elements such as `pic_width_max_in_luma_samples`, `pic_height_max_in_luma_samples`, `sps_log2_ctu_size_minus5`, `subpics_present_flag`, `sps_num_subpics_minus1`, `subpic_ctu_top_left_x`, `subpic_ctu_top_left_y`, `subpic_width_minus1`, `subpic_height_minus1`, `subpic_treated_as_pic_flag`, `loop_filter_across_subpic_enabled_flag`, `sps_subpic_id_present_flag`, `sps_subpic_id_signalling_present_flag`, `sps_subpic_id_len_minus1`, and `sps_subpic_id`. Additionally, some of the syntax elements included in the SPS syntax can exist in as many numbers as the number of subpicks. For example, the syntax elements subpic_ctu_top_left_x, subpic_ctu_top_left_y, subpic_width_minus1, subpic_height_minus1, subpic_treated_as_pic_flag, loop_filter_across_subpic_enabled_flag, and sps_subpic_id can exist as many as the number of subpicks. Specifically, these syntax elements can be represented as syntaxElement[I], where i can be 0 to (the number of subpicks - 1). In this case, syntaxElements can be as many syntax elements as there are subpicks.
[0493] A sub-image can be a unit lower than an image or frame. For example, an image or frame may include one or more sub-images. In this case, a sub-image can be a rectangular region. Specifically, a sub-image can refer to a rectangular region composed of one or more slices in an image. Furthermore, sub-images can be decoded independently. Therefore, even if the decoder only receives information about a specific sub-image in the image, the decoder can decode and reconstruct the received sub-image. Additionally, multiple sub-images in an image do not need to overlap.
[0494] refer to Figure 41 When the value of `subpics_present_flag` is 1, the decoder can parse `sps_num_subpics_minus1`, `subpic_ctu_top_left_x`, `subpic_ctu_top_left_y`, `subpic_width_minus1`, `subpic_height_minus1`, `subpic_treated_as_pic_flag`, and `loop_filter_across_subpic_enabled_flag`. Furthermore, `subpic_ctu_top_left_x`, `subpic_ctu_top_left_y`, `subpic_height_minus1`, `subpic_treated_as_pic_flag`, and `loop_filter_across_subpic_enabled_flag` can each be parsed as many times as `(sps_num_subpics_minus1+1)`.
[0495] Reference Figure 42 The syntax elements and sub-picture signaling methods of the above SPS syntax are described.
[0496] Figure 42 This is a diagram illustrating the syntax elements related to sub-images according to an embodiment of the present invention.
[0497] Figure 42 The `pic_width_max_in_luma_samples` element disclosed in the documentation is a syntax element indicating the maximum width of the image, and in this case, the maximum width of the image can be expressed in units of luminance samples. The value of `pic_width_max_in_luma_samples` can be non-zero and can be an integer multiple of a preset value. The preset value can be the larger of 8 and the minimum coding block size. In this case, the minimum coding block size can be determined based on luminance samples and can be described as `MinCbSizeY`.
[0498] Additionally, `pic_height_max_in_luma_samples` is a syntax element indicating the maximum height of the image, and in this case, the maximum height can be represented in units of luminance samples. Furthermore, the value of `pic_height_max_in_luma_samples` does not have to be 0 and can be an integer multiple of a preset value. The preset value can be the larger of 8 and the minimum coding block size. In this case, the minimum coding block size can be determined based on luminance samples and can be described as `MinCbSizeY`.
[0499] When an image uses sub-images, `pic_width_max_in_luma_samples` and `pic_height_max_in_luma_samples` refer to the image's width and height, respectively. In this case, the maximum width and maximum height of the image can be the same as the image's width and height, respectively.
[0500] `sps_log_2_ctu_size_minus5` can be a syntax element indicating the size of the coding tree block (CTB) of a coding tree unit (CTU). Specifically, `sps_log_2_ctu_size_minus5` can indicate the size of the luma coding tree block. Additionally, `sps_log_2_ctu_size_minus5` can have a value obtained by taking log2 of the CTB size in units of luma samples and then subtracting a preset value. If the size of the luma coding tree block is `CtbSizeY`, then `CtbSizeY` can be (1 << (sps_log2_ctu_size_minus5 + preset value)). The preset value can be 5. That is, when `sps_log2_ctu_size_minus5` is 0, 1, and 2, the `CtbSizeY` values can be 32, 64, and 128, respectively. Furthermore, the value of `sps_log2_ctu_size_minus5` can be 2 or less.
[0501] `subpics_present_flag` is a syntax element indicating the presence of subpicks. It indicates whether subpicks can be used or whether the number of subpicks can be signaled as a value greater than 1. In this case, subpicks can include `sps_num_subpics_minus1`, `subpic_ctu_top_left_x`, `subpic_ctu_top_left_y`, `subpic_width_minus1`, `subpic_height_minus1`, `subpic_treated_as_pic_flag`, `loop_filter_across_subpic_enabled_flag`, etc.
[0502] `sps_num_subpics_minus1` can be a syntax element indicating the number of subpicks. For example, the value obtained by adding a preset value to the value of `sps_num_subpics_minus1` can be the number of subpicks. In this case, the preset value can be 1. In this case, the number of subpicks can be signaled as a value of 1 or greater. Alternatively, the preset value can be 2. In this case, the number of subpicks can be signaled as a value of 2 or greater. The value of `sps_num_subpics_minus1` can be 0 or greater and 254 or less, and can be represented in 8 bits. Also, when the value of `sps_num_subpics_minus1` does not exist, its value can be inferred to be equal to 0.
[0503] Figure 42 The disclosed subpic_ctu_top_left_x, subpic_ctu_top_left_y, subpic_width_minus1, and subpic_height_minus1 are syntax elements that indicate the position and size of each subpic. subpic_ctu_top_left_x[i], subpic_ctu_top_left_y[i], subpic_width_minus1[i], and subpic_height_minus1[i] indicate the value corresponding to the i-th subpic. In this case, the value of i can be equal to or greater than 0 and less than or equal to the value of sps_num_subpics_minus1.
[0504] `subpic_ctu_top_left_x` indicates the x-coordinate (horizontal position) of the top-left position of a subpicture. Specifically, `subpic_ctu_top_left_x` can indicate the x-coordinate of the top-left CTU of the subpicture. In this case, the coordinate can be represented in CTU or CTB units. For example, the coordinate can be represented in CtbSizeY units. Additionally, `subpic_ctu_top_left_x` can be signaled as an unsigned integer value, and in this case, the number of bits can be Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY)). Also, when the value of `subpic_ctu_top_left_x` does not exist, its value can be inferred to be 0.
[0505] `subpic_ctu_top_left_y` indicates the y-coordinate (vertical position) of the top-left position of a subpicture. Specifically, `subpic_ctu_top_left_y` can indicate the y-coordinate of the top-left CTU of the subpicture. In this case, the coordinate can be represented in CTU or CTB units. For example, the coordinate can be represented in CtbSizeY units. Alternatively, `subpic_ctu_top_left_y` can be signaled as an unsigned integer value. In this case, the number of bits can be Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)). Also, when the value of `subpic_ctu_top_left_y` does not exist, its value can be inferred to be 0.
[0506] `subpic_width_minus1` can indicate the width of a subpicture. For example, the value obtained by adding a preset value to the value of `subpic_width_minus1` can be the width of the subpicture. In this case, the preset value can be 1. Additionally, the width of the subpicture can be represented in CTU or CTB units. For example, the width of the subpicture can be represented in CtbSizeY units. Furthermore, `subpic_width_minus1` can be signaled as an unsigned integer value, and in this case, the number of bits can be `Ceil(Log2(pic_width_max_in_luma_samples / CtbSizeY))`. Also, when the value of `subpic_width_minus1` does not exist, its value can be inferred to be equal to `(Ceil(pic_width_max_in_luma_samples / CtbSizeY)-1)`.
[0507] `subpic_height_minus1` indicates the height of a subpicture. For example, the value obtained by adding a preset value to the value of `subpic_height_minus1` can be the height of the subpicture. In this case, the preset value can be 1. Additionally, the height of the subpicture can be expressed in CTU or CTB units. For example, the height of the subpicture can be expressed in CtbSizeY units. Furthermore, `subpic_height_minus1` can be signaled as an unsigned integer value, and in this case, the number of bits can be Ceil(Log2(pic_height_max_in_luma_samples / CtbSizeY)). Also, when the value of `subpic_height_minus1` does not exist, its value can be inferred to be equal to (Ceil(pic_height_max_in_luma_samples / CtbSizeY)-1).
[0508] The Ceil(x) value described in this invention can be the smallest integer greater than or equal to x. The same operation as Ceil(Log2(x)) is performed in the above bit count calculation, which is the operation used to calculate the number of bits required to represent the value of x in binary format, where x is an integer greater than or equal to 0.
[0509] Furthermore, once the position and size of a sub-image are determined, the boundary positions of the sub-image can be calculated. This allows us to determine the locations within the sub-image.
[0510] Figure 43This is a diagram illustrating operators according to an embodiment of the present invention.
[0511] refer to Figure 43 Multiple division operations can be defined.
[0512] The ' / ' operator can represent integer division. That is, the result of ' / ' can be an integer. Specifically, to make the result of the ' / ' operation an integer, it can be truncated towards zero. For example, 7 / 4 and (-7) / (-4) have a value of 1. Furthermore, (-7) / 4 and 7 / (-4) have a value of -1. Integer division can be used in division about... Figure 42 In the operations or operations that infer the number of bits described in the text.
[0513] The division operator “÷” does not perform truncation or rounding. Therefore, the result of the “÷” operation can be or may not be an integer. For example, the values of 7÷4 and (-7)÷(-4) can be greater than 1. Furthermore, the values of (-7)÷4 and 7÷(-4) can be less than -1.
[0514] Figure 43 of The "-" operation is the same as the "÷" operation mentioned above.
[0515] Figure 44 These are pictures and sub-pictures illustrating embodiments of the present invention.
[0516] As mentioned above, an image can be divided into multiple sub-images. In this case, it can be done by... Figure 42 The publicly available signaling determines how to divide or configure sub-images. However, according to the reference... Figures 42 to 43 The described embodiments may contain ranges that cannot be represented by signaling. For example, when the width or height of the image is not divisible by CtbSizeY, there may be sub-image structures that cannot be represented. Specifically, when the width of the image is not divisible by CtbSizeY, the value of pic_width_max_in_luma_samples / CtbSizeY is truncated towards 0, thus potentially not representing the position of the rightmost CTB of the image. Similarly, when the height of the image is not divisible by CtbSizeY, the value of pic_height_max_in_luma_samples / CtbSizeY is truncated towards 0, thus potentially not representing the position of the bottommost CTB of the image. Therefore, when the width of the image is not divisible by CtbSizeY and the width of the sub-image is the same as the width of the image, the aforementioned signaling may not represent the width of the sub-image. Furthermore, when the width of the image is not divisible by CtbSizeY and the horizontal position of the sub-image is the rightmost CTB, the aforementioned signaling may not represent the top-left x-coordinate of the sub-image.
[0517] Specifically, refer to Figure 44 The image width can be 1032 luminance samples. Additionally, CtbSizeY can be 128 luminance samples. (For example, in...) Figure 44 In subpic 0, the width of the subpic can be the same as the width of the image. In this case, since the width of subpic 0 is 9 in CtbSizeY, the value of subpic_width_minus1 should represent 8. However, since pic_width_max_in_luma_samples / CtbSizeY(1032 / 128) is 8, ceil(log2(pic_width_max_in_luma_samples / CtbSizeY)) is 3. Therefore, Subpic_width_minus1 has the following problem: only values from 0 to 7 can be represented using 3 bits, while 8 may not be able to be represented. Additionally, there may be cases where the value of subpic_width_minus1 does not exist and is therefore inferred. In this case, since the image consists of a subpic, the width of the subpic should be inferred to be equal to the width of the image, but as mentioned above, when the width of the image is not divisible by CtbSizeY, the width of the subpic can be disregarded.
[0518] refer to Figure 44 For example, in subpicture 2, the top-left x-coordinate of the subpicture could be the rightmost CTB. In this case, the top-left x-coordinate of the subpicture should represent the 9th value in CtbSizeY units, and a value of 8 when the coordinate starts with 0. However, since pic_width_max_in_luma_samples / CtbSizeY is 8, ceil(log2(pic_width_max_in_luma_samples / CtbSizeY)) is 3. Therefore, subpic_ctu_top_left_x has the problem that only values from 0 to 7 can be represented using 3 bits, and 8 may not be able to be represented.
[0519] Figure 45 This is a diagram illustrating the syntax elements related to sub-images according to an embodiment of the present invention.
[0520] For reference Figure 44As mentioned above, when `pic_width_max_in_luma_samples` is not divisible by `CtbSizeY`, `pic_width_max_in_luma_samples / CtbSizeY` is truncated (rounded down), resulting in a reduced range of expression. Similarly, when `pic_height_max_in_luma_samples` is not divisible by `CtbSizeY`, `pic_height_max_in_luma_samples / CtbSizeY` is truncated (rounded down), and the range of expression is reduced. Therefore, this invention proposes a calculation method that does not involve truncation.
[0521] For example, since the top-left x and y coordinates of the sub-image, the width of the sub-image, and the height of the sub-image are expressed in units of CtbSizeY, they should be divisible by CtbSizeY. In this case, you can use... Figure 43 The "÷" operation described in the text.
[0522] `subpic_ctu_top_left_x` can represent the x-coordinate (horizontal position) of the top-left position of a subpicture. Specifically, `subpic_ctu_top_left_x` can represent the x-coordinate of the top-left CTU of the subpicture. In this case, the coordinate can be represented in CTU or CTB units. For example, it can be represented in CtbSizeY units. `subpic_ctu_top_left_x` can be signaled as an unsigned integer value, and the number of bits for `subpic_ctu_top_left_x` can be Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)). Also, when the value of `subpic_ctu_top_left_x` does not exist, its value can be inferred to be 0. In this case, the number of bits for `subpic_ctu_top_left_x` can be determined as Ceil(Log2(Ceil(pic_width_max_in_luma_samples÷CtbSizeY))).
[0523] `subpic_ctu_top_left_y` can represent the y-coordinate (vertical position) of the top-left position of a subpicture. Specifically, `subpic_ctu_top_left_y` can represent the y-coordinate of the top-left CTU of the subpicture. In this case, the coordinate can be represented in CTU or CTB units. For example, it can be represented in CtbSizeY units. `subpic_ctu_top_left_y` can be signaled as an unsigned integer value, and the number of bits for `subpic_ctu_top_left_y` can be Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)). Also, when the value of `subpic_ctu_top_left_y` does not exist, its value can be inferred to be 0. In this case, the number of bits for `subpic_ctu_top_left_y` can be determined as Ceil(Log2(Ceil(pic_height_max_in_luma_samples÷CtbSizeY))).
[0524] `subpic_width_minus1` can indicate the width of a subpicture. For example, the width of a subpicture can be indicated by adding a preset value to the value of `subpic_width_minus1`. In this case, the preset value can be 1. Additionally, the width of the subpicture can be represented in CTU or CTB units. For example, the width of the subpicture can be represented in CtbSizeY units. `subpic_width_minus1` can be signaled as an unsigned integer value, and the number of bits in `subpic_width_minus1` can be Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)). Also, when the value of `subpic_width_minus1` does not exist, its value can be inferred to be equal to (Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1). In this case, the number of bits in subpic_width_minus1 can be determined as Ceil(Log2(Ceil(pic_width_max_in_luma_samples÷CtbSizeY))).
[0525] `subpic_height_minus1` indicates the height of a subpicture. For example, the height of a subpicture can be represented by adding a preset value to the value of `subpic_height_minus1`. In this case, the preset value can be 1. Alternatively, the height of a subpicture can be represented in CTU or CTB units. For example, the height of a subpicture can be represented in CtbSizeY units. `subpic_height_minus1` can be signaled as an unsigned integer value, and the number of bits in `subpic_height_minus1` can be `Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY))`. Also, when the value of `subpic_height_minus1` does not exist, its value can be inferred to be equal to `(Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1)`. In this case, the number of bits in subpic_height_minus1 can be determined as Ceil(Log2(Ceil(pic_height_max_in_luma_samples÷CtbSizeY))).
[0526] By using Figure 45 The syntax elements disclosed in the code are used to expand the range of expressions for the position and size of sub-images. As mentioned above, when a syntax element is not present and its value is inferred, the number of bits can be calculated as Ceil(Log2(Ceil(x))). This may be advantageous in implementing devices (decoders / encoders) that can only receive y as an integer in Log2(y). The values of (pic_width_max_in_luma_samples÷CtbSizeY) and (pic_height_max_in_luma_samples÷CtbSizeY) above represent how many CtbSizeY units make up the width and height of the image, respectively. When using Figure 45 When the syntax element shown is used, it can represent Figure 44 The width of sub-image 0 and the x-coordinate of sub-image 2 are shown.
[0527] Figure 46 This is a flowchart illustrating a method for partition transformation blocks according to an embodiment of the present invention.
[0528] The following description is based on a reference. Figures 1 to 45 The described embodiments are methods and apparatus for partitioning transform blocks.
[0529] A video signal decoding device may include a processor that performs a method for partitioning transform blocks. First, the video signal decoding device may receive a bitstream including information about the transform block (e.g., syntax elements). The processor may determine a result value indicating the partitioning direction of the current transform block (TB) based on preset conditions (S4601). The processor may partition the current transform block into multiple transform blocks based on the result value (S4602). The processor may use the multiple transform blocks to decode the video signal (S4603).
[0530] Additionally, the video signal encoding device may include a processor that determines a result value indicating the partitioning direction of the current transform block (TB) based on preset conditions, divides the current transform block into multiple transform blocks based on the result value, and generates a bit stream including information about the multiple transform blocks.
[0531] In this case, the preset conditions may include conditions related to the color components of the current transform block.
[0532] In this case, the aforementioned preset conditions additionally include conditions related to the comparison between the width of the current transform block and the maximum transform block width, and the maximum transform block width can be determined based on the chroma format associated with the current transform block, the color components of the current transform block, and the maximum transform size.
[0533] In this case, the preset conditions also include conditions related to the result of comparing the first width value and the first height value, where the first width value is obtained by multiplying the width of the current transform block by a first value, and the first height value is obtained by multiplying the height of the current transform block by a second value. The first value and the second value are values related to the width and height of the current transform block, respectively. If the color component of the current transform block is luminance, the first value and the second value are set to 1, and if the color component of the current transform block is chrominance, the first value and the second value are determined based on the chrominance format associated with the current transform block.
[0534] When the width of the current transform block is greater than the maximum transform block width and the first width value is greater than the first height value, the resulting value can be determined to be 1, which is a value indicating that the partitioning direction is vertical. In this case, the width of each of the multiple transform blocks can be a value obtained by dividing the width of the transform block by 2, and the height of each of the multiple transform blocks can be the same as the height of the transform block.
[0535] When the width of the current transform block is less than or equal to the maximum transform block width, or when the first width value is less than or equal to the first height value, the resulting value can be determined as 0, which indicates that the partitioning direction is horizontal. In this case, the width of each of the multiple transform blocks can be the same as the width of the transform block, and the height of each of the multiple transform blocks can be the value obtained by dividing the height of the transform block by 2.
[0536] In this case, the preset conditions can be the conditions described by equations 2 to 4 above.
[0537] The maximum transform size can be determined based on the size of the code tree block (CTB) with the luminance component included in the code tree unit (CTU) associated with the current transform block. When the size of the code tree block is 32, the maximum transform size can be 32.
[0538] If the color component of the current transform block is chroma, the processor can parse a syntax element indicating whether the prediction method of the coded block associated with the current transform block is based on Block-Based Incremental Pulse Code Modulation (BDPCM). Subsequently, as a result of the parsing, if the prediction method of the coded block is not BDPCM, the processor can also parse syntax elements related to the prediction method of the coded block and determine the prediction method of the coded block based on the parsing result. In this case, the syntax element related to the prediction method of the coded block is a syntax element indicating at least one of Cross-Component Linear Model (CCLM), Planar Mode, DC Mode, Vertical Mode, Horizontal Mode, Diagonal Mode, and DM Mode.
[0539] The above embodiments of the present invention can be implemented by various means. For example, the embodiments of the present invention can be implemented by hardware, firmware, software, or a combination thereof.
[0540] In the case of hardware implementation, the method according to the embodiments of the present invention can be implemented by one or more of the following: application-specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field-programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, etc.
[0541] When implemented via firmware or software, the method according to embodiments of the invention can be implemented in the form of modules, processes, or functions that perform the above-described functions or operations. Software code can be stored in memory and driven by a processor. The memory can be located inside or outside the processor and can exchange data with the processor in various known ways.
[0542] Some embodiments may also be implemented in the form of a recording medium including computer-executable instructions, such as a program module executed by a computer. A computer-readable medium can be any available medium accessible to a computer and can include all volatile, non-volatile, removable, and non-removable media. Additionally, a computer-readable medium can include computer storage media and communication media. Computer storage media includes all volatile, non-volatile, removable, and non-removable media implemented in any method or technique for storing information such as computer-readable instructions, data structures, program modules, or other data. Typically, communication media includes computer-readable instructions, other data modulated data signals (such as data structures or program modules), or other transmission mechanisms, and includes any information transmission medium.
[0543] The above description of the present invention is for illustrative purposes only, and it will be understood that those skilled in the art to which this invention pertains can make changes to the invention without altering its technical concept or essential characteristics, and that the invention can be readily modified in other specific forms. Therefore, the above embodiments are illustrative and not limiting in any way. For example, each component described as a single entity can be distributed and implemented, and similarly, components described as distributed can also be implemented in an associated manner.
[0544] The scope of this invention is defined by the appended claims rather than the foregoing detailed description, and all changes or modifications derived from the meaning and scope of the appended claims and their equivalents shall be construed as being included within the scope of this invention.
Claims
1. A video signal decoding device, comprising a processor, in, The processor is configured to: Determine the resulting value indicating the partitioning direction of the current transform block (TB), wherein the partitioning direction is either vertical or horizontal. Based on the result value, the current transform block is partitioned into multiple transform blocks. Decode the current transform block based on the plurality of transform blocks. Wherein, when the width of the current transform block is less than or equal to the maximum transform block width, or when the first value is less than or equal to the second value, the resulting value is a value indicating that the partitioning direction is the horizontal direction. Wherein, when the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the resulting value is a value indicating that the partitioning direction is the vertical direction. in i) When the color component of the current transform block is the luminance component: The first value is obtained by multiplying the width of the current transform block by 1, and The second value is obtained by multiplying the height of the current transform block by 1. ii) When the color component of the current transform block is a chromaticity component: The first value is obtained by multiplying the width of the current transform block by SubWidthC, and The second value is obtained by multiplying the height of the current transform block by SubHeightC. Wherein, each of SubWidthC and SubHeightC is a variable determined based on the chroma format associated with the current transform block. Wherein, when the color component of the current transform block is the luminance component, the maximum transform block width is the maximum transform size, and Wherein, when the color component of the current transform block is the chroma component, the maximum transform block width is the maximum transform size divided by the SubWidthC.
2. The video signal decoding device according to claim 1, in, When the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:0, SubWidthC is 2 and SubHeightC is 2. Wherein, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:2, SubWidthC is 2 and SubHeightC is 1, and Specifically, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:4:4, SubWidthC is 1 and SubHeightC is 1.
3. The video signal decoding device according to claim 1, in, When the width of the current transform block is less than or equal to the width of the maximum transform block, or when the first value is less than or equal to the second value, the result value is determined to be 0.
4. The video signal decoding device according to claim 1, in, When the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the result value is determined to be 1.
5. The video signal decoding device according to claim 1, in, When the partitioning direction is the vertical direction: The width of each of the plurality of transform blocks is the width of the current transform block divided by 2, and The height of each of the plurality of transform blocks is the same as the height of the current transform block.
6. The video signal decoding device according to claim 1, in, When the partitioning direction is the horizontal direction: The width of each of the plurality of transform blocks is the same as the width of the current transform block, and The height of each of the plurality of transform blocks is the height of the current transform block divided by 2.
7. A video signal encoding device, comprising a processor, in, The processor is configured to: Obtain the bitstream to be decoded by the decoder using the decoding method. The decoding method includes: Determine the resulting value indicating the partitioning direction of the current transform block (TB), wherein the partitioning direction is either vertical or horizontal. Based on the result value, the current transform block is partitioned into multiple transform blocks, and Decode the current transform block based on the plurality of transform blocks. Wherein, when the width of the current transform block is less than or equal to the maximum transform block width, or when the first value is less than or equal to the second value, the resulting value is a value indicating that the partitioning direction is the horizontal direction. Wherein, when the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the resulting value is a value indicating that the partitioning direction is the vertical direction. in i) When the color component of the current transform block is the luminance component: The first value is obtained by multiplying the width of the current transform block by 1, and The second value is obtained by multiplying the height of the current transform block by 1. ii) When the color component of the current transform block is a chromaticity component: The first value is obtained by multiplying the width of the current transform block by SubWidthC, and The second value is obtained by multiplying the height of the current transform block by SubHeightC. Wherein, each of SubWidthC and SubHeightC is a variable determined based on the chroma format associated with the current transform block. Wherein, when the color component of the current transform block is the luminance component, the maximum transform block width is the maximum transform size, and Wherein, when the color component of the current transform block is the chroma component, the maximum transform block width is the maximum transform size divided by the SubWidthC.
8. The video signal encoding device according to claim 7, in, When the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:0, SubWidthC is 2 and SubHeightC is 2. Wherein, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:2, SubWidthC is 2 and SubHeightC is 1, and Specifically, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:4:4, SubWidthC is 1 and SubHeightC is 1.
9. The video signal encoding device according to claim 7, in, When the width of the current transform block is less than or equal to the width of the maximum transform block, or when the first value is less than or equal to the second value, the result value is determined to be 0.
10. The video signal encoding device according to claim 7, in, When the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the result value is determined to be 1.
11. The video signal encoding device according to claim 7, in, When the partitioning direction is the vertical direction: The width of each of the plurality of transform blocks is the width of the current transform block divided by 2, and The height of each of the plurality of transform blocks is the same as the height of the current transform block.
12. The video signal encoding device according to claim 7, in, When the partitioning direction is the horizontal direction: The width of each of the plurality of transform blocks is the same as the width of the current transform block, and The height of each of the plurality of transform blocks is the height of the current transform block divided by 2.
13. A method for obtaining a bitstream, the method comprising: Determine the resulting value indicating the partitioning direction of the current transform block (TB), wherein the partitioning direction is either vertical or horizontal; Based on the result value, the current transform block is partitioned into multiple transform blocks; and Obtain the bitstream including information about the plurality of transform blocks. Wherein, when the width of the current transform block is less than or equal to the maximum transform block width, or when the first value is less than or equal to the second value, the resulting value is a value indicating that the partitioning direction is the horizontal direction. Wherein, when the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the resulting value is a value indicating that the partitioning direction is the vertical direction. in i) When the color component of the current transform block is the luminance component: The first value is obtained by multiplying the width of the current transform block by 1, and The second value is obtained by multiplying the height of the current transform block by 1. ii) When the color component of the current transform block is a chromaticity component: The first value is obtained by multiplying the width of the current transform block by SubWidthC, and The second value is obtained by multiplying the height of the current transform block by SubHeightC. Wherein, each of SubWidthC and SubHeightC is a variable determined based on the chroma format associated with the current transform block. Wherein, when the color component of the current transform block is the luminance component, the maximum transform block width is the maximum transform size, and Wherein, when the color component of the current transform block is the chroma component, the maximum transform block width is the maximum transform size divided by the SubWidthC.
14. The method according to claim 13, in, When the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:0, SubWidthC is 2 and SubHeightC is 2. Wherein, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:2:2, SubWidthC is 2 and SubHeightC is 1, and Specifically, when the color component of the current transform block is a chroma component and the chroma format associated with the current transform block is 4:4:4, SubWidthC is 1 and SubHeightC is 1.
15. The method according to claim 13, in, When the width of the current transform block is less than or equal to the width of the maximum transform block, or when the first value is less than or equal to the second value, the result value is determined to be 0.
16. The method according to claim 13, in, When the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the result value is determined to be 1.
17. The method according to claim 13, in, When the partitioning direction is the vertical direction: The width of each of the plurality of transform blocks is the width of the current transform block divided by 2, and The height of each of the plurality of transform blocks is the same as the height of the current transform block.
18. The method according to claim 13, in, When the partitioning direction is the horizontal direction: The width of each of the plurality of transform blocks is the same as the width of the current transform block, and The height of each of the plurality of transform blocks is the height of the current transform block divided by 2.
19. A method for processing a video signal, the method comprising: Determine the resulting value indicating the partitioning direction of the current transform block (TB), wherein the partitioning direction is either vertical or horizontal; Based on the result value, the current transform block is partitioned into multiple transform blocks; and Decode the current transform block based on the plurality of transform blocks. Wherein, when the width of the current transform block is less than or equal to the maximum transform block width, or when the first value is less than or equal to the second value, the resulting value is a value indicating that the partitioning direction is the horizontal direction. Wherein, when the width of the current transform block is greater than the width of the maximum transform block and the first value is greater than the second value, the resulting value is a value indicating that the partitioning direction is the vertical direction. in i) When the color component of the current transform block is the luminance component: The first value is obtained by multiplying the width of the current transform block by 1, and The second value is obtained by multiplying the height of the current transform block by 1. ii) When the color component of the current transform block is a chromaticity component: The first value is obtained by multiplying the width of the current transform block by SubWidthC, and The second value is obtained by multiplying the height of the current transform block by SubHeightC. Wherein, each of SubWidthC and SubHeightC is a variable determined based on the chroma format associated with the current transform block. Wherein, when the color component of the current transform block is the luminance component, the maximum transform block width is the maximum transform size, and Wherein, when the color component of the current transform block is the chroma component, the maximum transform block width is the maximum transform size divided by the SubWidthC.