Mutual action of conversion size with code tool
By adapting transform coding tools and enabling larger coding units to use advanced transforms, the limitations of reduced maximum transform size in VVC are overcome, enhancing coding efficiency and flexibility in video encoding and decoding.
Patent Information
- Application Number
- JP2025046721
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-20
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2040-09-08
AI Technical Summary
The interaction between the maximum transform size and other transform coding tools in video encoding and decoding, particularly in the Versatile Video Coding (VVC) standard, is not adequately addressed, leading to limitations in coding efficiency and flexibility when the maximum transform size is reduced from 64 to 32.
Adapting the zero-setting process, multiple transform selection (MTS) size, chroma transform size, transform skip size, and block-based delta pulse code modulation (BDCM) to accommodate the reduced maximum transform size of 32, while enabling matrix-based intra prediction (MIP) and low-frequency non-separable transform (LFNST) for coding units up to 64x64, regardless of the maximum transform size.
Enhances coding efficiency by allowing larger coding units to utilize advanced transforms, improving performance by up to 0.6% for large dimensions and maintaining coding gains without additional tools, and providing flexibility to the encoder.
Smart Images

Figure 2025108440000001_ABST
Abstract
Description
Technical Field
[0001] At least one of the embodiments generally relates to a method or apparatus for encoding or decoding video, or compressing or decompressing video.
Background Art
[0002] To achieve high compression efficiency, video coding schemes typically employ prediction including motion vector prediction, and transforms that exploit the spatial and temporal redundancy of video content. Generally, intra prediction or inter prediction is used to exploit the correlation within or between frames, and then the difference between the original image and the predicted image, often called the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy coding, quantization, transformation, and prediction.
[0003] In the development of the Versatile Video Coding (VVC) standard, the maximum transform size is variable between 32 and 64. The maximum transform size interacts with other transform coding tools.
Summary of the Invention
[0004] At least one of the embodiments generally relates to a method or apparatus for encoding (encoding) or decoding (decoding) video, and more specifically, to a method or apparatus for the interaction between the maximum transform size and transform coding tools in a video encoder or video decoder.
[0005] According to a first aspect, a method is provided. The method includes steps for enabling a coding tool based on a maximum transform size, performing at least a portion of a discrete trigonometric transform on a subset of samples including a block, and encoding the block using the enabled coding tool.
[0006] According to a second aspect, a method is provided. The method includes steps for enabling an encoding tool based on a maximum transformation size, performing at least a portion of an inverse discrete trigonometric transform on a subset of samples including blocks, and decoding the blocks using the enabled encoding tool.
[0007] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to encode blocks of a video or decode a bitstream by performing any of the methods described above.
[0008] According to another general aspect of at least one embodiment, a device is provided that includes an apparatus according to any of the decoding embodiments and at least one of: (i) an antenna configured to receive a signal, the signal including video blocks; (ii) a band limiter configured to limit a received signal to a frequency band including video blocks; or (iii) a display configured to display an output representing video blocks.
[0009] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that includes data content generated according to any of the encoding embodiments or variations described.
[0010] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated according to any of the encoding embodiments or variations described.
[0011] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the encoding embodiments or variations described.
[0012] According to another general aspect of at least one embodiment, there is provided a computer program product including instructions that, when executed by a computer, cause the computer to perform any of the described decoding embodiments or variations.
[0013] The above and other aspects, features, and advantages of the general aspects will become apparent by reading the following detailed description of the exemplary embodiments with reference to the accompanying drawings.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0015] At least one of the present embodiments generally relates to a method or apparatus for video encoding or decoding, and more specifically, to a method or apparatus for the interaction between the maximum transform size and transform coding tools in a video encoder or video decoder.
[0016] The general aspects described herein are in the field of video compression. It is the interaction between the maximum transform size and other transform coding tools where the maximum transform size varies from 32 to 64 in the recent adoption of VVC. The value is calculated as follows.
[0017]
Number
[0018] The maximum transform size interacts with the following tools. 1 - Zero - setting process: First, the VVC performs zero - setting to reduce the complexity of large transform sizes. For the 2D DCT2 transform, only the top - left 32×32 coefficients are retained while the rest are set to zero. That is, for 64×64, 64×32, and 32×32, DCT2 calculates the first 32 coefficients in both the horizontal and vertical directions, and non - DCT2 transforms (DST7 and DCT8) perform zero - setting to retain the top - left 16×16 coefficients. The adoption of JVET - 00545 does not resolve how zero - setting is performed when the maximum transform size is 32.
[0019] The capture of the draft text is shown below, and the zero - setting is shaded.
[0020]
Table 1
[0021] This means that when tu_mts_idx is greater than zero, i.e., when the MTS transform is used (DST7, DST7), the zero - setting width and height are set to 16, while when tu_mts_idx is zero (DCT2), the zero - setting is set to 32. 2 - MTS size: MTS or multiple transform selections are transform tools adopted in the VVC where the selection between DST7 and DCT8 pairs is permitted as other transform pairs that enhance the DCT2 transform pair. MTS is executed for block sizes from 4×4 to 32×32. That is, half - size DCT2. The adoption of JVET - 00545 does not resolve how the MTS size is considered when the maximum transform size is 32.
[0022] The capture of the draft text is shown below, and the mts size is masked.
[0023]
Table 2-1
[0024]
Table 2-2
[0025] This indicates that the MTS is signaled when both the width and height are less than 32, regardless of MaxTbSizeY. 3 - Chroma conversion size: In VVC, the chroma size is half of the luma size. That is, the chroma conversion block is allowed to be 2×2 to 32×32, while the luma size is 4×4 to 64×64. With the adoption of JVET - 00545, how the chroma size is fixed when the maximum conversion size is 32 is not resolved.
[0026] In the VVC spec, the chroma size is calculated as follows.
[0027]
Table 3
[0028] Under common testing condition (CTC), a 4:2:0 chroma format is used and the maximum transform size (MaxTbSizeY) is 64. Therefore, the maximum transform size is 32 for the chroma in CTC. However, when MaxTbSizeY is 32, the maximum chroma size is 16 according to the current SPEC. 4 - Transform skip size: In VVC, transform skip is performed for block sizes having the same range as DCT2. In other words, transform skip is performed for block sizes from 4×4 to 64×64. With the adoption of JVET - 00545, how the transform skip size is fixed when the maximum transform size is 32 is not resolved.
[0029] In the VVC specification, the maximum transform skip size is defined as follows. log2_transform_skip_max_size_minus2 specifies the maximum block size used for transform skip and shall be in the range of 0 to 3.
[0030] If not present, the value of log2_transform_skip_max_size_minus2 is assumed to be equal to 0.
[0031] The variable MaxTsSize is set equal to 1<<(log2_transform_skip_max_size_minus2 + 2).
[0032] That is, the maximum MaxTsSize can take values from 4 to 32 regardless of the MaxTbSizeY value. 5 - BDCM size: BD - PCM is block - based delta pulse code modulation. It is a coding tool for the transform skip residual. It is currently permitted under the same size conditions as transform skip. That is, the maximum MaxTsSize. The following text shows the conditions (masking) for coding the BDPCM flag.
[0033]
Table 4
[0034] 6-MIP In VVC Draft 6, MIP (Matrix-based Intra Prediction) is an intra prediction mode in which the prediction signal is generated by multiplying several trained prediction matrices with a certain shift to the reference samples. The mode is signaled when the CU size is less than or equal to the maximum allowable transform size dimension. This limitation was necessary to limit memory requirements and coding complexity. This is because MIP is a matrix-based method and the prediction matrix is larger for larger blocks.
[0035] First, the maximum transform size (MaxTbSizeY) is always kept as 64 in VTM5.0. However, in VTM6.0, this value can be either 64 or 32. Samples of the draft text of VTM6.0 are provided below (the shaded part indicates the MIP part).
[0036]
Table 5-1
[0037]
Table 5-2
[0038] The MaxTbSizeY value was fixed at 64. However, with the adoption of JVET-00545, MaxTbSizeY can be either 64 or 32.
[0039] Intuitively, when MaxTbSizeY is 32, MIP is signaled up to a CU size of 32×32. This prevents larger MIPs from using CUs, thus limiting the coding efficiency. The presently described embodiments propose to enable MIPs for CUs up to 64×64, regardless of the maximum transform size. This is done by enabling TU tiling when the CU size is larger than MaxTbSizeY.
[0040] First, the maximum transform size is always maintained as 64 in VTM 5.0. However, in the recent adoption of JVET-00545, the maximum transform size (MaxTbSizeY) can be either 64 or 32, controlled by the SPS flag (sps_sbt_max_size_64_flag). When this occurs, it is necessary to adapt the zero-setting process, MTS size, chroma transform size, transform skip size, and BDCM size to this change.
[0041] General embodiments propose to adapt the signaling of the following tools according to the maximum transform size: zero-setting process, MTS size, chroma transform size, transform skip size, and BDCM. The affected codec modules are the intra-coding designs (160) and 260 of FIGS. 1 and 2.
[0042] Embodiment 1: Zero-Setting Process In this embodiment, the zero-setting process depends on the maximum transform size. In this way, when the maximum size is 32 instead of 64, the zero-setting size is reduced by half. This is shown (in italics) in the following text
[0043]
Table 6
[0044] This can also be done independently for DCT2 transform and other MTS transforms (DST7 and DCT8). That is, if you want to do it only for DCT2:
[0045]
Table 7
[0046] Otherwise, for DST7 / DCT8 only
[0047]
Table 8
[0048] Embodiment 2: MTS Size MTS signaling is permitted up to a size of 32×32. This is independent of MaxTbSizeY, regardless of whether it is 64 or 32. To connect with MaxTbSizeY, the signaling of MTS can be resized up to MaxTbSizeY / 2×MaxTbSizeY / 2. This is indicated in italics in the following spec:
[0049]
Table 9-1
[0050]
Table 9-2
[0051] Note that this directly affects the subblock transform (SBT) tool. SBT is a transform unit splitting tool for inter-blocks that semantically selects transforms from DCT2, DST7, and DCT8. According to this specification, the variable implicitMtsEnabled is derived as follows. - If -sps_mts_enabled_flag is equal to 1 and one of the following conditions is true, implicitMtsEnabled is set to 1: - IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT - cu_sbt_flag is equal to 1 and Max(nTbW, nTbH) is 32 or less - sps_explicit_mts_intra_enabled_flag is equal to 0, CuPredMode[0][xTbY][yTbY] is equal to MODE_INTRA, lfnst_idx[x0][y0] is equal to 0, and intra_mip_flag[x0][y0] is equal to 0 - Otherwise, implicitMtsEnabled is equal to 0.
[0052] The variable trTypeHor that specifies the horizontal transform kernel and the variable trTypeVer that specifies the vertical transform kernel are derived as follows. - If cIdx is greater than 0, trTypeHor and trTypeVer are equal to 0. - Otherwise, if ImplicitMts enabled is equal to 1, the following applies. - If IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT or sps_explicit_mts_intra_enabled_flag is equal to 0 and CuPredMode[0][xTbY][yTbY] is equal to MODE_INTRA, trTypeHor and trTypeVer are derived as follows. trTypeHor=(nTbW>=4&&nTbW<=16)?1:0(8-975) trTypeVer=(nTbH>=4&&nTbH<=16)?1:0(8-976) - Otherwise (when cu_sbt_flag is equal to 1), trTypeHor and trTypeVer are specified in Tables 8 to 15 according to cu_sbt_horizontal_flag and cu_sbt_pos_flag. - Otherwise, trTypeHor and trTypeVer are specified in Table 8-14 according to tu_mts_idx[xTbY][yTbY].
[0053]
Table 10
[0054] Conversion type 2 means DCT8, and 1 means DST7.
[0055] That is, when MTS is limited to MaxTbSizeY / 2 and MaxTbSizeY is 32, DST7 and DCT8 of size 32×32 are not supported, so the above table cannot be used. Instead, DCT2 needs to be used. The corresponding specification changes are as follows.
[0056] The variable implicitMtsEnabled is derived as follows. - When sps_mts_enabled_flag is equal to 1 and one of the following conditions is true, implicitMtsEnabled is set equal to 1: - IntraSubPartitionsSplitType is not equal to ISP_NO_SPLIT - cu_sbt_flag is equal to 1 and Max(nTbW, nTbH) is less than or equal to MaxTbSizeY / 2 - sps_explicit_mts_intra_enabled_flag is equal to 0, CuPredMode[0][xTbY][yTbY] is equal to MODE_INTRA, lfnst_idx[x0][y0] is equal to 0, and intra_mip_flag[x0][y0] is equal to 0
[0057] Embodiment 3: Chroma Conversion Size According to the VVC specification, the maximum chroma conversion width and height can be half of the maximum luminance 1. However, since the maximum luminance conversion size can be 32, the maximum size of chroma can be 16. This is a small number and is considered not useful in practice. Therefore, in this embodiment, the minimum of the chroma size is fixed at 32. The specification can be changed as follows (italic). maxTbWidth=(cIdx==0)?MaxTbSizeY:max(MaxTbSizeY / SubWidthC,32)(8-41) maxTbHeight=(cIdx==0)?MaxTbSizeY:max(MaxTbSizeY / SubHeightC,32)(8-42)
[0058] Embodiment 4: Transform Skip Size The transform skip flag can be signaled up to a size of 32×32. This is independent of the maximum transform block size, regardless of whether it is 64 or 32. To connect with the maximum transform size, the text is modified as follows. log2_transform_skip_max_size_minus2 specifies the maximum block size used for transformation and shall be in the range of 0 to MaxTbLog2SizeY - 3.
[0059] MaxTbLog2SizeY is calculated as follows (according to the VVC specification). MaxTbLog2SizeY=sps_max_luma_transform_size_64_flag?6:5(7-28)
[0060] Embodiment 5: BDCM BDPCM uses the same signaling conditions as transform skip. Therefore, the above Embodiment 4 is also applicable to BDPCM.
[0061] Embodiment 6: MIP The described general aspect proposes to enable MIP for CUs of up to 64×64 regardless of the maximum conversion size. This is done by enabling TU tiling when the CU size is larger than MaxTbSizeY. This is to improve the coding efficiency when MaxTbSizeY is 32.
[0062] The basic idea of the present invention is to permit MIP for CUs of a maximum size of 64x64 regardless of MaxTbSizeY. This is to improve the coding performance when MaxTbSizeY is set to 32 by enabling MIP when the CU size is larger than 32×32. Experimentally, it has been shown that MIP functions better for arrays with large dimensions. The following results are generated by obtaining the VTM software as an anchor, and the tests are on VTM without MIP.
[0063] [Table 11]
[0064] Obviously, MIP gives a 0.6% gain for class A1 (large dimensions) and a 0.3% gain for class C (small dimensions). Therefore, when MaxTbSizeY is 32, enabling MIP for 64×64 CUs provides a coding gain for arrays with large dimensions. Furthermore, since the same structure of MIP is maintained, no additional tools are required.
[0065] Compared with the current design, even when the conversion size is smaller, more flexibility for the encoder is enabled by implementing MIP for CUs of up to 64×64.
[0066] This can be achieved by TU tiling. That is, the CU is distributed into multiple TUs and MIP is executed independently. There are two methods for doing the following. 1 - TU tiling, then MIP prediction TU Tiling after 2-MIP Prediction
[0067] That is, considering that the CUs of size 64×64, 32×64 or 64x32 and MaxTbSizeY are 32, the first option is to divide the CU into 32×32 TUs, perform MIP in 32×32 blocks, generate a prediction signal, and encode the residuals. In that case, the prediction signal can be generated using reference samples from the reconstructed 32×32 blocks. The second option is to perform MIP in a large CU (64×64, 32×64 or 64×32), then divide the CU into TUs of size 32×32, and encode the residuals. The second option is more consistent with the current design of VTM. This is because, for conventional intra prediction (angular, DC or planar), the prediction signal is generated with the same size as the TU, and as a result, the reference samples can use the reconstructed blocks to improve the prediction of adjacent blocks.
[0068] The corresponding specifications are shown in italics below.
[0069] [Table 12-1]
[0070] [Table 12-2]
[0071] The specification text already supports TU tiling when the transform unit is larger than the maximum transform size, which is shown in the following text (shaded).
[0072] [Table 13]
[0073] Embodiment 7: LFNST Index In the specification text, LFNST (Low Frequency Non-Separable Transform) is permitted up to MaxTbSizeY. The original motivation was to avoid the latency issue when decoding a large CU of size 128×128 where the LFNST index is decoded after decoding the TU residual. Therefore, it was decided to make LFNST possible up to the maximum transform size, which was initially 64. With the adoption of JVET-00545, the latency issue is not significant when the CU size is 64×64 and MaxTbSizeY is 32. Therefore, when MaxTbSIzeY is set to 32, it is possible to enable the LFNST index in this case to improve the coding gain.
[0074] The corresponding specification changes are indicated in italics.
[0075] [Table 14]
[0076] Furthermore, LFNST is only permitted when the primary transform is DCT2 (condition: tu_mts_idx [x0][y0]==0), so it is necessary to check this condition for multiple TUs’. The changes are as follows.
[0077] [Table 15]
[0078] [Table 16]
[0079] That is, define a variable MTS_notDCT2 to check whether any of the TUs is not using DCT2. If so, LFNST is not permitted.
[0080] An embodiment of method 400 under the general aspects described herein is shown in FIG. 4. The method starts at start block 401, and control proceeds to block 410 to enable coding tools based on the maximum transform size. Control proceeds from block 410 to block 420 and performs at least a portion of a discrete trigonometric transform on a subset of samples including the blocks. Control proceeds from block 420 to block 430 and encodes the blocks using the enabled coding tools. Transformed coefficients of the blocks are determined using the transformed subset of samples.
[0081] An embodiment of method 500 under the general aspects described herein is shown in FIG. 5. The method starts at start block 501, and control proceeds to block 510 to enable coding tools based on the maximum transform size. Control proceeds from block 510 to block 520 and performs at least a portion of an inverse discrete trigonometric transform on a subset of samples including the blocks. Control proceeds from block 520 to block 530 and decodes the blocks using the enabled coding tools.
[0082] FIG. 6 shows an embodiment of an apparatus 600 for compressing, encoding, or decoding video using various coding tools depending on the maximum transform size. The apparatus includes a processor 1410 and can be interconnected to a memory 1420 through at least one port. Both the processor 1410 and the memory 1420 can also have one or more additional interconnections to external connections.
[0083] Furthermore, processor 610 is configured to insert or receive information in the bitstream and compress, encode, or decode using any of the described aspects.
[0084] This document describes various aspects, including tools, functions, embodiments, models, methods, etc. Many of these aspects are specifically described and may sometimes be read as limiting the present invention, at least to show individual features. However, this is for the purpose of clarifying the description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined or interchanged to provide further aspects. Further, these aspects can be combined with or interchanged with aspects described in previous applications.
[0085] The aspects described and contemplated in this document can be implemented in many different forms. The following FIGS. 1, 2, and 3 provide some embodiments, but other embodiments are also contemplated, and the descriptions of FIGS. 1, 2, and 3 do not limit the scope of the implementation forms. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the methods, apparatuses, and methods described, and / or a computer-readable storage medium storing a bitstream generated according to any of the methods described.
[0086] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstructed" is used on the encoder side and the term "decoded" is used on the decoder side.
[0087] This specification describes various methods, each of which includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the method to operate correctly, the order and / or use of specific steps and / or actions can be changed or combined.
[0088] Using the various methods and other aspects described in this document, modules of the video encoder 100 and video decoder 200 as shown in FIGS. 1 and 2, such as, for example, the intra prediction module, the entropy encoding module, and / or the decoding module (160, 360, 145, 330), can be modified. Further, aspects of the present disclosure are not limited to VVC or HEVC, and can be applied to other standards and recommendations, and extensions of any such standards and recommendations (including VVC and HEVC), whether existing or to be developed in the future. Unless otherwise specified or technically impossible, the aspects described in this document can be used individually or in combination.
[0089] In this document, various numerical values are used, such as, for example, {{1,0}, {3,1}, {1,1}}. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0090] FIG. 1 shows the encoder 100. Although variations of this encoder 100 are conceivable, for the sake of clarity, the encoder 100 will be described below without describing all the variations that may be expected.
[0091] The video sequence may undergo pre-encoding processing (101) before being encoded, for example, applying a color conversion to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), performing remapping of the input picture components to obtain a more resilient signal distribution for compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the pre-processing and added to the bitstream.
[0092] In encoder 100, the picture is encoded by the encoder elements described below. The picture to be encoded is divided (102) into units, for example, called coding units (CUs), and processed. Each unit is encoded using, for example, either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction (160) is performed. In the inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) whether to use either the intra mode or the inter mode for encoding the unit, and indicates the intra / inter decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (110) the block predicted from the original image block.
[0093] Next, the prediction residual is transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy encoded (145) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is directly encoded without applying the transformation process or the quantization process.
[0094] The encoder decodes the encoded blocks to provide reference data for further prediction. To decode the prediction residuals, the quantized transform coefficients are inverse quantized (140) and inverse transformed (150). The decoded prediction residuals and the predicted blocks are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed picture, for example, to perform deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).
[0095] FIG. 2 shows a block diagram of video decoder 200. In decoder 200, the bitstream is decoded by the decoder elements described below. Video decoder 200 generally executes a decoding path that is reverse to the encoding path as described in FIG. 1. Encoder 100 generally also performs decoding of the video as part of the encoding of the video data.
[0096] The input to the decoder includes a video bitstream, which can be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how the picture is partitioned. Thus, the decoder can partition the picture according to the decoded picture partitioning information (235). To decode the prediction residuals, the transform coefficients are inverse quantized (240) and inverse transformed (250). The decoded prediction residuals and the predicted blocks are combined (255) to reconstruct the image block. The predicted blocks can be obtained from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275) (270). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (280).
[0097] The decoded picture can further undergo post-decoding processing (285), for example, inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping that performs the reverse of the remapping process executed in the pre-encoding processing (101). In the post-decoding processing, metadata derived in the pre-encoding processing and signaled in the bitstream can be used.
[0098] FIG. 3 shows a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device including various components described below and is configured to execute one or more of the aspects described in this document. Examples of such a device include various electronic devices such as a personal computer, a laptop computer, a smartphone, a tablet computer, a digital multimedia set-top box, a digital television receiver, a personal video recording system, a connected home appliance, and a server, but are not limited thereto. The elements of system 1000 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input ports and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.
[0099] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described in this document, for example. The processor 1010 can include an embedded memory, an input / output interface, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, which can include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 can include, by way of non-limiting example, an internal storage device, an external storage device, and / or a network-accessible storage device.
[0100] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform an encoding function and / or a decoding function. As is known, a device can include one or both of an encoding module and a decoding module. Further, the encoder / decoder module 1030 can be implemented as a separate element of the system 1000 or can be incorporated within the processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0101] The program code loaded into the processor 1010 or the encoder / decoder 1030 to execute the various aspects described in this document can be stored in the storage device 1040 and then loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of various items during the execution of the processes described in this document. Such stored items include, but are not limited to, input video, decoded video or a part of the decoded video, bitstream, matrix, variable, and further equations, mathematical expressions, operations, and intermediate or final results from the processing of operation logic.
[0102] In some embodiments, the internal memory of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and to provide a working memory for the processing required during encoding or decoding. However, in another embodiment, an external memory of the processing device (the processing device can be, for example, either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, for example, dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as a working memory for video encoding operations and decoding operations such as MPEG-2, HEVC, or VVC (Versatile Video Coding).
[0103] Inputs to the elements of system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals wirelessly transmitted, for example, by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0104] In various embodiments, the input device of block 1130 has respective input processing elements known in the art. For example, the RF section can be associated with elements necessary to (i) select a desired frequency (also referred to as selecting a signal or band-limiting a signal to a particular frequency band), (ii) down-convert the selected signal, (iii) band-limit again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in certain embodiments, for example, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and re-filtering to a desired frequency band. In various embodiments, the order of the above-described (and other) elements is rearranged, some of these elements are deleted, and / or other elements that perform similar or different functions are added. Adding elements can include inserting elements between existing elements, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0105] Furthermore, the USB terminal and / or the HDMI terminal can each include an interface processor for connecting the system 1000 to other electronic devices via a USB connection and / or an HDMI connection. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, can be implemented, for example, in a separate input processing IC or in the processor 1010 as needed. Similarly, aspects of the USB or HDMI interface processing can be implemented, as needed, in a separate interface IC or in the processor 1010. The demodulated, error-corrected, and further demultiplexed stream is provided to various processing elements, such as the processor 1010 and an encoder / decoder 1030 that operates in combination with memory and storage elements, for processing the data stream as needed to present it on the output device.
[0106] The various elements of the system 1000 can be provided within an integrated housing, where the various elements are interconnected using an internal bus known in the art, including a suitable connection configuration 1140, such as an I2C bus, wiring, and a printed circuit board, and can transmit data to each other.
[0107] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or a network card, and the communication channel 1060 can be implemented, for example, within a wired medium and / or a wireless medium.
[0108] In various embodiments, data is streamed to system 1000 using a wireless network such as IEEE 802.11. The wireless signals of these embodiments are received via, for example, communication channel 1060 and communication interface 1050 adapted for Wi-Fi communication. The communication channel 1060 of these embodiments is generally connected to an access point or router that provides access to an external network including the Internet to enable streaming applications and other over-the-top communications. In another embodiment, a set-top box that distributes data via the HDMI connection of input block 1130 is used to provide streaming data to system 1000. In yet another embodiment, streaming data is provided to system 1000 using the RF connection of input block 1130.
[0109] System 1000 can provide output signals to various output devices including display 1100, speaker 1110, and other peripheral devices 1120. The other peripheral devices 1120 include, in various examples of the embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 can be integrated into a single unit with other components of system 1000, for example, in an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver, for example, a timing controller (T Con) chip.
[0110] Alternatively, for example, if the RF section of input 1130 is part of a separate set-top box, display 1100 and speaker 1110 can be separated from one or more of the other components. In various embodiments where display 1100 and speaker 1110 are external components, the output signal can be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0111] The embodiments can be implemented by computer software implemented by the processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type suitable for the technical environment and can be implemented using any suitable data storage technology, such as, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 can be of any type suitable for the technical environment and can include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0112] Various implementations include decoding. As used in this application, "decoding" can include all or part of the processing performed on the received encoded sequence, for example, to generate a final output suitable for display. In various embodiments, such processing can include, for example, one or more of the processing generally performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processing can, in addition to or alternatively, include processing performed by the decoder of various implementations described in this application, such as, for example, extracting the weight indexes used for various intra prediction reference arrays.
[0113] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader decoding process will become apparent based on the specific context of the description and is believed to be well understood by those skilled in the art.
[0114] The various implementations include encoding. As used in this application, "encoding", similar to the above description regarding "decoding", can include all or part of the processes performed on an input video sequence, for example, to generate an encoded bitstream. In various embodiments, such processes can include, for example, one or more of the processes commonly performed by an encoder, such as splitting, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes can include, in addition to or alternatively to these, for example, processes performed by the encoders of the various implementations described in this application, such as weighting of the intra prediction reference array.
[0115] As a further example, in one embodiment, "encoding" refers to only entropy encoding, in another embodiment, "encoding" refers to only differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader encoding process will become apparent based on the context of the particular description and is considered well understood by those skilled in the art.
[0116] Note that the syntactic elements used in this specification are terms for explanation. Therefore, they do not exclude the use of other syntactic element names.
[0117] When a figure is presented as a flowchart, it should be understood that the figure also provides a block diagram of the corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that the figure also provides a flowchart of the corresponding method / process.
[0118] In various embodiments, rate distortion calculation or rate distortion optimization is referred to. During encoding, typically, often a constraint on the computational complexity is given and the balance or trade-off between rate and distortion is considered. Rate distortion optimization is usually formulated to minimize a rate distortion function which is a weighted sum of rate and distortion. There are various approaches to solving the rate distortion optimization problem. For example, these approaches can be based on an extensive test of all encoding options including all considered modes or encoding parameter values, involving their encoding costs and a complete evaluation of the associated distortion of the reconstructed signal after encoding and decoding. Also, to reduce the encoding complexity, faster approaches can be used, in particular, using an approximation of distortion calculation based on a predicted or prediction residual signal instead of the reconstructed signal. These two approaches can also be used in combination, for example, using the approximate distortion for only some of the possible encoding options and the complete distortion for other encoding options. In another approach, only a subset of the possible encoding options is evaluated. More generally, many approaches employ any of various techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.
[0119] The implementations and aspects described in this specification can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (for example, described only as a method), the implementation of the described features can also be implemented in other forms (for example, an apparatus or a program). The apparatus can be implemented, for example, in appropriate hardware, software, and firmware. The method can be implemented, for example, in a processor, which refers to a general processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Further, the processor can include, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant ( "portable / personal digital assistant, PDA"), and other devices that facilitate the communication of information between end users.
[0120] References to "an embodiment" or "embodiment" or "an implementation" or "implementation", and other variations thereof, mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and other variations that appear in various places throughout this document do not necessarily all refer to the same embodiment.
[0121] Furthermore, this document may refer to "determining" various information. Determining information can include, for example, one or more of estimating information, calculating information, predicting information, or obtaining information from memory.
[0122] Furthermore, this document may refer to "accessing" various information. Accessing information can include, for example, one or more of receiving information, obtaining information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0123] Furthermore, this document may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information can include, for example, one or more of accessing information or obtaining information (e.g., from a memory). Further, "receiving" generally involves in some form during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information.
[0124] The use of any of " / ", "and / or", "at least one of", for example, in the cases of "A / B", "A and / or B", "at least one of A and B", is understood to be intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related technical fields, this can be extended for as many listed items as there are.
[0125] Also, as used herein, the term "signaling" particularly means indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals a particular one of a plurality of weights used in the intra prediction reference array. Thus, in some embodiments, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has that particular parameter and other parameters, implicit signaling, which does not perform the transmission but simply enables the decoder to recognize and select that particular parameter, can be used. By avoiding the transmission of actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above description pertains to the verb form of the word "signal", but the word "signal" can also be used as a noun herein.
[0126] As will be apparent to those skilled in the art, in an implementation form, for example, various signals can be generated that are formatted to convey information that can be stored or transmitted. Such information can include, for example, instructions for executing a method or data generated by one of the described implementation forms. For example, a signal can be formatted to convey the bitstream of the described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information conveyed by the signal can be, for example, analog information or digital information. The signal can be transmitted via various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0127] Embodiments can include, alone or in combination, one or more of the following features or entities across various different claim categories and types. · Enabling matrix-based intra prediction of coding units up to a determined size, regardless of the maximum transform size. · Enabling tiling of transform units when the size of a coding unit is larger than the maximum transform size. · Enabling low-frequency non-separable transform (LFNST) up to a determined size. · For any transform unit that indicates whether DCT2 is not used as the transform, if so indicated, including in the bitstream or checking the bitstream for syntax elements that do not permit LFNST. · A bitstream or signal that includes one or more of the described syntax elements or variations thereof. Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof. A television, set-top box, mobile phone, tablet, or other electronic device that performs in-loop filtering according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that performs in-loop filtering according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). A television, set-top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal including an encoded image and performs in-loop filtering according to any of the described embodiments. A television, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives a signal including an encoded image (e.g., using an antenna) and performs in-loop filtering according to any of the described embodiments.
[0128] A variety of other generalized and specific inventions and claims are also supported and contemplated throughout the present disclosure.
Claims
1. A method for video encoding, comprising: activating an encoding tool based on a maximum transform size; performing at least a portion of a discrete trigonometric transform on a subset of samples including a block; encoding the block using the activated encoding tool.
2. An apparatus, comprising: a processor configured to: activate an encoding tool based on a maximum transform size; perform at least a portion of a discrete trigonometric transform on a subset of samples including a block; encode the block using the activated encoding tool.
3. A method for video decoding, comprising: activating an encoding tool based on a maximum transform size; performing at least a portion of an inverse discrete trigonometric transform on a subset of samples including a block; decoding the block using the activated encoding tool.
4. An apparatus, comprising: a processor configured to: activate an encoding tool based on a maximum transform size; perform at least a portion of an inverse discrete trigonometric transform on a subset of samples including a block; decode the block using the activated encoding tool.
5. The method according to claim 1 or claim 3, or the apparatus according to claim 2 or claim 4, wherein the encoding tool is matrix-based intra prediction.
6. The method according to claim 1 or claim 3, or the apparatus according to claim 2 or claim 4, wherein the encoding tool is transform unit tiling.
7. The method according to claim 1 or claim 3, or the apparatus according to claim 2 or claim 4, wherein the encoding tool is low-frequency non-separable transform (LFNST) up to a determined transform size.
8. The method or apparatus according to claim 7, wherein the determined transform size is 32×32.
9. The method or apparatus according to claim 7, wherein the LFNST is not permitted if any transform unit does not use DCT2.
10. The encoding tool according to the method of claim 1 or claim 3, or the apparatus of claim 2 or claim 4, represented by a bit stream.
11. The method or apparatus according to claim 7, wherein the determined conversion size is 64×64.
12. A device, the apparatus according to any one of claims 4 to 11, and at least one of: (i) an antenna configured to receive a signal, the signal including a video block; (ii) a band limiter configured to limit the received signal to a frequency band including the video block; and (iii) a display configured to display an output representing the video block.
13. A non-transitory computer-readable medium including data content for reproduction using a processor, the data content being generated according to the method of claim 1 and any one of claims 5 to 11, or by the apparatus of claim 2 and any one of claims 5 to 11.
14. A signal including video data generated according to the method of claim 1 and any one of claims 5 to 11, or by the apparatus of claim 2 and any one of claims 5 to 11, for reproduction using a processor.
15. A computer program product including instructions that, when the program is executed by a computer, cause the computer to execute the method of claim 1, claim 3, and any one of claims 5 to 11.
Citation Information
Patent Citations
Maximum transform size control
WO2020180769A1
Block size dependent use of video coding mode
WO2021018083A1