Encoder for harmonizing matrix-based intra prediction and secondary conversion core selection, decoder, and handling method

The method addresses the challenge of harmonizing MIP and RST tools in video coding by determining intra-prediction modes and selecting appropriate transform cores, enhancing compression ratios with minimal quality loss.

JP2025186372APending Publication Date: 2025-12-23HUAWEI TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025152225
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-04-17
Filing Date
2025-09-12
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing video coding technologies face challenges in achieving high compression ratios with minimal sacrifice in picture quality, particularly in the harmonization of Matrix-based Intra Prediction (MIP) and Reduced Secondary Transform (RST) tools.

Method used

A method and apparatus for determining an intra-prediction mode and selecting a secondary transform based on the intra-prediction mode, disabling secondary transforms for MIP mode, and using predefined tables to select appropriate transform cores for non-MIP modes, including a newly trained transform core set for MIP mode.

Benefits of technology

Enhances video compression by harmonizing MIP and RST tools, improving compression ratios with minimal impact on picture quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025186372000001_ABST
    Figure 2025186372000001_ABST
Patent Text Reader

Abstract

To provide an encoder for harmonizing matrix-based intra prediction and secondary conversion core selection, a decoder and a handling method.SOLUTION: A coding method implemented by a decoding device or an encoding device includes: a step 1601 for determining an intra-prediction mode for a current block; and a step 1603 for determining a selection of secondary conversion for the current block according to the intra-prediction mode determined for the current block.SELECTED DRAWING: Figure 15
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Technical Field

[0002] FIELD OF THE INVENTION Embodiments of the present disclosure relate generally to the field of picture processing, and more particularly to intra prediction. [Background technology]

[0003] Video coding (video encoding and decoding) is used in a wide range of digital video applications, such as broadcast digital TV, video transmission over the Internet and mobile networks, real-time conversation applications such as video chat and video conferencing, DVD and Blu-ray discs, video content collection and editing systems, and video cameras for security applications.

[0004] The amount of video data required to render even a relatively short video can be substantial, which can create difficulties when the data is streamed or otherwise communicated over communications networks with limited bandwidth capacity. Thus, video data is typically compressed before being communicated over modern telecommunications networks. Video size can also be an issue when the video is stored on a storage device, as memory resources may be limited. Video compression devices often use software and / or hardware at the source to code the video data before transmission or storage, thereby reducing the amount of data needed to represent a digital video image. The compressed data is then received at the destination by a video decompression device, which decodes the video data. Due to limited network resources and the ever-increasing demand for higher video quality, improved compression and decompression techniques that improve compression ratios with little or no sacrifice in picture quality are desirable. Summary of the Invention

[0005] The present disclosure seeks to alleviate or solve the above-mentioned problems.

[0006] Embodiments of the present application provide apparatus and methods for encoding and decoding according to the independent claims.

[0007] The present disclosure discloses a coding method implemented by a decoding device or an encoding device, which includes steps of determining an intra-prediction mode of a current block; and determining a selection of a secondary transform for the current block based on the intra-prediction mode determined for the current block.

[0008] In this manner, the method of this disclosure determines the intra-prediction mode of the current block, and determines whether and how to perform a secondary transform of the current block based on the determined intra-prediction mode.

[0009] In the above-described method, determining the selection of the secondary transform core for the secondary transform of the current block may be based on an intra-prediction mode index of the current block.

[0010] In the above-described method, if the current block is not predicted using MIP, Matrix-based Intra Prediction, mode, a secondary transform core may be selected for the secondary transform of the current block.

[0011] Thus, when an intra-predicted block is predicted using MIP mode, for example, the value of the MIP flag may be used to indicate whether the block is predicted using MIP mode or not, and the secondary transform is disabled for this intra-predicted block; in other words, the value of the secondary transform index is set to 0 or the secondary transform index does not need to be decoded from the bitstream.

[0012] In this way, a harmonization between MIP and RST tools is achieved in terms of secondary conversion core selection.

[0013] The above method may further include: disabling a secondary transform of the current block when the current block is predicted using the MIP mode.

[0014] In the above-described method, disabling the secondary transformation of the current block may include setting a value of the secondary transformation indication information for the current block to a default value.

[0015] In the above method, whether the current block is predicted using the MIP mode may be indicated according to the value of the MIP indication information.

[0016] In the above method, if the current block is not predicted using MIP mode, the secondary transform may be selected according to the following table. [Table 1] where: sps_lfnst_enabled_flag equal to 1 specifies that lfnst_idx may be present in intra-coding unit syntax, and sps_lfnst_enabled_fag equal to 0 specifies that lfnst_idx is not present in intra-coding unit syntax; intra_mip_flag[x0][y0] equal to 1 specifies that the intra prediction type for the luma sample is matrix-based intra prediction. intra_mip_flag[x0][y0] equal to 0 specifies that the intra prediction type for the luma sample is not matrix-based intra prediction; if intra_mip_flag[x0][y0] is not present, it is inferred to be equal to 0. lfnst_idx specifies whether and which one of the two low-frequency non-separable transform kernels in the selected transform set is used. lfnst_idx equal to 0 specifies that no low-frequency non-separable transform is used in the current coding unit. If lfnst_idx is not present, it is inferred to be equal to 0; transform_skip_flag[x0][y0][cIdx] specifies whether a transform is applied to the associated transform block.

[0017] The above-mentioned method may further include: obtaining an intra-prediction mode index for the current block according to a matrix-based intra-prediction, MIP, mode index of the current block and a size of the current block; and selecting a secondary transform core for secondary transform of the current block based on the intra-prediction mode index of the current block.

[0018] Thus, during the process of transform core selection for secondary transforms, if a block is predicted using MIP mode, one of the secondary transform core sets is considered to be used for this block.

[0019] In the above-mentioned method, the intra-prediction mode index of the current block may be obtained according to a mapping relationship between the MIP mode index and the size of the current block, and the mapping relationship may be represented according to a predefined table.

[0020] The above method may further include using a secondary transform core for a secondary transform of the current block if the current block is predicted using a matrix-based intra prediction, MIP, mode.

[0021] In the above-described method, the secondary conversion core may be one of the secondary conversion cores used for the non-MIP mode.

[0022] In the above methods, the secondary conversion cores may be different from any of the secondary conversion cores used for the non-MIP modes.

[0023] In the above-described method, if the current block is predicted using MIP mode, a lookup table may be used to map MIP mode indices to normal intra-mode indices, and the secondary transform core set may be selected based on the normal intra-mode indices.

[0024] In the above method, MIP mode indices may be mapped to regular intra mode indices based on the following table: [Table 2] Here, the selection of the secondary transformation set may be performed according to the following table: [Table 3]

[0025] The method may further include: providing four transform core sets having transform core set indices 0, 1, 2, and 3, respectively, where each transform core set of the four transform core sets can include two transforms; providing a fifth transform core set having a transform core set index 4, where the fifth transform core set has the same dimensions as the transform core sets having core set indices 0 to 3, and the fifth transform core set is newly trained based on the same machine learning method and input training set for MIP mode; and selecting a reduced secondary transform (RST) matrix by determining the transform core set of the five transform core sets to be applied to a current block according to an intra prediction mode of the current block, as follows: if the current intra block is a multi-prediction block using CCLM, a cross-component linear model (CLM), ... If the current intra block is predicted using MIP mode, select the transform core set with transform core set index 0; if the current intra block is predicted using MIP mode, select the transform core set with transform core set index 4; otherwise, select the transform core set using the value of the current block's intra prediction mode index and the following table: [Table 4]

[0026] The described method may further include: providing four transform core sets having transform core set indices 0, 1, 2, and 3, respectively, where each transform core set of the four transform core sets can include two transforms; providing a fifth transform core set having a transform core set index 4, where the fifth transform core set has the same dimensions as the transform core sets having core set indices 0 to 3, and the fifth transform core set is newly trained based on the same machine learning method and input training set for MIP mode; and selecting a reduced secondary transform (RST) matrix by determining the transform core set of the five transform core sets to be applied to a current block according to an intra prediction mode of the current block, as follows: if the current intra block is a multi-prediction block, the current intra block is a multi-prediction block using CCLM, Cross-Component Linear Model (CCM), or a multi-prediction block using CCLM, Cross-Component Linear Model (CCL). If the current intra block is predicted using MIP mode, select the transform core set with transform core set index 0; if the current intra block is predicted using MIP mode, select the transform core set with transform core set index 4; otherwise, select the transform core set using the value of the current block's intra prediction mode index and the following table: [Table 5]

[0027] Thus, during the process of transform core selection for secondary transforms, if a block is predicted using MIP mode, the trained secondary transform core set is considered to be used for this block. The trained secondary transform core set may be different from the transform core set in the above example.

[0028] The present disclosure further provides an encoder comprising processing circuitry for performing the above method when implemented by an encoding device.

[0029] The present disclosure further provides a decoder comprising processing circuitry for performing the above method when implemented by a decoding device.

[0030] The present disclosure further provides a computer program product comprising program code for performing the above method.

[0031] The present disclosure further provides a decoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, wherein the programming, when executed by the processors, configures the decoder to perform the above-described method when implemented by a decoding device.

[0032] The present disclosure further provides an encoder comprising one or more processors and a non-transitory computer-readable storage medium coupled to the processors and storing programming for execution by the processors, wherein the programming, when executed by the processors, configures the encoder to perform the above-described method when implemented by an encoding device.

[0033] The present disclosure further provides a decoder having a determination unit configured to determine an intra-prediction mode of a current block; and a selection unit configured to determine a selection of a secondary transformation of the current block based on the intra-prediction mode determined for the current block.

[0034] The present disclosure further provides an encoder having a determination unit configured to determine an intra-prediction mode of a current block; and a selection unit configured to determine a selection of a secondary transform of the current block based on the intra-prediction mode determined for the current block.

[0035] These and other objects are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, the description and the drawings.

[0036] The method according to the first aspect of the invention may be performed by an apparatus according to the third aspect of the invention. Further features and implementations of the method according to the third aspect of the invention correspond to the features and implementations of the apparatus according to the first aspect of the invention.

[0037] The method according to the second aspect of the invention may be performed by an apparatus according to the fourth aspect of the invention. Further features and implementations of the method according to the fourth aspect of the invention correspond to the features and implementations of the apparatus according to the second aspect of the invention.

[0038] According to a fifth aspect, the present invention relates to an apparatus for decoding a video stream, comprising a processor and a memory, the memory storing instructions for causing the processor to perform the method according to the first aspect.

[0039] According to a sixth aspect, the present invention relates to an apparatus for encoding a video stream, comprising a processor and a memory, the memory storing instructions for causing the processor to perform the method according to the second aspect.

[0040] According to a seventh aspect, there is proposed a computer-readable storage medium storing instructions that, when executed, cause one or more configured processors to code video data, the instructions causing said one or more processors to perform a method according to the first aspect or the second aspect or any possible embodiment of the first or second aspect.

[0041] According to an eighth aspect, the present invention relates to a computer program comprising program code for performing, when the computer program is run on a computer, a method according to the first aspect or the second aspect or any possible embodiment of the first aspect or the second aspect.

[0042] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0043] In the following, embodiments of the invention will be explained in more detail with reference to the accompanying drawings and figures. [Figure 1A] 1 is a block diagram illustrating an example of a video coding system configured to implement embodiments of the present invention. [Figure 1B] FIG. 2 is a block diagram illustrating another example of a video coding system configured to implement embodiments of the present invention. [Figure 2] FIG. 1 is a block diagram illustrating an example of a video encoder configured to implement embodiments of the present invention. [Figure 3] FIG. 1 is a block diagram illustrating an exemplary structure of a video decoder configured to implement embodiments of the present invention; [Figure 4] FIG. 1 is a block diagram illustrating an example of an encoding or decoding device. [Figure 5] FIG. 10 is a block diagram illustrating another example of an encoding device or a decoding device. [Figure 6] FIG. 10 is a block diagram illustrating an example of MIP mode matrix multiplication for a 4x4 block. [Figure 7] FIG. 10 is a block diagram illustrating an example of MIP mode matrix multiplication for an 8x8 block. [Figure 8] FIG. 10 is a block diagram illustrating an example of MIP mode matrix multiplication for an 8x4 block. [Figure 9]FIG. 10 is a block diagram illustrating an example of MIP mode matrix multiplication for a 16×16 block. [Figure 10] FIG. 1 is a block diagram illustrating an exemplary secondary conversion process. [Figure 11] 1 is a block diagram illustrating an exemplary secondary transform core multiplication process of an encoding and decoding device. [Figure 12] FIG. 1 is a block diagram illustrating an exemplary quadratic transform core dimension reduction from 16×64 to 16×48. [Figure 13] FIG. 1 is a block diagram illustrating an exemplary MIP MPM reconstruction based on the positions of neighboring blocks. [Figure 14] FIG. 10 is a block diagram illustrating another exemplary MIP MPM reconstruction based on the positions of neighboring blocks. [Figure 15] 1 illustrates a method implemented by a decoding device or an encoding device according to the present disclosure. [Figure 16] 1 shows an encoder according to the present disclosure. [Figure 17] 1 shows a decoder according to the present disclosure. [Figure 18] 31 is a block diagram illustrating an exemplary structure of a content supply system 3100 for implementing a content distribution service. [Figure 19] FIG. 2 is a block diagram showing the structure of an example of a terminal device.

[0044] In the following, identical reference signs refer to identical or at least functionally equivalent features, unless expressly specified otherwise. DETAILED DESCRIPTION OF THE INVENTION

[0045] In the following description, reference is made to the accompanying drawings, which form a part of this disclosure and which show, by way of illustration, specific aspects of embodiments of the present invention or in which embodiments of the present invention may be practiced. It is understood that embodiments of the present invention may be practiced in other aspects and may include structural or logical changes not shown in the drawings. Therefore, the following detailed description is not to be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.

[0046] For example, it is understood that disclosure related to a described method may also be true of a corresponding apparatus or system configured to perform that method, and vice versa. For example, when one or more particular method steps are described, a corresponding apparatus may include one or more units, e.g., functional units, for performing the described one or more method steps (e.g., one unit performing one or more steps, or multiple units each performing one or more of the steps), even if such one or more units are not explicitly described or shown. On the other hand, for example, when a particular apparatus is described based on one or more units, e.g., functional units, a corresponding method may include a step for performing the function of the one or more units (e.g., one step for performing the function of the one or more units, or multiple steps for performing one or more functions of the multiple units, respectively), even if such one or more steps are not explicitly described or shown. Furthermore, it is understood that features of various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise noted.

[0047] Video coding typically refers to the processing of a sequence of pictures to form a video or video sequence. Instead of the term "picture," the terms "frame" or "image" are sometimes used synonymously in the field of video coding. Video coding (or coding in general) has two parts: video encoding and video decoding. Video encoding is performed at the source side and typically involves processing the original video picture (e.g., by compression) to reduce the amount of data required to represent the video picture (for more efficient storage and / or transmission). Video decoding is performed at the destination side and typically involves the reverse process compared to the encoder to reconstruct the video picture. Embodiments referring to "coding" a video picture (or pictures in general) are understood to relate to "encoding" or "decoding" the video picture or respective video sequence. The combination of the encoding and decoding parts is also called a codec (coding and decoding).

[0048] In the case of lossless video coding, the original video picture can be reconstructed, i.e., the reconstructed video picture has the same quality as the original video picture (assuming there are no transmission or other data losses during storage or transmission). In the case of lossy video coding, further compression, e.g., by quantization, is performed to reduce the amount of data representing the video picture, which cannot be perfectly reconstructed at the decoder, i.e., the quality of the reconstructed video picture is lower or worse than the quality of the original video picture.

[0049] Some video coding standards belong to the group of "lossy hybrid video codecs" (i.e., they combine spatial and temporal prediction in the sample domain with 2D transform coding to apply quantization in the transform domain). Each picture in a video sequence is typically divided into a set of non-overlapping blocks, and coding is typically performed at the block level. In other words, in an encoder, video is typically processed, i.e., encoded, at the block (video block) level. This is done, for example, by generating a predictive block using spatial (intra-picture) and / or temporal (inter-picture) prediction, subtracting the predictive block from a current block (the block currently being processed / to be processed) to obtain a residual block, transforming the residual block, and quantizing the residual block in the transform domain to reduce the amount of data to be transmitted (compression). In a decoder, the reverse process is applied to the encoded or compressed block to reconstruct the current block for representation. Additionally, the encoder replicates the decoder processing loop, so that both generate the same predictions (eg, intra-prediction and inter-prediction) and / or reconstructions for processing, ie, coding, of subsequent blocks.

[0050] In the following embodiment of a video coding system 10, a video encoder 20 and a video decoder 30 are described based on FIGS.

[0051] 1A is a schematic block diagram illustrating an example coding system 10, e.g., video coding system 10 (or coding system 10 for short), that can utilize the techniques of the present application. A video encoder 20 (or encoder 20 for short) and a video decoder 30 (or decoder 30 for short) of video coding system 10 represent example devices that can be configured to perform the techniques according to various examples described herein.

[0052] 1A, coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 that decodes the encoded picture data 13. Source device 12 includes an encoder 20 and may additionally or optionally include a picture source 16, a preprocessor (or pre-processing unit) 18, for example, a picture preprocessor 18, and a communication interface or unit 22.

[0053] Picture source 16 may comprise or be any kind of picture capture device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any kind of other device for obtaining and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). Picture source may also be any kind of memory or storage for storing any of the above pictures.

[0054] To distinguish between the preprocessor 18 and the processing performed by the preprocessing unit 18, the picture or picture data 17 may be referred to as a raw picture or raw picture data 17.

[0055] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain a preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It will be understood that the preprocessing unit 18 may be an optional component.

[0056] A video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21 (as described in further detail below, e.g., with reference to FIG. 2).

[0057] The communications interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) over the communications channel 13 to another device, for example the destination device 14 or any other device, for storage or direct reconstruction.

[0058] The destination device 14 has a decoder 30 (e.g., a video decoder 30) and may additionally or optionally have a communications interface or communications unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0059] The communications interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or a further processed version thereof), e.g., directly from the source device 12 or from any other source, e.g., a storage device, e.g., an encoded picture data storage device, and to provide the encoded picture data 21 to a decoder 30.

[0060] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., a direct wired or wireless connection, or via any type of network, e.g., a wired or wireless network or any combination thereof, or any type of private and public network, or any type of combination thereof.

[0061] The communications interface 22 may be configured, for example, to package the encoded picture data 21 into a suitable format, e.g., packets, and / or process the encoded picture data using any type of transmission encoding or processing for transmission over a communications link or network.

[0062] The counterpart communication interface 28 of the communication interface 22 may be configured, for example, to receive the transmitted data and process the transmitted data using any type of corresponding transmission decoding or processing and / or unpackaging to obtain encoded picture data 21.

[0063] Both communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces, as indicated by the arrow for communication channel 13 pointing from source device 12 to destination device 14 in FIG. 1A, or as bidirectional communication interfaces, and may be configured to send and receive messages, e.g., to set up a connection, receive, acknowledge, and exchange messages, e.g., to set up a communication link and / or any other information related to data transmission, e.g., encoded picture data transmission.

[0064] The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or decoded pictures 31 (further details are described below, e.g., with reference to Figure 3 or Figure 5).

[0065] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), e.g., the decoded picture 31, to obtain post-processed picture data 33, e.g., the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare the decoded picture data 31 for display, e.g., by a display device 34.

[0066] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g., to a user or viewer. The display device 34 may be or include any type of display for presenting the reconstructed picture, e.g., an integrated or external display or monitor. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0067] 1A depicts source device 12 and destination device 14 as separate devices, an embodiment of the devices may include both or both functionality: source device 12 or corresponding functionality and destination device 14 or corresponding functionality. In such an embodiment, source device 12 or corresponding functionality and destination device 14 or corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0068] As will be apparent to those skilled in the art based on the above description, the functionality of different units or the presence and (exact) division of functions within source device 12 and / or destination device 14 as shown in FIG. 1A may vary depending on the actual device and application.

[0069] Encoder 20 (e.g., video encoder 20), or decoder 30 (e.g., video decoder 30), or both encoder 20 and decoder 30 may be implemented via processing circuitry such as that shown in FIG. 1B, e.g., one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, dedicated video coding, or any combination thereof. Encoder 20 may be implemented via processing circuitry 46 to embody various modules discussed with respect to encoder 20 of FIG. 2 and / or any other encoder system or subsystem described herein. Decoder 30 may be implemented via processing circuitry 46 to embody various modules discussed with respect to decoder 30 of FIG. 3 and / or any other decoder system or subsystem described herein. The processing circuitry may be configured to perform various operations, as described below, as shown in FIG. 5. If the techniques are implemented partially in software, a device may store instructions for the software on a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either video encoder 20 or video decoder 30 may also be integrated into a single device as part of a combined encoder / decoder (codec), for example, as shown in FIG. 1B.

[0070] Source device 12 and destination device 14 may include any of a wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or content delivery server), a broadcast receiver device, a broadcast transmitter device, etc., and may use no operating system or any type of operating system.

[0071] In some cases, source device 12 and destination device 14 may be equipped for wireless communication. Thus, source device 12 and destination device 14 may be wireless communication devices. In some cases, video coding system 10 shown in FIG. 1A is merely an example, and the present technology may be applied to video coding scenarios (e.g., video encoding or video decoding) that do not necessarily involve data communication between an encoding device and a decoding device. In other examples, data may be retrieved from local memory, streamed over a network, etc. A video encoding device may encode data and store it in memory, and / or a video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other but simply encode data to memory and / or retrieve data from memory and decode it.

[0072] For ease of description, embodiments of the present invention are described herein with reference to, for example, High Efficiency Video Coding (HEVC) or Versatile Video Coding (VVC) reference software, the next-generation video coding standards developed by the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG) Joint Collaboration Team on Video Coding (JCT-VC). Those skilled in the art will understand that embodiments of the present invention are not limited to HEVC or VVC.

[0073] Encoders and encoding methods FIG. 2 shows a schematic block diagram of an exemplary video encoder 20 configured to implement the techniques of the present application. In the example of FIG. 2, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoded picture buffer (DPB) 230, a mode selection unit 260, an entropy encoding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter prediction unit 244, an intra prediction unit 254, and a segmentation unit 262. The inter prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in FIG. 2 may also be referred to as a hybrid video encoder or a video encoder using a hybrid video codec.

[0074] The residual calculation unit 204, the transform processing unit 206, the quantization unit 208, and the mode selection unit 260 may be referred to as forming a forward signal path of the encoder 20, while the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the buffer 216, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may be referred to as forming a backward signal path of the video encoder 20. The backward signal path of the video encoder 20 corresponds to the signal path of a decoder (see video decoder 30 in FIG. 3). The inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the loop filter 220, the decoded picture buffer (DPB) 230, the inter prediction unit 244, and the intra prediction unit 254 may also be referred to as forming an “embedded decoder” of the video encoder 20.

[0075] Pictures and picture divisions (pictures and blocks) Encoder 20 may be configured to receive, for example, via input 201, a picture 17 (or picture data 17), e.g., a picture of a video or a sequence of pictures forming a video sequence. The received picture or picture data may be a preprocessed picture 19 (or preprocessed picture data 19). For simplicity, the following description refers to picture 17. Picture 17 may also be referred to as a current picture or a picture to be coded (particularly in video coding, to distinguish the current picture from other pictures, e.g., previously encoded and / or decoded pictures of the same video sequence, i.e., a video sequence that also includes the current picture).

[0076] A (digital) picture is or can be considered as a two-dimensional array or matrix of intensity-valued samples. The samples in the array may be called pixels (short for picture element) or picture elements. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the picture. For color representation, three color components are typically used; that is, a picture is represented by or can contain three sample arrays. In an RBG format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, e.g., YCbCr, which contains a luminance component denoted Y (sometimes L is used instead) and two chrominance components denoted Cb and Cr. The luminance (or luma for short) component Y represents brightness or gray-level intensity (e.g., as in a grayscale picture), while the two chrominance (or chroma for short) components Cb and Cr represent chromaticity or color information components. Thus, a picture in YCbCr format contains a luminance sample array of luminance sample values ​​(Y) and two chrominance sample arrays of chrominance values ​​(Cb and Cr). A picture in RGB format can be converted or translated to YCbCr format, and vice versa. This process is also known as color conversion or translation. If a picture is monochrome, it may contain only a luminance sample array. Thus, a picture can be, for example, an array of luma samples in monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0077] Embodiments of video encoder 20 may include a picture partition unit (not shown in FIG. 2) configured to partition picture 17 into multiple (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC) or coding tree blocks (CTBs) or coding tree units (CTUs) (H.265 / HEVC and VVC). The picture partition unit may be configured to use the same block size and a corresponding grid defining the block size for all pictures of a video sequence, or to vary the block size between pictures or between subsets or groups of pictures, and to partition each picture into corresponding blocks.

[0078] In further embodiments, the video encoder may be configured to directly receive blocks 203 of picture 17, e.g., one, some, or all of the blocks forming picture 17. Picture blocks 203 may also be referred to as current picture blocks or picture blocks to be coded.

[0079] Similar to picture 17, picture block 203 is, or can be, considered as a two-dimensional array or matrix of samples having intensity values ​​(sample values), albeit with smaller dimensions than picture 17. In other words, block 203 may include, for example, one sample array (e.g., a luma array in the case of a monochrome picture 17, or a luma or chroma array in the case of a color picture), or three sample arrays (e.g., a luma array and two chroma arrays in the case of a color picture 17), or any other number and / or type of arrays depending on the applied color format. The number of samples in the horizontal and vertical directions (or axes) of block 203 defines the size of block 203. Thus, a block may be, for example, an M×N (M columns by N rows) array of samples or an M×N array of transform coefficients.

[0080] An embodiment of video encoder 20 such as that shown in FIG. 2 may be configured to encode picture 17 block by block, eg, encoding and prediction is performed block by block 203.

[0081] 2 may be further configured to partition and / or encode pictures using slices (also referred to as video slices), where a picture may be partitioned into or encoded using one or more (typically non-overlapping) slices, each of which may include one or more blocks (e.g., CTUs).

[0082] 2 may be further configured to divide and / or encode a picture using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), where a picture may be divided into or encoded using one or more (typically non-overlapping) tile groups, each of which may contain, for example, one or more blocks (e.g., CTUs) or one or more tiles. Each tile may be, for example, rectangular and may contain one or more blocks (e.g., CTUs), e.g., full or partial blocks.

[0083] Residual calculation The residual calculation unit 204 may be configured to calculate a residual block 205 (also referred to as residual 205) based on the picture block 203 and the prediction block 265 (further details about the prediction block 265 will be described below), for example by subtracting sample values ​​of the prediction block 265 from sample values ​​of the picture block 203 sample by sample (pixel by pixel) to obtain the residual block 205 in the sample domain.

[0084] conversion The transform processing unit 206 may be configured to apply a transform, such as a discrete cosine transform (DCT) or a discrete sine transform (DST), to the sample values ​​of the residual block 205 to obtain transform coefficients 207 in the transform domain. The transform coefficients 207 may also be referred to as transform residual coefficients and represent the residual block 205 in the transform domain.

[0085] The transform processing unit 206 may be configured to apply an integer approximation of a DCT / DST, such as the transform specified for H.265 / HEVC. Compared to an orthogonal DCT transform, such an integer approximation is typically scaled by a factor. To preserve the norm of the residual blocks processed by the forward and inverse transforms, an additional scaling factor is applied as part of the transform process. The scaling factor is typically selected based on certain constraints, such as the scaling factor being a power of two for shift operations, the bit depth of the transform coefficients, a trade-off between accuracy and implementation cost, etc. A specific scaling factor may be specified, for example, for the inverse transform, e.g., by the inverse transform processing unit 212 (and the corresponding inverse transform, e.g., by the inverse transform processing unit 312 in the video decoder 30), and a corresponding scaling factor for the forward transform, e.g., by the transform processing unit 206 in the encoder 20, may be specified accordingly.

[0086] An embodiment of video encoder 20 (and, in particular, transform processing unit 206) may be configured to output transform parameters, e.g., one or more transform types, e.g., directly or encoded or compressed via entropy encoding unit 270, so that, for example, video decoder 30 may receive and use the transform parameters for decoding.

[0087] quantization The quantization unit 208 may be configured to quantize the transform coefficients 207, for example by applying scalar quantization or vector quantization, to obtain quantized coefficients 209. The quantized coefficients 209 may also be referred to as quantized transform coefficients 209 or quantized residual coefficients 209.

[0088] The quantization process may reduce the bit depth associated with some or all of the transform coefficients 207. For example, an n-bit transform coefficient may be rounded to an m-bit transform coefficient during quantization, where n is greater than m. The degree of quantization may be modified by adjusting the quantization parameter (QP). For example, for scalar quantization, different scaling may be applied to achieve finer or coarser quantization. A smaller quantization step size corresponds to finer quantization, while a larger quantization step size corresponds to coarser quantization. The applicable quantization step size may be indicated by the quantization parameter (QP). The quantization parameter may, for example, be an index into a predefined set of applicable quantization step sizes. For example, a small quantization parameter may correspond to finer quantization (small quantization step size) and a large quantization parameter may correspond to coarser quantization (large quantization step size), or vice versa. Quantization may involve division by a quantization step size, and corresponding and / or inverse dequantization, e.g., by the inverse quantization unit 210, may involve multiplication by the quantization step size. Some standards, e.g., HEVC, embodiments may be configured to use a quantization parameter to determine the quantization step size. Generally, the quantization step size may be calculated based on the quantization parameter using a fixed-point approximation of a formula involving division. Additional scaling factors may be introduced for quantization and dequantization to restore the norm of the residual block, which may be modified due to the scaling used in the fixed-point approximation of the formula for the quantization parameter and the quantization step size. In one example implementation, the scaling of the inverse transform and dequantization may be combined. Alternatively, customized quantization tables may be used and signaled from the encoder to the decoder, e.g., in the bitstream. Quantization is a lossy operation, and the loss increases with increasing quantization step size.

[0089] Embodiments of video encoder 20 (and, in particular, quantization unit 208) may be configured to output a quantization parameter (QP), e.g., directly or encoded via entropy encoding unit 270, so that, for example, video decoder 30 can receive and apply the quantization parameter for decoding.

[0090] inverse quantization Inverse quantization unit 210 is configured to apply the inverse quantization of quantization unit 208 to the quantized coefficients, e.g., by applying the inverse of the quantization scheme applied by quantization unit 208, based on or using the same quantization step size as quantization unit 208, to obtain dequantized coefficients 211. The dequantized coefficients 211 may also be referred to as dequantized residual coefficients 211 and correspond to transform coefficients 207, although they are typically not identical to the transform coefficients due to loss due to quantization.

[0091] Inverse transformation The inverse transform processing unit 212 is configured to apply an inverse transform of the transform applied by the transform processing unit 206, such as an inverse discrete cosine transform (DCT) or an inverse discrete sine transform (DST) or other inverse transform, to obtain a reconstructed residual block 213 (or corresponding dequantized coefficients 213) in the sample domain. The reconstructed residual block 213 may also be referred to as a transform block 213.

[0092] Reconstruction The reconstruction unit 214 (e.g., adder or summer 214) is configured to add the transform block 213 (i.e., the reconstructed residual block 213) to the prediction block 265 to obtain the reconstructed block 215 in the sample domain, for example, by adding the sample values ​​of the reconstructed residual block 213 and the sample values ​​of the prediction block 265 sample by sample.

[0093] filtering Loop filter unit 220 (or "loop filter" 220 for short) is configured to filter reconstructed block 215 to obtain filtered block 221, or generally, to filter reconstructed samples to obtain filtered samples. The loop filter unit is configured, for example, to smooth pixel transitions or otherwise improve video quality. Loop filter unit 220 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, such as a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 220 is shown in FIG. 2 as an in-loop filter, in other configurations, loop filter unit 220 may be implemented as a post-loop filter. Filtered block 221 may also be referred to as a filtered reconstructed block 221.

[0094] Embodiments of video encoder 20 (and, in particular, loop filter unit 220) may be configured to output loop filter parameters (such as sample adaptive offset information), e.g., directly or encoded via entropy encoding unit 270, so that, for example, decoder 30 can receive and apply the same loop filter parameters or respective loop filters for decoding.

[0095] Decoded Picture Buffer The decoded picture buffer (DPB) 230 may be a memory that stores reference pictures, or generally reference picture data, for encoding video data by the video encoder 20. The DPB 230 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The decoded picture buffer (DPB) 230 may be configured to store one or more filtered blocks 221. The decoded picture buffer 230 may further be configured to store other previously filtered blocks, e.g., previously reconstructed and filtered blocks 221, of the same current picture or of a different picture, e.g., a previously reconstructed picture, and may provide a complete previously reconstructed, i.e., decoded, picture (and corresponding reference blocks and samples) and / or a partially reconstructed current picture (and corresponding reference blocks and samples), e.g., for inter-prediction. The decoded picture buffer (DPB) 230 may be configured to store one or more unfiltered reconstructed blocks 215, for example if the reconstructed blocks 215 are not filtered by the loop filter unit 220, or unfiltered reconstructed samples in general, or any other further processed version of the reconstructed blocks or samples.

[0096] Mode Selection (Segmentation and Prediction) The mode selection unit 260 includes a partitioning unit 262, an inter-prediction unit 244, and an intra-prediction unit 254, and is configured to receive or obtain original picture data, e.g., original block 203 (current block 203 of current picture 17), and reconstructed picture data, e.g., filtered and / or unfiltered reconstructed samples or blocks, of the same (current) picture and / or from one or more previously decoded pictures, e.g., from a decoded picture buffer 230 or other buffer (e.g., a line buffer, not shown). The reconstructed picture data is used as reference picture data for prediction, e.g., inter-prediction or intra-prediction, to obtain a prediction block 265 or predictor 265.

[0097] The mode selection unit 260 may be configured to determine or select a partitioning (including no partitioning) and prediction mode (e.g., intra or inter prediction mode) for the current block prediction mode and generate a corresponding prediction block 265 used for calculating the residual block 205 and reconstructing the reconstruction block 215.

[0098] Embodiments of the mode selection unit 260 may be configured to select a partitioning and prediction mode (e.g., from those supported by or available to the mode selection unit 260) that provides the best match, i.e., the smallest residual (smallest residual means better compression for transmission or storage) or the smallest signaling overhead (smallest signaling overhead means better compression for transmission or storage), or that considers or balances both. The mode selection unit 260 may also be configured to determine the partitioning and prediction mode based on rate-distortion optimization (RDO), i.e., select the prediction mode that provides the smallest rate-distortion. Terms such as “best,” “minimum,” “optimum,” etc. in this context do not necessarily refer to an overall “best,” “minimum,” “optimum,” etc., but may also refer to the termination or fulfillment of a selection criterion, such as a value above or below a threshold, or other constraint, potentially leading to a “non-optimal selection,” but reducing complexity and processing time.

[0099] In other words, the partitioning unit 262 may be configured to partition the block 203 into smaller block partitions or sub-blocks (which still constitute blocks), e.g., using quadtree partitioning (QT), binary partitioning (BT), or ternary tree partitioning (TT), or any combination thereof, recursively, and perform prediction on each of the block partitions or sub-blocks, where the mode selection includes selecting a tree structure of the partitioned block 203, and a prediction mode is applied to each of the block partitions or sub-blocks.

[0100] The partitioning (eg, by partitioning unit 260) and prediction processes (eg, by inter-prediction unit 244 and intra-prediction unit 254) performed by exemplary video encoder 20 are described in more detail below.

[0101] Partitioning The splitting unit 262 can partition (or divide) the current block 203 into smaller partitions, e.g., square or rectangular sized blocks. These smaller blocks (which may also be called subblocks) may then be further divided into even smaller partitions. This is also called tree splitting or hierarchical tree splitting; for example, a root block at root tree level 0 (hierarchical level 0, depth 0) may be recursively split into two or more blocks at the next, lower tree level, e.g., tree level 1 (hierarchical level 1, depth 1), which may then be split again into two or more blocks at the next, lower level, e.g., tree level 2 (hierarchical level 2, depth 2), and so on, until the splitting is terminated because a termination criterion is met, e.g., a maximum tree depth or a minimum block size is reached. Blocks that are not further split are also called leaf blocks or leaf nodes of the tree. A tree using a split into two partitions is called a binary tree (BT), a tree using a split into three partitions is called a ternary tree (TT), and a tree using a split into four partitions is called a quad tree (QT).

[0102] As mentioned above, the term "block" as used herein may refer to a portion of a picture, particularly a square or rectangular portion. For example, with reference to HEVC and VVC, a block may be or correspond to a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), and a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a transform block (TB), or a prediction block (PB).

[0103] For example, a coding tree unit (CTU) may be or include a CTB of luma samples for a picture having three sample arrays, two corresponding CTBs of chroma samples, or a CTB of samples for a picture coded using three distinct color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding tree block (CTB) may be an N x N block of samples for some value of N, such that the division of a component into CTBs is a partition. A coding unit (CU) may be or include a coding block of luma samples for a picture having three sample arrays, two corresponding coding blocks of chroma samples, or a coding block of samples for a picture coded using three distinct color planes and syntax structures used to code a monochrome picture or sample. Correspondingly, a coding block (CB) may be an M x N block of samples for some value of M and N, such that the division of a CTB into coding blocks is a partition.

[0104] In an embodiment, for example, according to HEVC, coding tree units (CTUs) may be divided into CUs by using a quadtree structure referred to as a coding tree. The decision of whether to code a picture region using inter-picture (temporal) or intra-picture (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four PUs according to a PU partition type. Within one PU, the same prediction process is applied, and related information is transmitted to the decoder for each PU. After obtaining a residual block by applying a prediction process based on the PU partition type, the CU can be divided into transform units (TUs) according to another quadtree structure similar to the coding tree for the CU.

[0105] In an embodiment, according to the latest video coding standard currently under development, for example, referred to as Versatile Video Coding (VVC), a combined quadtree and binary tree (QTBT) partitioning is used to partition a coding block. In the QTBT block structure, a CU can have either a square or rectangular shape. For example, a coding tree unit (CTU) is first partitioned using a quadtree structure. The quadtree leaf nodes are further partitioned using a binary tree or ternary tree (or triangular) structure. The partitioned tree leaf nodes are called coding units (CUs), and their segmentation is used for prediction and transform processing without further partitioning. This means that in the QTBT coding block structure, CUs, PUs, and TUs have the same block size. In parallel, multiple partitions, for example, ternary tree partitioning, may also be used with the QTBT block structure.

[0106] In one example, mode select unit 260 of video encoder 20 may be configured to perform any combination of the partitioning techniques described herein.

[0107] As described above, video encoder 20 is configured to determine or select a best or optimal prediction mode from a (e.g., predetermined) set of prediction modes, which may include, for example, intra-prediction modes and / or inter-prediction modes.

[0108] Intra prediction The set of intra prediction modes may include, for example, 35 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, as defined in HEVC, or may include, for example, 67 different intra prediction modes, e.g., non-directional modes such as DC (or average) mode and planar mode, or directional modes, as defined in VVC.

[0109] The intra prediction unit 254 is configured to use reconstructed samples of neighboring blocks of the same current picture to generate an intra prediction block 265 according to an intra prediction mode from a set of intra prediction modes.

[0110] Intra prediction unit 254 (or, generally, mode select unit 260) is further configured to output the intra prediction parameters (or, generally, information indicating the selected intra prediction mode for the block) in the form of syntax element 266 to entropy encoding unit 270 for inclusion in encoded picture data 21, so that, for example, video decoder 30 can receive the prediction parameters and use them for decoding.

[0111] Inter Prediction The set (or possible) inter-prediction modes depend on available reference pictures (i.e., previous, at least partially decoded pictures, e.g., stored in DBP 230) and other inter-prediction parameters, such as whether the entire reference picture is used to find the best-matching reference block or only a portion of the reference picture, e.g., a search window area around that region of the current block, is used, and / or whether pixel interpolation, e.g., half / semi-pixel and / or quarter-pixel interpolation, is applied.

[0112] In addition to the above prediction modes, skip mode and / or direct mode may also be applied.

[0113] The inter prediction unit 244 may include a motion estimation (ME) unit and a motion compensation (MC) unit (neither of which are shown in FIG. 2). The motion estimation unit may be configured to receive or obtain a picture block 203 (such as the current picture block 203 of the current picture 17) and a decoded picture 231, or at least one or more previously reconstructed blocks, e.g., reconstructed blocks of one or more other / different previously decoded pictures 231, for motion estimation. For example, a video sequence may include the current picture and the previously decoded picture 231, or in other words, the current picture and the previously decoded picture 231 may be part of or form a sequence of pictures that form a video sequence.

[0114] The encoder 20 may be configured to, for example, select a reference block from multiple reference blocks of the same or different ones of multiple other pictures, and provide the reference picture (or reference picture index) and / or an offset (spatial offset) between the position (x, y coordinates) of the reference block and the position of the current block as an inter-prediction parameter to the motion estimation unit, which offset is also called a motion vector (MV).

[0115] The motion compensation unit is configured to, for example, obtain, e.g., receive, inter prediction parameters and perform inter prediction based on or using the inter prediction parameters to obtain an inter prediction block 265. The motion compensation performed by the motion compensation unit may include performing interpolation, possibly to sub-pixel accuracy, to fetch or generate a prediction block based on motion / block vectors determined by motion estimation. Interpolation filtering may generate additional pixel samples from known pixel samples, thus potentially increasing the number of candidate prediction blocks that can be used to code the picture block. Upon receiving a motion vector for the PU of the current picture block, the motion compensation unit may locate the prediction block to which the motion vector points in one of the reference picture lists.

[0116] The motion compensation unit may also generate syntax elements associated with the blocks and video slices for use by video decoder 30 in decoding picture blocks of the video slices. In addition to or instead of slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be generated or used.

[0117] Entropy Coding The entropy encoding unit 270 is configured to apply, for example, an entropy encoding algorithm or scheme (e.g., a variable length coding (VLC) scheme, a context-adaptive VLC scheme (CAVLC), an arithmetic coding scheme, binarization, context-adaptive binary arithmetic coding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or other entropy encoding method or technique) or bypass (uncompressed) to the quantized coefficients 209, inter-prediction parameters, intra-prediction parameters, loop filter parameters, and / or other syntax elements to obtain encoded picture data 21. The encoded picture data 21, for example, in the form of an encoded bitstream 21, can be output via output 272, so that, for example, video decoder 30 can receive and use those parameters for decoding. The encoded bitstream 21 can be transmitted to video decoder 30 or stored in memory for later transmission or retrieval by video decoder 30.

[0118] Other structural variations of the video encoder 20 can be used to encode the video stream. For example, a non-transform-based encoder 20 can quantize the residual signal directly for certain blocks or frames without the transform processing unit 206. In another implementation, the encoder 20 can have the quantization unit 208 and the inverse quantization unit 210 combined into a single unit.

[0119] Decoder and decoding method 3 shows an example of a video decoder 30 configured to implement the techniques of the present application. The video decoder 30 is configured to receive encoded picture data 21 (e.g., encoded bitstream 21), for example, encoded by encoder 20, to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, for example, data representing picture blocks and associated syntax elements of an encoded video slice (and / or tile group or tile).

[0120] 3, decoder 30 includes an entropy decoding unit 304, an inverse quantization unit 310, an inverse transform processing unit 312, a reconstruction unit 314 (e.g., summer 314), a loop filter 320, a decoded picture buffer (DBP) 330, a mode application unit 360, an inter prediction unit 344, and an intra prediction unit 354. Inter prediction unit 344 may be or may include a motion compensation unit. Video decoder 30, in some examples, may perform a decoding path that is generally the reverse of the encoding path described with respect to video encoder 100 from FIG. 2.

[0121] As described with respect to encoder 20, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, loop filter 220, decoded picture buffer (DPB) 230, inter prediction unit 344, and intra prediction unit 354 are also referred to as forming an “embedded decoder” of video encoder 20. Thus, inverse quantization unit 310 may be functionally identical to inverse quantization unit 110, inverse transform processing unit 312 may be functionally identical to inverse transform processing unit 212, reconstruction unit 314 may be functionally identical to reconstruction unit 214, loop filter 320 may be functionally identical to loop filter 220, and decoded picture buffer 330 may be functionally identical to decoded picture buffer 230. Accordingly, the descriptions provided for the respective units and functions of video encoder 20 correspondingly apply to the respective units and functions of video decoder 30.

[0122] Entropy Decoding The entropy decoding unit 304 is configured to parse the bitstream 21 (or generally, the encoded picture data 21) and, e.g., perform entropy decoding on the encoded picture data 21 to obtain, e.g., quantized coefficients 309 and / or decoded coding parameters (not shown in FIG. 3 ), such as inter-prediction parameters (e.g., reference picture indices and motion vectors), intra-prediction parameters (e.g., intra-prediction modes or indices), transform parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to the encoding schemes described with respect to the entropy encoding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit 360 and other parameters to other units of the decoder 30. Video decoder 30 may receive syntax elements at the video slice level and / or the video block level. In addition to or as an alternative to slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used.

[0123] inverse quantization Inverse quantization unit 310 may be configured to receive a quantization parameter (QP) (or generally, information regarding inverse quantization) and quantized coefficients from encoded picture data 21 (e.g., by parsing and / or decoding by entropy decoding unit 304), and apply inverse quantization to decoded quantized coefficients 309 based on the quantization parameter to obtain dequantized coefficients 311. Dequantized coefficients 311 are sometimes referred to as transform coefficients 311. The inverse quantization process may use the quantization parameter determined by video encoder 20 for each video block in a video slice (or tile or tile group) to determine the degree of quantization, and similarly, the degree of dequantization to be applied.

[0124] Inverse transformation The inverse transform processing unit 312 may be configured to receive the dequantized coefficients 311, also referred to as transform coefficients 311, and apply a transform to the dequantized coefficients 311 to obtain reconstructed residual blocks 213 in the sample domain. The reconstructed residual blocks 213 may also be referred to as transform blocks 313. The transform may be an inverse transform, e.g., an inverse DCT, an inverse DST, an inverse integer transform, or a conceptually similar inverse transform process. The inverse transform processing unit 312 may further be configured to receive transform parameters or corresponding information from the encoded picture data 21 (e.g., by parsing and / or decoding by the entropy decoding unit 304) to determine the transform to apply to the dequantized coefficients 311.

[0125] Reconstruction The reconstruction unit 314 (e.g., an adder or summer 314) may be configured to add the reconstructed residual block 313 to the prediction block 365 to obtain a reconstructed block 315 in the sample domain, e.g., by adding sample values ​​of the reconstructed residual block 313 and the prediction block 365.

[0126] filtering Loop filter unit 320 (in the coding loop or after the coding loop) is configured to filter reconstructed block 315 to obtain filtered block 321, e.g., to smooth pixel transitions or otherwise improve video quality. Loop filter unit 320 may include one or more loop filters, such as a deblocking filter, a sample adaptive offset (SAO) filter, or one or more other filters, e.g., a bilateral filter, an adaptive loop filter (ALF), a sharpening filter, a smoothing filter, or a collaborative filter, or any combination thereof. Although loop filter unit 320 is shown in FIG. 3 as an in-loop filter, in other configurations, loop filter unit 320 may be implemented as a post-loop filter.

[0127] Decoded Picture Buffer The decoded video blocks 321 of the picture are then stored in a decoded picture buffer 330. The decoded picture buffer 330 stores the decoded picture 331 as a reference picture for subsequent motion compensation for other pictures and / or for respective output or display.

[0128] Decoder 30 is configured to output decoded pictures 311 for presentation or viewing to a user, for example via output 312.

[0129] prediction The inter prediction unit 344 may be identical to the inter prediction unit 244 (in particular, the motion compensation unit), and the intra prediction unit 354 may be identical in function to the inter prediction unit 254, and performs the division or partitioning decision and prediction based on the division and / or prediction parameters or respective information received from the encoded picture data 21 (e.g., by parsing and / or decoding, e.g., by the entropy decoding unit 304). The mode application unit 360 may be configured to perform block-wise prediction (intra prediction or inter prediction) based on the reconstructed picture, block, or respective sample (filtered or unfiltered) to obtain a prediction block 365.

[0130] When a video slice is coded as an intra-coded (I) slice, intra prediction unit 354 of mode application unit 360 is configured to generate a prediction block 365 for a picture block of the current video slice based on a signaled intra prediction mode and data from previously decoded blocks of the current picture. When a video picture is coded as an inter-coded (i.e., B or P) slice, inter prediction unit 344 (e.g., a motion compensation unit) of mode application unit 360 is configured to generate a prediction block 365 for a video block of the current video slice based on motion vectors and other syntax elements received from entropy decode unit 304. For inter prediction, the prediction block may be generated from one of the reference pictures in one of the reference picture lists. Video decoder 30 may construct the reference frame lists, List 0 and List 1, using a default construction technique based on the reference pictures stored in DPB 330. The same or similar may apply for or by embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or as an alternative to slices (e.g., video slices). For example, video may be coded using I, P, or B tile groups and / or tiles.

[0131] Mode application unit 360 is configured to determine prediction information for video blocks of the current video slice by parsing motion vectors or related information and other syntax elements, and use the prediction information to generate a prediction block for the current video block to be decoded. For example, mode application unit 360 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) to use for coding the video blocks of the video slice, the inter-prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more of the reference picture lists for the slice, motion vectors for each inter-encoded video block of the slice, inter-prediction status for each inter-coded video block of the slice, and other information for decoding video blocks in the current video slice. The same or similar may apply for or by embodiments that use tile groups (e.g., video tile groups) and / or tiles (e.g., video tiles) in addition to or instead of slices (e.g., video slices). For example, video may be coded using I, P, or B tile groups and / or tiles.

[0132] 3 may be configured to divide and / or decode pictures using slices (also referred to as video slices), where a picture can be divided into or decoded using one or more (typically non-overlapping) slices, each of which may include one or more blocks (e.g., CTUs).

[0133] 3 may be configured to divide and / or decode a picture using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), where a picture can be divided into or decoded using one or more (typically non-overlapping) tile groups, each of which may be, for example, rectangular in shape and contain one or more blocks (e.g., CTUs), e.g., full or partial blocks.

[0134] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 may generate an output video stream without the loop filtering unit 320. For example, a non-transform-based decoder 30 may inverse quantize the residual signal directly for certain blocks or frames without the inverse transform processing unit 312. In another embodiment, the video decoder 30 may have the inverse quantization unit 310 and the inverse transform processing unit 312 combined into a single unit.

[0135] It should be understood that in the encoder 20 and the decoder 30, the processing result of the current step may be further processed and then output to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the processing result of the interpolation filtering, motion vector derivation, or loop filtering.

[0136] It should be noted that further operations may be applied to the derived motion vectors of the current block (including, but not limited to, control point motion vectors in affine mode, sub-block motion vectors in affine, planar, and ATMVP modes, temporal motion vectors, etc.). For example, the value of a motion vector is constrained to a predefined range according to its representation bits. If the representation bits of a motion vector are bitDepth, the range is -2^(bitDepth-1) to 2^(bitDepth-1)-1. For example, if bitDepth is set equal to 16, the range is -32768 to 32767, and if bitDepth is set equal to 18, the range is -131072 to 131071. For example, the values ​​of derived motion vectors (e.g., MVs of four 4x4 sub-blocks in one 8x8 block) are constrained such that the maximum difference between the integer parts of the four 4x4 sub-block MVs is less than or equal to N pixels, e.g., less than or equal to 1 pixel. Here we provide two methods for constraining motion vectors according to bitDepth.

[0137] Method 1: Remove the overflow MSB (Most Significant Bit) with a flow operation

number

number

[0138] Method 2: Remove the overflow MSB by clipping the value.

number

number

[0139] 4 is a schematic diagram of a video coding device 400 according to an embodiment of the present disclosure. The video coding device 400 is suitable for implementing the disclosed embodiments described herein. In an embodiment, the video coding device 400 may be a decoder, such as the video decoder 30 of FIG. 1A, or an encoder, such as the video encoder 20 of FIG. 1A.

[0140] Video coding device 400 includes an ingress port 410 (or input port 410) and a receiver unit (Rx) 420 for receiving data; a processor, logic unit, or central processing unit (CPU) 430 for processing data; a transmitter unit (Tx) 440 and an egress port 450 (or output port 450) for transmitting data; and a memory 460 for storing data. Video coding device 400 may also have optical-to-electrical (OE) and electrical-to-optical (EO) components coupled to ingress port 410, receiver unit 420, transmitter unit 440, and egress port 450 for inputting and outputting optical or electrical signals.

[0141] The processor 430 is implemented by hardware and software. The processor 430 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGA, ASIC, and DSP. The processor 430 communicates with the ingress port 410, the receiver unit 420, the transmitter unit 440, the egress port 450, and the memory 460. The processor 430 includes a coding module 470. The coding module 470 implements the above-disclosed embodiments. For example, the coding module 470 implements, processes, prepares, or provides various coding operations. Thus, the inclusion of the coding module 470 substantially improves the functionality of the video coding device 400 and enables transformation of the video coding device 400 into different states. Alternatively, the coding module 470 is implemented as instructions stored in the memory 460 and executed by the processor 430.

[0142] Memory 460 may include one or more disks, tape drives, and solid-state drives, may be used as overflow data storage, may store programs when the programs are selected for execution, and may store instructions and data read during program execution. Memory 460 may be, for example, volatile and / or nonvolatile, and may be read-only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).

[0143] FIG. 5 is a simplified block diagram of a device 500 that may be used as either or both of source device 12 and destination device 14 from FIG. 1, according to an example embodiment.

[0144] Processor 502 in device 500 may be a central processing unit. Alternatively, processor 502 may be any other type of device or devices now existing or later developed that are capable of manipulating or processing information. While the disclosed implementations can be practiced using a single processor as shown, such as processor 502, advantages in speed and efficiency can be achieved using multiple processors.

[0145] The memory 504 in the device 500 may, in some implementations, be a read-only memory (ROM) device or a random-access memory (RAM) device. Any other suitable type of storage device may be used as the memory 504. The memory 504 may include code and data 506 that is accessed by the processor 502 using a bus 512. The memory 504 may further include an operating system 508 and application programs 510, which include at least one program that allows the processor 502 to perform the methods described herein. For example, the application programs 510 may include applications 1-N, which further include a video coding application that performs the methods described herein.

[0146] The device 500 may also include one or more output devices, such as a display 518. The display 518, in one example, may be a touch-sensitive display that combines a display with touch-sensitive elements operable to sense touch input. The display 518 may be coupled to the processor 502 via the bus 512.

[0147] Although shown here as a single bus, bus 512 of device 500 may be comprised of multiple buses. Additionally, secondary storage 514 may be directly coupled to other components of device 500 or may be accessed over a network, and may include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 500 may be implemented in a wide variety of configurations.

[0148] The 14th JVET Conference held in Geneva adopted the contribution JVET-N0217: Affine Linear Weighted Intra Prediction (ALWIP).

[0149] ALWIP introduces three new sets of intra modes, which are: 35 modes for 4x4 blocks. 19 modes for 8x4, 4x8, 8x8 blocks. 11 modes for other cases where width and height are both less than or equal to 64 samples.

[0150] Correspondingly, in ALWIP, the variable for block size type (sizeId) is defined as follows: If the block size is 4x4, the block size type sizeId is 0. Otherwise, if the block size is 8x4, 4x8, or 8x8, the block size type sizeId is 1. Otherwise, if the block size is not one of the above and the block width and height are both less than 64, the block size type sizeId is 2.

[0151] These modes generate the luma intra prediction signal by matrix-vector multiplication and offset addition from reference samples one line to the left and above the current block. For this reason, affine linear weighted intra prediction is also called matrix-based intra prediction (MIP). For the following text, the terms MIP and ALWIP are used interchangeably, and both describe the tools in JVET-N0217.

[0152] To predict samples for a rectangular block of width W and height H, Affine Linear Weighted Intra Prediction (ALWIP) takes as input one line of H reconstructed neighboring boundary samples to the left of the block and one line of W reconstructed neighboring boundary samples above the block. If reconstructed samples are not available, they are generated in the same way as regular intra prediction. The generation of the predicted signal is based on three steps: 1. Of the boundary samples, 4 samples are extracted by averaging if W=H=4, and 8 samples in all other cases. 2. Using the averaged samples as input, a matrix-vector multiplication followed by the addition of an offset is performed. The result is a reduced prediction signal on a subsampled set of samples in the original block. 3. Predictions at the remaining positions are generated from the predictions on the subsampled set by linear interpolation, which is a single-step linear interpolation in each direction.

[0153] The whole process of averaging, matrix-vector multiplication, and linear interpolation is illustrated for different shapes in Figures 6 to 9. Note that the remaining shapes are treated the same as the one in the illustrated case.

[0154] Figure 6 illustrates the process for a 4x4 block. Given a 4x4 block, ALWIP takes two averages along each axis of the boundary. The resulting four input samples enter a matrix-vector multiplication. The matrix is ​​taken from set S0. After adding an offset, this gives 16 final predicted samples. No linear interpolation is required to generate the predicted signal. Thus, a total of (4 16) / (4 4) = 4 multiplications are performed per sample.

[0155] Figure 7 illustrates this process for an 8x8 block. Given an 8x8 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is ​​taken from set S1. This gives 16 samples on the odd positions of the prediction block. Thus, a total of (8 16) / (8 8) = 2 multiplications per sample are performed. After adding the offset, these samples are vertically interpolated using the reduced top boundary. This is followed by horizontal interpolation using the original left boundary. Thus, a total of two multiplications are required per sample to compute the ALWIP prediction.

[0156] Figure 8 illustrates this process for an 8x4 block. Given an 8x4 block, ALWIP takes four averages along the horizontal axis of the boundary and the four original boundary values ​​on the left boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is ​​taken from set S1. This gives 16 samples on each odd horizontal position of the prediction block, at each vertical position. Thus, a total of (8 16) / (8 4) = 4 multiplications per sample are performed. After adding the offset, these samples are horizontally interpolated using the original left boundary. Thus, a total of four multiplications per sample are required to compute the ALWIP prediction.

[0157] If transposed, it is treated accordingly.

[0158] Figure 9 illustrates this process for a 16x16 block. Given a 16x16 block, ALWIP takes four averages along each axis of the boundary. The resulting eight input samples enter a matrix-vector multiplication. The matrix is ​​taken from set S2. This gives 64 samples on the odd positions of the prediction block. Thus, a total of (8 64) / (16 16) = 2 multiplications per sample are performed. After adding the offset, these samples are vertically interpolated using the eight averages from the top boundary. This is followed by horizontal interpolation using the original left boundary. Thus, a total of two multiplications are required per sample to compute the ALWIP prediction.

[0159] For larger shapes, the procedure is essentially the same, and it is easy to check that the number of multiplications per sample is less than four.

[0160] For W × 8 blocks where W > 8, samples are provided at odd horizontal positions and every vertical position, so only horizontal interpolation is required. In this case, (8 64) / (W 8) = 64 / W multiplications are performed per sample to compute the reduced prediction. For W > 16, the additional multiplications per sample required for linear interpolation are less than two. Thus, the total number of multiplications per sample is four or less.

[0161] Finally, for W×4 blocks where W>8, A k Let W be the matrix obtained by removing each row corresponding to the odd-numbered entries along the horizontal axis of the downsampled block. Thus, the output size is 32, and again, only horizontal interpolation remains to be performed. To compute the reduced prediction, (8 32) / (W 4) = 64 / W multiplications are performed per sample. For linear interpolation, no additional multiplications are needed for W=16, but for W>16, fewer than two multiplications per sample are needed. Thus, the total number of multiplications is less than or equal to four.

[0162] If transposed, it is treated accordingly.

[0163] In the JVET-N0217 contribution, the Most Probable Mode (MPM) list approach is also applied for MIP intra-mode coding. Currently there are two MPM lists used for blocks: 1. If the current block uses normal intra mode (i.e., not MIP intra mode), the 6-MPM list is used. 2. If the current block uses MIP intra mode, the 3-MPM list is used.

[0164] Both of the above two MPM lists are constructed based on the intra prediction modes of neighboring blocks, so the following cases are possible: 1. The current block is normally intra predicted, while one or more of its neighboring blocks is MIP intra predicted, or 2. The current block is subjected to MIP intra prediction, while one or more of its neighboring blocks is subjected to normal intra prediction.

[0165] Under such circumstances, the neighboring intra-prediction modes are derived indirectly using a look-up table.

[0166] In one example, if the current block is normally intra predicted, but the block above it (A) as shown in FIG. 13 is applied with MIP intra prediction, The following lookup table 1 is used: The normal intra prediction mode is derived based on the block size and type of the block above and the MIP intra prediction mode of the block above. Similarly, when the left (L) block as shown in Figure 13 is to be subjected to MIP intra prediction, the normal intra prediction mode is derived based on the block size and type of the left block and the MIP intra prediction mode of the left block. [Table 6]

[0167] In one example, when a current block is subjected to MIP intra prediction and the block above it (A) as shown in Figure 14 is predicted using normal intra prediction mode, the following lookup table 2 is used. The MIP intra prediction mode is derived based on the block size type and normal intra prediction mode of the block above. Similarly, when a left (L) block as shown in Figure 14 is subjected to normal intra prediction based on the block size type and normal intra prediction mode of the left block, the MIP intra prediction mode is derived. [Table 7]

[0168] In JEM, a secondary transform is applied between the forward primary transform and quantization (at the encoder), and between dequantization and the inverse primary transform (at the decoder). As shown in Figure 10, the 4x4 (or 8x8) secondary transform that is performed depends on the block size. For example, for small blocks (i.e., min(width, height)<8), a 4x4 secondary transform is applied, and for blocks larger than per 8x8 block (i.e., min(width, height)>4), an 8x8 secondary transform is applied.

[0169] The application of the non-separable transformation is explained below using an example input. To apply a non-separable transform, we use a 4×4 input block X

number

number

number

[0170] Indivisible transformations are

number

number

number

[0171] In VVC 5.0, as a new coding tool with the following features, the reduced secondary transform (RST) according to proposal JVET-N0193 was adopted.

[0172] The main idea of the reduced transform (RT) is to map an N - dimensional vector to an R - dimensional vector in a different space. Here, R / N (R < N) is the reduction factor. The RT matrix is an R×N matrix as follows:

Equation

[0173] A reduction factor of 4 (1 / 4 size) RST8x8 is applied. Thus, instead of the traditional 8x8 non-separable transform matrix size of 64x64, a 16x64 direct matrix is ​​used. In other words, a 64x16 inverse RST matrix is ​​used at the decoder side to generate the core (primary) transform coefficients in the 8x8 top-left regions. Forward RST8x8 uses a 16x64 (or 8x64 for 8x8 blocks) matrix to generate non-zero coefficients only in the top-left 4x4 region within a given 8x8 region. In other words, when RST is applied, the 8x8 region except for the top-left 4x4 region has only zero coefficients. For RST4x4, a 16x16 (or 8x16 for 4x4 blocks) direct matrix multiplication is applied.

[0174] Inverse RST is conditionally applied when the following two conditions are met: Block size is greater than or equal to a given threshold (W>=4 && H>=4) ·Conversion skip mode·flag equals 0 If both the width (W) and height (H) of a transform coefficient block are greater than 4, RST8x8 is applied to an 8x8 region in the top-left corner of the transform coefficient block. Otherwise, RST4x4 is applied to a min(8,W) x min(8,H) region in the top-left corner of the transform coefficient block.

[0175] If the RST index is equal to 0, then the RST is not applied. Otherwise, the RST is applied and the kernel is chosen using the RST index. Furthermore, RST is applied for both luma and chroma for intra CUs in both intra and inter slices. When dual tree is enabled, RST indices for luma and chroma are signaled separately. For inter slices (dual tree disabled), a single RST index is signaled and used for both luma and chroma.

[0176] Intra Sub-Partitions (ISP) is an intra prediction mode in VVC4.0. When ISP mode is selected, RST is disabled and RST index is not signaled. Even if RST is applied to all feasible partition blocks, the performance improvement is small. Furthermore, disabling RST for ISP-predicted residuals may reduce encoding complexity.

[0177] The RST matrix is ​​selected from four transform sets, each containing two transforms. The transform set to be applied is determined from the intra prediction mode as follows:

[0178] If one of the three CCLM (Cross-component linear model; in this mode, the chroma components are predicted from the luma component) modes is indicated, then transform set 0 is selected. Otherwise, transform set selection is performed according to the following table: [Table 8] The index IntraPredMode used to access Table 3 has a range of [-14, 83]. This is the transform mode index used for wide-angle intra prediction.

[0179] An example of a set of transformations is shown below.

number

number

number

number

number

number

number

number

number

number

number

number

[0180] In the case of RST8x8 or RST4x4, there may arise a time when all TUs have the size of 4x4 TUs or 8x8 TUs in terms of the number of multiplications, so the top 8x64 and 8x16 matrices (i.e., the first 8 transformation basis vectors from the top in each matrix) are applied to the 8x8 and 4x4 TUs, respectively.

[0181] For blocks larger than 8x8 TUs (both width and height are greater than 8), RST8x8 (i.e., a 16x64 matrix) is applied to the top-left 8x8 region. For 8x4 TUs or 4x8 TUs, RST4x4 (i.e., a 16x16 matrix) is applied to the top-left 4x4 region (in one example, RST4x4 is not applied to other 4x4 regions). For 4xN or Nx4 TUs (N ≥ 16), RST4x4 is applied to the two adjacent top-left 4x4 blocks.

[0182] Due to the simplifications mentioned above, in some cases there are eight multiplications per sample.

[0183] To reduce the secondary transformation matrix size, 16x48 matrices are applied in the same transformation set configuration, as shown in Figure 12, where each 16x48 matrix takes 48 input samples from three 4x4 blocks in the top left 8x8 block (in one example, the bottom right 4x4 blocks are excluded).

[0184] VVC5.0 discloses the MIP and RST tools, both of which are applicable to intra blocks. However, the two tools are not harmonized in terms of secondary transform core selection. In other words, when MIP is not a normal intra mode, that is, when an intra block is predicted using the MIP mode, the RST transform core selection method is not defined in both the adopted proposals JVET-N0217 and JVET-N0193. The following solution solves the above problem:

[0185] Figure 15 illustrates a method according to the present disclosure. Figure 15 illustrates a coding method implemented by a decoding device or an encoding device. The decoding device may be the decoder 30 as described above. Similarly, the encoding device may be the encoder 20 as described above. In Figure 15, in step 1601, the method includes determining an intra-prediction mode for the current block. In a next step 1603, the method includes determining a secondary transform selection for the current block based on the intra-prediction mode determined for the current block. This will be described in further detail below.

[0186] Figure 16 illustrates an encoder 20 according to the present disclosure. In Figure 16, the encoder 20 includes a determination unit 2001 configured to determine an intra-prediction mode for a current block. The encoder 20 of Figure 17 further illustrates a selection unit 2003 configured to determine a selection of a secondary transform for the current block based on the intra-prediction mode determined for the current block.

[0187] Figure 17 illustrates a decoder 30 according to the present disclosure. In Figure 17, the decoder 30 includes a determination unit 3001 configured to determine an intra-prediction mode for a current block. The decoder 30 of Figure 17 further illustrates a selection unit 3003 configured to determine a selection of a secondary transform for the current block based on the intra-prediction mode determined for the current block.

[0188] In the following, the coding method of FIG. 15 implemented by the decoding device or encoding device of FIG. 17 and FIG. 16, respectively, will be described in further detail.

[0189] Solution 1. According to solution 1, matrix-based intra prediction and contractive quadratic transform are excluded for the same intra prediction block.

[0190] In one example, if an intra-predicted block is predicted using MIP mode (in one example, the value of the MIP flag may be used to indicate whether the block is predicted using MIP mode), the secondary transform is disabled for this intra-predicted block. In other words, the value of the secondary transform index is set to 0, or the secondary transform index does not need to be decoded from the bitstream.

[0191] If the intra-predicted block is not predicted using MIP mode, the transform core for the secondary transform is selected based on the method described in JVET-N0193. Specification text changes compared to JVET-N0193 are highlighted in gray. [Table 9] 7.4.3.1 Semantics of the sequence parameter set RBSP …… sps_st_enabled_flag equal to 1 specifies that st_idx may be present in the residual coding syntax for the intra-coded unit. sps_st_enabled_flag equal to 0 specifies that st_idx is not present in the residual coding syntax for the intra-coded unit. …… 7.4.7.5 Semantics of Coding Units …… st_idx[x0][y0] specifies which secondary transformation kernel is applied between two candidate kernels in the selected transformation set. st_idx[x0][y0] equal to 0 specifies that no secondary transformation is applied. The array indices x0, y0 specify the position (x0, y0) of the top-left sample of the considered transformation block relative to the top-left sample of the picture. intra_mip_flag[x0][y0] equal to 1 specifies that the intra prediction type for luma samples is affine linear weighted intra prediction. intra_lwip_flag[x0][y0] equal to 0 specifies that the intra prediction type for luma samples is not affine linear weighted intra prediction.

[0192] Solution 2 According to solution 2, during the process of transform core selection for secondary transforms, if a block is predicted using MIP mode, one of the secondary transform core set is considered to be used for this block.

[0193] In one embodiment, When the current block is predicted in MIP mode, transform set 0 is used as the selected secondary transform core set.

[0194] The RST matrix is ​​chosen from four transform sets, each containing two transforms. The transform set applied to a block is determined according to the intra prediction mode, as follows: If the current intra block is predicted using CCLM mode, transformation set 0 is selected; Otherwise, if the current intra block is predicted using MIP mode, transform set 0 is selected; Otherwise (if the current intra block is not predicted using CCLM or MIP mode), the selection of the transformation set is done according to the following table: [Table 10] The index range of IntraPredMode is between -14 and 83 (inclusive), which is the transformed mode index used for wide-angle intra prediction.

[0195] In this solution, one transform set is used when the current block is predicted using MIP mode, in one example transform set 0 is used, but other transform sets can also be used in this solution.

[0196] Solution 3 According to Solution 3, during the process of transform core selection for secondary transform, if a block is predicted using MIP mode, the trained secondary transform core set is considered to be used for this block. The trained secondary transform core set may be different from the transform core set in the above example.

[0197] In one embodiment: When the current block is predicted in MIP mode, transform set 4 (the newly trained one) is used as the selected secondary transform core set. Transformation set 4 has the same dimensions (i.e., 16 x 16 and 16 x 48) as transformation sets 0-3, newly trained based on the same machine learning method and input training set, specifically for the MIP mode.

[0198] The RST matrix is ​​chosen from four transformation sets, each containing two transformations. The set of transforms applied to a block is determined according to the intra prediction mode as follows: If the current intra block is predicted using CCLM mode, transformation set 0 is selected; Otherwise, if the current intra block is predicted using MIP mode, transform set 4 is selected; Otherwise (if the current intra block is not predicted using CCLM or MIP mode), the selection of the transformation set is done according to the following table: [Table 11] The index range of IntraPredMode is between -14 and 83 (inclusive), which is the transformed mode index used for wide-angle intra prediction.

[0199] In this solution, if the current block is predicted using MIP mode, a new trained transformation set (eg, transformation set 4) is used.

[0200] Solution 4 According to Solution 4, during the process of transform core selection for secondary transform, if a block is predicted using MIP mode, a lookup table is used to map the MIP mode index to a normal intra mode index, and then a set of secondary transform cores is selected based on this normal intra mode index.

[0201] In one embodiment: If the current block is predicted using MIP mode, the MIP mode index is mapped to a regular intra mode index based on Table 6. In this example, Table 6 is the same as the MIP MPM lookup table. [Table 12]

[0202] Then, the selection of the secondary transformation set is performed according to the following table: [Table 13]

[0203] For example, if the current block is predicted using MIP mode index 10 and the value of the block size type sizeID for the current block is 0, the mapped normal intra mode index is 18 based on Table 6, and secondary transform set 2 is selected based on Table 7.

[0204] In this solution, if a block is predicted using MIP mode, a mapping method from MIP mode indexes to regular intra-mode indexes is used, and the secondary transform core selection is based on the mapped regular intra-mode indexes.

[0205] 18 is a block diagram showing a content delivery system 3100 for implementing a content delivery service. The content delivery system 3100 includes a capture device 3102, a terminal device 3106, and optionally a display 3126. The capture device 3102 communicates with the terminal device 3106 through a communication link 3104. The communication link may include the communication channel 13 described above. The communication link 3104 may include, but is not limited to, WIFI, Ethernet, cable, wireless (3G / 4G / 5G), USB, or any combination thereof.

[0206] The capture device 3102 may generate data and encode it using the encoding method described in the above embodiments. Alternatively, the capture device 3102 may deliver the data to a streaming server (not shown), which encodes the data and transmits the encoded data to the terminal device 3106. The capture device 3102 may include, but is not limited to, a camera, a smartphone or pad, a computer or laptop, a video conferencing system, a PDA, an in-vehicle device, or any combination thereof. For example, the capture device 3102 may include the source device 12 described above. If the data includes video, a video encoder 20 included in the capture device 3102 may actually perform the video encoding process. If the data includes audio (i.e., voice), an audio encoder included in the capture device 3102 may actually perform the audio encoding process. For some practical scenarios, the capture device 3102 delivers the encoded video and audio data by multiplexing them together. In other practical scenarios, for example in a videoconferencing system, the encoded audio data and the encoded video data are not multiplexed: the capture device 3102 delivers the encoded audio data and the encoded video data separately to the terminal device 3106.

[0207] In the content delivery system 3100, the terminal device 3106 receives and plays the encoded data. The terminal device 3106 may be a device capable of receiving and restoring data, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a set-top box (STB) 3116, a video conferencing system 3118, a video surveillance system 3120, a personal digital assistant (PDA) 3122, a vehicle-mounted device 3124, or any combination thereof, capable of decoding the encoded data described above. For example, the terminal device 3106 may include the destination device 14 described above. If the encoded data includes video, the video decoder 30 included in the terminal device is prioritized to perform video decoding. If the encoded data includes audio, the audio decoder included in the terminal device is prioritized to perform audio decoding.

[0208] For terminal devices with a display, such as a smartphone or pad 3108, a computer or laptop 3110, a network video recorder (NVR) / digital video recorder (DVR) 3112, a TV 3114, a personal digital assistant (PDA) 3122, or a vehicle mounted device 3124, the terminal device can provide the decoded data to its display. For terminal devices without a display, such as an STB 3116, a video conferencing system 3118, or a video surveillance system 3120, an external display 3126 is contacted thereto to receive and display the decoded data.

[0209] When each device in this system performs encoding or decoding, the picture encoding device or picture decoding device shown in the above-mentioned embodiment can be used.

[0210] 19 is a diagram illustrating the structure of an example of the terminal device 3106. After the terminal device 3106 receives the stream from the capture device 3102, a protocol progression unit 3202 analyzes the transmission protocol of the stream. This protocol includes, but is not limited to, Real Time Streaming Protocol (RTSP), Hypertext Transfer Protocol (HTTP), HTTP Live Streaming Protocol (HLS), MPEG-DASH, Real Time Transport Protocol (RTP), Real Time Messaging Protocol (RTMP), or any kind of combination thereof. After the protocol progression unit 3202 processes the stream, a stream file is generated. The file is output to the demultiplexing unit 3204. The demultiplexing unit 3204 can separate the multiplexed data into encoded audio data and encoded video data. As mentioned above, in some practical scenarios, for example, in a video conferencing system, the encoded audio data and the encoded video data are not multiplexed. In this situation, the encoded data is sent to the video decoder 3206 and the audio decoder 3208 without passing through the demultiplexing unit 3204.

[0211] Through the demultiplexing process, a video elementary stream (ES), an audio ES, and optionally subtitles are generated. A video decoder 3206, which includes the video decoder 30 as described in the above embodiments, decodes the video ES using the decoding method as shown in the above embodiments to generate video frames and provides this data to a synchronization unit 3212. An audio decoder 3208 decodes the audio ES to generate audio frames and provides this data to a synchronization unit 3212. Alternatively, the video frames may be stored in a buffer (not shown in FIG. 19) before being provided to the synchronization unit 3212. Similarly, the audio frames may be stored in a buffer (not shown in FIG. 19) before being provided to the synchronization unit 3212.

[0212] The synchronization unit 3212 synchronizes the video and audio frames and provides the video / audio to a video / audio display 3214. For example, the synchronization unit 3212 synchronizes the presentation of video and audio information. Information may be coded in the syntax using timestamps for the presentation of the coded audio and visual data and for the delivery of the data stream itself.

[0213] If subtitles are included in the stream, the subtitle decoder 3210 decodes the subtitles, synchronizes them with the video and audio frames, and provides the video / audio / subtitles to the video / audio / subtitle display 3216.

[0214] Mathematical Operators The mathematical operators used in this application are similar to those used in the C programming language. However, the results of integer division and arithmetic shift operations are more precisely defined, and additional operations such as exponentiation and real division are defined. Numbering and counting conventions generally start from 0. For example, "first" is equivalent to 0th, "second" is equivalent to 1st, etc.

[0215] Arithmetic operators The following operators are defined as follows: + Addition - subtraction (as a two-argument operator) or negation (as a unary prefix operator) * Multiplication, including matrix multiplication x y Exponentiation. Specifies x to the power y. In other contexts, such notation is used for superscripts where interpretation as a power is not intended. / Integer division. The result is truncated towards zero. For example, 7 / 4 and -7 / -4 round down to 1, and -7 / 4 and 7 / -4 round down to -1. ÷ Used to indicate division in mathematical expressions where truncation or rounding is not intended.

number

number

[0216] Logical operators The following logical operators are defined as follows: x && y Boolean logic "and" of x and y x||y Boolean logic "or" of x and y ! Boolean logic "not" x?y:z Evaluates to the value of y if x is true or not equal to 0, otherwise evaluates to the value of z.

[0217] Relational operators The following relational operators are defined as follows: > Greater than >= Greater than or equal to < Less than <= Less than or equal to == Equal != Not equal to When a relational operator is applied to a syntax element or variable that has been assigned the value "na" (not applicable), the value "na" is treated as the unique value for that syntax element or variable. The value "na" is not considered equal to any other value.

[0218] Bitwise Operators The following bitwise operators are defined as follows: & Bitwise "and". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on binary arguments that contain fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. | Bitwise "or". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on binary arguments that contain fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. ^ Bitwise "exclusive or". When acting on integer arguments, it acts on the two's complement representation of the integer value. When acting on binary arguments that contain fewer bits than the other arguments, the shorter argument is extended by adding additional significant bits equal to 0. x>>y Arithmetic right shift of y binary digits of the two's complement integer representation of x. This function is defined only for non - negative integer values of y. The bits shifted into the most significant bits (MSB) as a result of the right shift have the same value as the MSB of x before the shift operation. x<<y Arithmetic left shift of y binary digits of the two's complement integer representation of x. This function is defined only for non - negative integer values of y. The bits shifted into the least significant bits (LSB) as a result of the left shift have a value equal to 0.

[0219] Assignment operator The following arithmetic operators are defined as follows: = Assignment operator ++ Increment, i.e., x++ is equivalent to x=x + 1; when used in an array index, it is evaluated to the value of the variable before the increment operation. -- Decrement, i.e., x-- is equivalent to x=x - 1; when used in an array index, it is evaluated to the value of the variable before the decrement operation. += Increment by the specified amount. I.e., x+=3 is equivalent to x=x + 3, and x+=(-3) is equivalent to x=x+(-3). -= Decrement by the specified amount, i.e. x-=3 is equivalent to x=x-3, and x-=(-3) is equivalent to x=x-(-3).

[0220] Range Notation The following notation is used to specify a range of values: x=y…zx takes integer values ​​starting from y up to and including z, where x, y, and z are integers and z is greater than y.

[0221] mathematical functions The following mathematical functions are defined:

number

number

number

number

number

[0222] The table below specifies the precedence of operations from highest to lowest; a higher position in the table indicates a higher precedence.

[0223] For operators that are also used in the C programming language, the precedence order used herein is the same as that used in the C programming language.

[0224] Table: Operation precedence from highest (top of table) to lowest (bottom of table) [Table 14] Text description of logical operations In the text, mathematically it has the following form: if(condition0) statement 0 else if(condition 1) statement 1 … else / * Informational comments for the remaining conditions * / statement n A logical statement written in may also be written as follows: …as follows / …the following applies: -If condition 0, then statement 0 - Otherwise, if condition 1, then statement 1 -… -Otherwise (informative comment on the remaining conditions) statement n.

[0225] Each "if... then..., otherwise, if..., otherwise" statement in the text is introduced by "... it is" or "... the following applies" and is immediately followed by "if...". The final condition of an "if... then..., otherwise, if..., otherwise" may always be "otherwise". Intervening "if... then..., otherwise, if..., otherwise" statements can be identified by matching "... it is" or "... it applies" with the final "otherwise".

[0226] In the text, mathematically it has the following form: if(condition0a & condition0b) statement 0 else if(condition 1a || condition 1b) statement 1 … else statement n A logical statement written in may also be written as follows: …as follows / …the following applies: -Statement 0 if all of the following conditions are true: -Condition 0a -condition 0b - Otherwise, if one or more of the following conditions are true, then statement 1 -Condition 1a -Condition 1b -… - If not, statement n.

[0227] In the text, mathematically it has the following form: if(condition0) statement 0 if(condition1) statement 1 A logical statement written in may also be written as follows: If condition 0, then statement 0 If condition 1, then statement 1.

[0228] Although embodiments of the present disclosure have been described primarily in terms of video coding, it should be noted that embodiments of coding system 10, encoder 20, and decoder 30 (and correspondingly, system 10), as well as other embodiments described herein, may be configured for processing or coding of still pictures, i.e., processing or coding of individual pictures independent of any preceding or subsequent pictures, as in video coding. Generally, when picture processing coding is limited to a single picture 17, only inter-prediction units 244 (encoder) and 344 (decoder) may not be available. All other features (also referred to as tools or techniques) of video encoder 20 and video decoder 30 may be used equally well for still picture processing. The other functions are, for example, residual calculation 204 / 304, transform 206, quantization 208, inverse quantization 210 / 310, (inverse) transform 212 / 312, segmentation 262 / 362, intra prediction 254 / 354, and / or loop filtering 220, 320, and entropy coding 270, and entropy decoding 304.

[0229] Embodiments of, for example, the encoder 20 and the decoder 30, and the functions described herein with reference to, for example, the encoder 20 and the decoder 30, may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on a computer-readable medium or transmitted over a communication medium as one or more instructions or code and executed by a hardware-based processing unit. Computer-readable media may include computer-readable storage media, which correspond to tangible media such as data storage media, or communication media, including any medium that facilitates transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, computer-readable media may generally correspond to (1) tangible computer-readable storage media that are non-transitory, or (2) communication media such as a signal or carrier wave. Data storage media may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0230] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio waves, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio waves, and microwaves are included within the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, and instead refer to non-transitory, tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically while discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0231] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor," as used herein, may refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Furthermore, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or may be incorporated into a combined codec. Additionally, these techniques may be implemented entirely in one or more circuits or logic elements.

[0232] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including a wireless handset, an integrated circuit (IC), or a set of ICs (e.g., a chipset). This disclosure describes various components, modules, or units to highlight functional aspects of an apparatus configured to perform the disclosed techniques, but these do not necessarily require realization by different hardware units. Rather, as described above, the various units may be combined within a codec hardware unit or may be provided by a collection of interoperating hardware units including one or more processors, as described above, in association with suitable software and / or firmware.

[0233] The present disclosure discloses the following further aspects.

[0234] A first aspect of a coding method implemented by a decoding device or an encoding device, comprising: if the current block is not predicted using matrix-based intra prediction, MIP, mode, selecting a secondary transform core for a secondary transform of the current block based on an intra prediction mode index of the current block.

[0235] A second aspect of the method according to the first aspect, further comprising disabling a secondary transform of the current block if the current block is predicted using MIP mode.

[0236] A third aspect of the method according to the second aspect, wherein the step of disabling the secondary transformation of the current block includes setting a value of the secondary transformation indication information for the current block to a default value.

[0237] A fourth aspect of the method according to any one of the first to third aspects, wherein whether the current block is predicted using MIP mode is indicated according to a value of MIP indication information.

[0238] A fifth aspect of a coding method implemented by a decoding device or an encoding device, comprising: obtaining an intra-prediction mode index of a current block according to a matrix-based intra-prediction, MIP, mode index of the current block and a size of the current block; selecting a secondary transform core for a secondary transform of the current block based on an intra-prediction mode index of the current block.

[0239] A sixth aspect of the method according to the fifth aspect, wherein the intra prediction mode index of the current block is obtained according to a mapping relationship between the MIP mode index and the size of the current block, and the mapping relationship is indicated according to a predefined table.

[0240] A seventh aspect of a coding method implemented by a decoding device or an encoding device, comprising: using a secondary transform core for a secondary transform of a current block when the current block is predicted using matrix-based intra prediction, MIP, mode.

[0241] An eighth aspect of the method according to the seventh aspect, wherein the secondary conversion core is one of the secondary conversion cores used for a non-MIP mode.

[0242] A ninth aspect of the method according to the seventh aspect, wherein the secondary conversion core is different from any one of the secondary conversion cores used for non-MIP modes.

[0243] A tenth aspect of an encoder (20) having processing circuitry for carrying out the method of any one of the first to ninth aspects.

[0244] An eleventh aspect of a decoder (30) having processing circuitry for carrying out the method of any one of the first to ninth aspects.

[0245] A twelfth aspect of a computer program product comprising program code for carrying out a method according to any one of the first to ninth aspects.

[0246] A thirteenth aspect of a decoder comprising: one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the decoder to perform a method according to any one of the first to ninth aspects.

[0247] A thirteenth aspect of the encoder, comprising: one or more processors; and a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor, the programming, when executed by the processor, configuring the encoder to perform a method according to any one of the first to ninth aspects.

Claims

1. 1. A coding method implemented by an encoding device or a decoding device, comprising: Disabling a secondary transform of the current block if the current block is predicted using a matrix-based intra prediction (MIP) mode; and if the current block is not predicted using an MIP mode, when determining a selection of a secondary transform for the current block, selecting a secondary transform core set index for the current block based on an intra-prediction mode index for the current block. method.

2. You can now disable secondary transformations for blocks: setting the value of the secondary conversion instruction information for the current block to a default value; The method of claim 1.

3. 3. The method according to claim 1, wherein whether the current block is predicted using the MIP mode is indicated by the value of the MIP indication information.

4. 4. The method according to claim 1, wherein the secondary conversion core is one of the secondary conversion cores used for a non-MIP mode.

5. The method further comprises the step of: Table 15 5. The method of claim 1, wherein the method is performed according to the following formula:

6. 6. The method according to claim 1, wherein the number of secondary conversion core sets is four.

7. A computer readable medium storing computer instructions for carrying out the method of any one of claims 1 to 6.

8. one or more processors; a non-transitory computer-readable storage medium coupled to the processor and storing programming for execution by the processor; wherein the programming, when executed by the processor, configures the encoder to perform the method of any one of claims 1 to 6 when implemented by an encoding device. Encoder (20).

9. a decision unit configured to disable a secondary transform of the current block if the current block is predicted using a matrix-based intra prediction (MIP) mode; a selection unit configured to select a secondary transform core set index for the secondary transform of the current block based on an intra-prediction mode index of the current block when determining a selection of the secondary transform of the current block if the current block is not predicted using an MIP mode; Decoder (30).

10. A storage medium storing an encoded bitstream for a video signal, the encoded bitstream including a plurality of syntax elements, the plurality of syntax elements including an intra-prediction mode index and MIP indication information of a current block, the intra-prediction mode index and MIP indication information of the current block being used to select a secondary transform core set for the current block if the current block is not predicted using an MIP mode; storage medium.

Citation Information

Patent Citations

  • Transform coding based on matrix-based intra prediction

    WO2020207493A1

  • Matrix derivation in intra coding mode

    WO2020211807A1

  • Encoding device, decoding device, encoding method, and decoding method

    WO2020213677A1

  • Transform for matrix-based intra-prediction in image coding

    WO2020213944A1

  • Transform in intra prediction-based image coding

    WO2020213945A1