Lossless mode for generic video coding
By disabling lossy tools and using only lossless tools, enabling lossless mode in the video encoding and decoding system, the problem of difficulty in realizing lossless codec in the prior art is solved, and efficient lossless codec effect is achieved.
Patent Information
- Application Number
- CN202510148653.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-06-20
- Filing Date
- 2020-06-16
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video codec systems have difficulties in implementing lossless codecs, especially when dealing with compatibility issues between lossy codec tools and lossless tools.
It is proposed to disable the lossless tool by designing lossy tools, use only lossless tools, and enable lossless mode in the video codec system, and adapt the tool to enable lossless or near lossless codec, ensuring that secondary lossless codec is applied after residual codec.
It realizes lossless encoding and decoding in the video encoding and decoding system, improves the accuracy and reliability of encoding and decoding, and ensures the complete reconstruction of video data.
Smart Images

Figure CN120075442A_ABST
Abstract
Description
[0001] This divisional application is a divisional application with an application date of June 16, 2020, an application number of 202080056767.X, and an invention title of "Lossless Mode for Versatile Video Coding". Technical Field
[0002] The present disclosure belongs to the field of video compression, and at least one embodiment more specifically relates to a lossless mode for Versatile Video Coding (VVC). Background Art
[0003] To achieve high compression efficiency, image and video coding and decoding schemes typically employ prediction and transformation to leverage the spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image block and the predicted image block (usually represented as a prediction error or prediction residual) is transformed, quantized, and entropy-coded. During encoding, the original image block is typically segmented / divided into sub-blocks (e.g., quadtree segmentation may be used). To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to prediction, transformation, quantization, and entropy-coding. Summary of the Invention
[0004] A lossless coding and decoding mode is proposed in a video coding and decoding system including multiple coding and decoding tools, some of which are lossy by design and some of which can be adapted to be lossless or near-lossless. To enable the lossless mode in such a video coding and decoding system, it is proposed to disable the tools that are lossy by design and only use the lossless tools, to adapt some tools to enable lossless coding and decoding, and to adapt some tools to enable near-lossless coding and decoding, such that a secondary lossless coding and decoding can be applied after the residual coding and decoding, thereby providing lossless coding and decoding.
[0005] In a specific embodiment, a method for determining the type of residual coding and decoding includes, in the case where information indicates the use of transform skip residual coding, obtaining a flag representing a special mode, and when the flag representing the special mode is true, determining that conventional residual coding and decoding must be used instead of the transform skip residual coding that should be used.
[0006] According to a first aspect, a method for determining the type of residual coding and decoding includes, in the case where information indicates the use of transform skip residual coding, obtaining a flag representing a special mode, and when the flag representing the special mode is true, determining that conventional residual coding and decoding must be used instead of the transform skip residual coding that should be used.
[0007] According to a second aspect, a video encoding method includes, for a video block, determining a type of residual encoding and decoding according to the method of the first aspect.
[0008] According to a third aspect, a video decoding method includes, for a video block, determining a type of residual encoding and decoding according to the method of the first aspect.
[0009] According to a fourth aspect, a video encoding apparatus includes an encoder configured to determine a type of residual encoding and decoding according to the method of the first aspect.
[0010] According to a fifth aspect, a video decoding apparatus includes a decoder configured to determine a type of residual encoding and decoding according to the method of the first aspect.
[0011] One or more embodiments of the present embodiment further provide a non-transitory computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to at least a part of any of the above methods. One or more embodiments further provide a computer program product including instructions for performing at least a part of any of the above methods. Description of the Drawings
[0012] Figure 1A A block diagram of a video encoder according to an embodiment is shown.
[0013] Figure 1B A block diagram of a video decoder according to an embodiment is shown.
[0014] Figure 2 A block diagram of an example of a system in which various aspects and embodiments are implemented is shown.
[0015] Figure 3 An example of a simplified block diagram of a decoding process for PCM, lossless, and transform skip modes is shown.
[0016] Figure 4 Horizontal and vertical transforms at each SBT position are shown.
[0017] Figure 5 The LMCS architecture is shown from the perspective of the decoder.
[0018] Figure 6 The application of a secondary transform is shown.
[0019] Figure 7 A simplified secondary transform (RST) is shown.
[0020] Figure 8 Forward and inverse simplified transforms are shown.
[0021] Figure 9Shows an example of forward RST8x8 processing using a 16x48 matrix.
[0022] Figure 10A and 10B Shows an example flow chart according to at least one embodiment.
[0023] Figure 11 Shows the application of a forward then reverse function for luminance reshaping.
[0024] Figure 12 Shows an example flow chart of an embodiment including the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual decoding when transform skip is inferred to be true for lossless coding / decoding blocks.
[0025] Figure 13 Shows an example flow chart of an embodiment including the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual decoding when transform skip is inferred to be false for lossless coding / decoding blocks.
[0026] Figure 14 Shows an example flow chart of an embodiment including the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual decoding when transform skip is always parsed for lossless coding / decoding blocks.
[0027] Figure 15 Shows an example of a simplified block diagram of quadratic lossless coding / decoding.
[0028] Figure 16 Shows an example flow chart of the parsing process for region_transquant_bypass_flag and split_cu_flag when using region-level signaling.
[0029] Figure 17 Shows different partitions of luminance and chrominance in the case of a dual-tree. Detailed Description
[0030] Various embodiments relate to a post-processing method for a predicted value of sampling of an image block, the value being predicted according to an intra prediction angle, wherein the sampled value is modified after the prediction such that it is determined based on a weighted difference between the value of a left reference sample and the obtained predicted value of the sample, wherein the left reference sample is determined based on the intra prediction angle. An encoding method, a decoding method, an encoding apparatus, and a decoding apparatus based on the post-processing method are proposed.
[0031] In addition, although principles related to specific drafts of the VVC (Versatile Video Coding) or HEVC (High Efficiency Video Coding) specifications are described, the present aspects are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically excluded, the aspects described in the present application can be used alone or in combination.
[0032] Figure 1A A video encoder 100 is shown. Variations of the encoder 100 are anticipated, but for clarity, the encoder 100 is described below without describing all anticipated variations. Before being encoded, a video sequence may undergo pre-encoding processing (101), for example, applying a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components in order to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.
[0033] In the encoder 100, pictures are encoded by encoder elements as described below. For example, the picture to be encoded is segmented (102) and processed in units of CUs. Each unit is encoded using, for example, an intra or inter mode. When a unit is encoded in the intra mode, it performs intra prediction (160). In the inter mode, motion estimation (175) and compensation (170) are performed. The encoder determines (105) which of the intra mode or inter mode to use for encoding the unit and indicates this intra / inter determination, for example, by a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.
[0034] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded and decoded (145) to output a bitstream. The encoder may skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass the transformation and quantization, i.e., directly code and decode the residual without applying the transformation or quantization processing.
[0035] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse-transformed (150) to decode the prediction residuals. In the case of combining (155) the decoded prediction residuals and the predicted blocks, the image block is reconstructed. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset), Adaptive Loop Filter (ALF) filtering to reduce coding artifacts. The filtered image is stored in the reference picture buffer (180).
[0036] Figure 1B A block diagram of the video decoder 200 is shown. In the decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding pass that is reciprocal to the encoding pass described in Figure 1A The encoder 100 generally also performs video decoding as part of encoding the video data. In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 100. First, the bitstream is entropy decoded (230) to obtain transform coefficients, motion vectors, and other coding information. The picture segmentation information indicates how the picture is segmented. Thus, the decoder can partition (235) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (240) and inverse-transformed (250) to decode the prediction residuals. The image block is reconstructed by combining (255) the decoded prediction residuals and the predicted block. The predicted block can be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (280).
[0037] The decoded picture may further undergo post-decoding processing (285), such as inverse color transformation (e.g., from YCbCr 4:2:0 to RGB 4:4:4), or performing the inverse of the inverse remapping process performed in the pre-encoding process (101). The post-decoding processing may use metadata derived from the pre-encoding process and signaled in the bitstream.
[0038] Figure 2A block diagram illustrating an example of a system in which various aspects and embodiments are implemented. System 1000 may be implemented as a device including various components described below and configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described in this document.
[0039] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing various aspects described, for example, in this document. Processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., volatile storage device and / or non-volatile storage device). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drive, and / or optical disk drive. As a non-limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0040] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded video or decoded video, and encoder / decoder module 1030 may include its own processor and memory. Encoder / decoder module 1030 represents the (multiple) modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software known to those skilled in the art.
[0041] The program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to execute the various aspects described in this document can be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for the processor 1010 to run. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 can store one or more of the various items during the processing described in this document. Such stored items can include, but are not limited to, input video, decoded video or portions of the decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.
[0042] In some embodiments, the memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 1010 or the encoder / decoder module 1030) can be used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by the Joint Video Experts Team JVET).
[0043] As shown in block 1130, inputs can be provided to the elements of the system 1000 through various input devices. Such input devices include, but are not limited to: (i) an RF section that receives, for example, an RF signal transmitted over the air by a broadcaster, (ii) a COMP input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High-Definition Multimedia Interface (HDMI) input terminal. Figure 2 Other examples not shown include composite video.
[0044] In various embodiments, the input device of block 1130 has associated therewith corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal band to a band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select (e.g.) a signal band (which may be referred to as a channel in some embodiments), (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a stream of desired data packets. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0045] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via a USB and / or HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, within a separate input processing IC or within processor 1010. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed within a separate interface IC or within processor 1010. The demodulated, error-corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, which operate in combination with memory and storage elements to process the data stream as needed for presentation on an output device.
[0046] The various elements of system 1000 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected using a suitable connection arrangement 1140 (e.g., internal buses known in the art, including an inter-integrated circuit (I2C) bus, wiring, and printed circuit boards) and data may be sent therebetween.
[0047] System 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to send and receive data over the communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or a network card, and the communication channel 1060 can be implemented within, for example, a wired and / or wireless medium.
[0048] In various embodiments, a wireless network such as a Wi-Fi network (e.g., IEEE 802.11, where IEEE refers to the Institute of Electrical and Electronics Engineers) is used to stream or otherwise provide data to system 1000. The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to an external network including the Internet, for allowing streaming applications and other over-the-top communication. Other embodiments use a set-top box to provide streaming data to system 1000, and the set-top box transmits data via an HDMI connection of the input block 1130. Still other embodiments use an RF connection of the input block 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0049] System 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used in a television, a tablet computer, a laptop computer, a cellular phone (mobile phone), or other devices. The display 1100 can also be integrated with other components (e.g., as in a smart phone) or separated (e.g., an external monitor for a laptop computer). In various examples of embodiments, other peripheral devices 1120 include one or more of a stand-alone digital video disk (or digital versatile disk) (DVR for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more of the peripheral devices 1120 that function based on the output of system 1000. For example, a disk player performs the function of playing the output of system 1000.
[0050] In various embodiments, signaling using communication protocols such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention is used to communicate control signals between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120. The output devices may be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. In an electronic device (such as, for example, a television), display 1100 and speaker 1110 may be integrated with other components of system 1000 in a single unit. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0051] Display 1100 and speaker 1110 may alternatively be separated from one or more of the other components, for example if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, output signals may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0052] These embodiments may be implemented by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, an embodiment may be implemented by one or more integrated circuits. As a non-limiting example, memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as optical storage devices, magnetic storage devices, semiconductor-based storage devices, fixed memory, and removable memory. As a non-limiting example, processor 1010 may be of any type suitable for the technical environment and may include one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0053] The lossless mode is available in HEVC. In this mode, transform and quantization bypass is indicated at the CU level by a flag at the start of the CU syntax structure. If bypass is enabled, the transform and quantizer scaling operations are skipped, and the residual signal is directly coded and decoded without any degradation. Thus, this mode achieves a perfect reconstruction of the lossless representation of the coded block. The sample differences are coded as if they were at the level of the quantized transform coefficients, i.e., the transform block coding is reused with the transform sub-blocks, coefficient scanning, and last significant coefficient signaling. This mode can be useful, for example, for local coding of graphical content where quantization artifacts may be highly visible or completely intolerable. The encoder can also switch to this mode if the rate-distortion cost of normal transform coding (usually when using low quantization parameters) just exceeds the rate cost of bypass coding. For CUs coded in lossless mode, the post-filter is disabled.
[0054] If activated in the Sequence Parameter Set (SPS), PCM coding can be indicated at the CU level. If it is valid for the considered CU, prediction, quantization, and transform are not applied. Instead, the sample values of the samples in the corresponding coded block are directly coded into the bitstream with the PCM sample bit depth configured in the SPS. The granularity of the application of PCM coding can be configured between the maximum luminance coding tree block size and the minimum of 32×32 at the high end and the minimum luminance coding block size at the low end. If a coding unit is coded in PCM mode, the size of other coding units in the same coding tree unit must not be smaller than the size of the PCM unit. Since by definition PCM coding achieves a lossless representation of the corresponding block, the bits spent for PCM coding of the coding unit can be considered as an upper bound on the amount of bits required to code the CU. Thus, in cases where the application of transform-based residual coding would exceed this limit (e.g., for unusually noisy content), the encoder can switch to PCM coding.
[0055] The lossless mode of HEVC is activated using the transquant_bypass_enabled_flag flag coded in the Picture Parameter Set (PPS) (see Table 1 below). This flag enables the coding of the cu_transquant_bypass_flag at the CU level (Table 2). cu_transquant_bypass_flag Specifies bypass of quantization, transform processing, and loop filter.
[0056] Table 1: PPS syntax in HEVC for signaling lossless coding at the picture level
[0057]
[0058]
[0059] Table 2: Coding Unit Syntax for Lossless Coding and Decoding in HEVC
[0060]
[0061] If cu_transquant_bypass_flag is true for all CUs, lossless coding can be used for the entire frame, but it is also possible to perform lossless coding only on a region. This is usually the case for mixed content with overlapping text and graphics and natural video. The text and graphics regions can be losslessly coded to maximize readability, while the natural content can be coded in a lossy manner.
[0062] Furthermore, quantization is disabled for lossless coding since it is lossy by definition. In HEVC, the transform process (DCT-2 and DST-4 are used in HEVC) is lossy due to rounding operations. If the reconstructed signal is lossless, post-filters such as the deblocking filter and sample adaptive offset are not useful.
[0063] Figure 3 An example of a simplified block diagram of the decoding process for PCM, lossless, and transform skip modes is shown. The input to the decoding process 300 is the bitstream. In step 310, data is decoded from the bitstream. In particular, this provides the quantized transform coefficients and the transform type. In step 320, inverse quantization is applied to the quantized transform coefficients. In step 330, the inverse transform is applied to the resulting transform coefficients. The resulting residual is added to the prediction signal from intra or inter prediction from step 340 in step 350. The result is a reconstructed block. When the I_PCM mode is selected in step 305, the sample values are directly decoded without entropy decoding. When the lossless coding mode is selected, steps 320 and 330 are skipped. When the transform skip mode is selected, step 330 is skipped.
[0064] In the recent Versatile Video Coding (VVC) Test Model 4 (VTM4), large transform sizes up to 64×64 are enabled, which are mainly used for higher resolution videos such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, high-frequency transform coefficients are zeroed out so that only low-frequency coefficients are retained, thus losing that information. More precisely, this information is lost because it is not encoded in the resulting bitstream and is thus not available to the decoder. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 columns of the transform coefficients are retained. Similarly, when N equals 64, only the top 32 rows of the transform coefficients are retained. In both of these examples, the remaining 32 columns or 32 rows of the transform coefficients are lost. When using the transform skip mode for large blocks, the entire block is used without zeroing out any values. In fact, this includes setting the maximum size of the transform (set to 5), as shown in the syntax excerpted from the current VVC specification in Table 3 below (the relevant lines are set in bold). Thus, the maximum actual transform size is used, which is set to 32 (2^5) in the current VVC specification. This parameter will be named "max_actual_transf_size" below.
[0065] Table 3
[0066]
[0067] In addition to the DCT-2 used in HEVC, VVC also adds a Multiple Transform Selection (MTS) scheme for residual coding of inter- and intra-coded blocks. The newly introduced transform matrices are DST-7 and DCT-8. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at the SPS level, the MTS index at the CU level is signaled to indicate the separable transform pair used for the CU among DCT-2, DST-7, and DCT-8. MTS only applies to luminance. The CU-level MTS index is signaled when the following conditions are met: both the width and height are less than or equal to 32, and the CBF flag equals 1.
[0068] To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, high-frequency transform coefficients are zeroed out. Only the coefficients within the 16x16 low-frequency region are retained.
[0069] The block size limit for transform skip is the same as that for MTS, which stipulates that transform skip applies to CUs when both the block width and height are equal to or less than 32.
[0070] For an inter-predicted CU where cu_cbf equals 1, cu_sbt_flag can be signaled to indicate whether to decode the entire residual block or a sub-part of the residual block. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded / decoded with the inferred adaptive transform while the other parts of the residual block are set to zero.
[0071] Figure 4 The horizontal and vertical transforms for each SBT position are shown. The sub-block transform is a position-dependent transform applied to the luminance transform block. Two positions of SBT-H (horizontal) and SBT-V (vertical) are associated with different core transforms. More specifically, the horizontal and vertical transforms for each SBT position are specified in Figure 4 . For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the corresponding transform is set to DCT-2. Thus, the sub-block transform jointly specifies the TU tiling, cbf, and horizontal and vertical transforms of the residual block, which can be considered as a syntax shortcut in the case where the main residual of the block is on one side of the block.
[0072] In VTM4, a coding / decoding tool called Luminance Mapping and Chrominance Scaling (LMCS) is added before the loop filter as a new processing block. LMCS has two main parts: 1) in-loop mapping of the luminance component based on an adaptive piecewise linear model; 2) for the chrominance component, applying luminance-dependent chrominance residual scaling.
[0073] Figure 5 The LMCS architecture is shown from the perspective of the decoder. Figure 5 The light gray shaded blocks in indicate the positions where the processing is applied in the mapped domain; and these include inverse quantization, inverse transform, luminance intra prediction, and addition of luminance prediction and luminance residual. Figure 5 The unshaded blocks in indicate the positions where the processing is applied in the original (i.e., unmapped) domain; and these include loop filters such as deblocking, ALF, and SAO, motion compensation prediction, chrominance intra prediction, addition of chrominance prediction and chrominance residual, and storage of the decoded picture as a reference picture. Figure 5 The dark gray shaded blocks in are the new LMCS functional blocks, including forward and inverse mapping of the luminance signal and luminance-dependent chrominance scaling processing. Like most other tools in VVC, LMCS can be enabled / disabled at the sequence level using an SPS flag.
[0074] Also known as non-separable quadratic transform (NSST) or reduced quadratic transform (RST), the quadratic transform is applied between the forward main transform and quantization (at the encoder) and between dequantization and the inverse main transform (at the decoder side).
[0075] Figure 6 Shows the application of the second transformation. In JEM, as shown in the figure, for each 8×8 block, a 4×4 second transformation is applied to the small block (i.e., min(width, height) < 8), and an 8×8 second transformation is applied to the larger block (i.e., min(width, height) > 4).
[0076] The application of the non-separable transformation is described below by taking the input as an example. The 4x4 input block is represented as the following matrix:
[0077]
[0078] To apply the non-separable transformation, this 4x4 input block X is first represented as a vector
[0079] The non-separable transformation is calculated as where represents the transformation coefficient vector, and T is a 16×16 transformation matrix. The 16×1 coefficient vector is then reorganized into a 4×4 block using the scan order (horizontal, vertical or diagonal) of the block. The coefficients with smaller indices will be placed in the 4x4 coefficient block together with the smaller scan indices. There are a total of 35 transformation sets, and each transformation set uses 3 non-separable transformation matrices (kernels). The mapping from the intra prediction mode to the transformation set is predefined. For each transformation set, the selected non-separable second transformation candidate is further specified by an explicitly signaled second transformation index. This index is signaled once in the bitstream after the transformation coefficients for each CU within each frame.
[0080] Figure 7 Shows the reduced second transformation (RST). Using this technique, 16x48 and 16x16 matrices are used for 8×8 and 4×4 blocks respectively. For ease of annotation, the 16x48 transformation is denoted as RST8x8, and the 16x16 transformation is denoted as RST4x4. The main idea of the reduced transformation (RT) is to map an N-dimensional vector to an R-dimensional vector in a different space, where R / N (R < N) is the reduction factor.
[0081] Figure 8 Shows the forward and inverse reduced transformations. The RT matrix is an R×N matrix as follows:
[0082]
[0083] where the R rows of the transformation are the R bases of the N-dimensional space. The inverse transformation matrix of RT is the transpose of its forward transformation.
[0084] Figure 9An example of forward RST8x8 processing using a 16x48 matrix is shown. In the configuration adopted, a 16x48 matrix is applied instead of a 16x64 matrix with the same transform set configuration, and each matrix obtains 48 input data from three 4x4 blocks (excluding the lower right 4x4 block) in the upper left 8x8 block ( Figure 4 ). With the help of dimensionality reduction, the memory usage for storing all RST matrices is reduced from 10KB to 8KB, while the performance drops reasonably.
[0085] In addition, in VTM5, a codec tool called chroma residual joint coding and decoding has been adopted. When this tool is activated, a single joint residual block is used to describe the residuals of both the Cb and Cr blocks in the same transform unit, as shown in Equation 1.
[0086] res joint =(res Cb -res Cr ) / 2
[0087] Equation 1: Joint residual calculated from Cb and Cr residuals
[0088] Then, the Cb and Cr signals are reconstructed by subtracting the joint residual of Cb and adding it to Cr, as shown in Equation 2.
[0089]
[0090] Equation 2: Reconstructing Cb and Cr signals from joint residual coding and decoding
[0091] The coding and decoding flag is at the TU level to enable the joint coding and decoding of chroma residuals. If this flag is disabled, separate coding and decoding of Cb and Cr residuals are used.
[0092] The embodiments described below are designed taking the foregoing into consideration. Figure 1A The encoder 100, Figure 1B The decoder 200, and Figure 2 The system 1000 are adapted to implement at least one of the following embodiments.
[0093] In at least one embodiment, the present application is directed to a lossless coding and decoding mode in a video coding and decoding system, which includes a plurality of coding and decoding tools, some of which are lossy by design and some of which can be adapted to be lossless or near lossless. To enable the lossless mode in a video coding and decoding system such as VVC, the following strategies are proposed:
[0094] - Disable the tools that are lossy by design and only use lossless tools,
[0095] - Adapt some tools to enable lossless coding and decoding,
[0096] - Adapt some tools to enable near-lossless encoding and decoding (with a limited and small difference from the original signal), so that secondary lossless encoding and decoding can be applied after residual encoding and decoding, thereby providing lossless encoding and decoding.
[0097] Lossless encoding and decoding can be processed at the frame level or the region level.
[0098] Figure 10A and 10B shows an example flowchart according to at least one embodiment. The figure shows an example of general processing for determining whether a given tool should be disabled at the frame level, block level, whether it should be adapted to be lossless, or whether secondary lossless encoding and decoding should be performed to use the tool adapted to be near-lossless. Such processing can be implemented, for example, in Figure 1A the encoder device 100.
[0099] In the first step ( Figure 10A 400), it is evaluated whether the tool is lossy by design.
[0100] · If the tool is lossy, a second check ( Figure 10B 402) tests whether the tool can be disabled at the block level (in this case, only control can be performed at the frame level).
[0101] ο If the tool cannot be disabled at the block level, the value of the flag transquant_bypass_enabled_flag (usually signaled at the PPS level) is checked in Figure 10B step 404. If transquant_bypass_enabled_flag is true, the tool is disabled in Figure 10B step 410 (this applies, for example, to LMCS). If transquant_bypass_enabled_flag is false, the tool is enabled in Figure 10B step 411. Once step 410 or 411 is applied, the processing ends.
[0102] ο If the tool can be disabled at the block level, the value of the CU-level flag cu_transquant_bypass_flag is checked in Figure 10B step 405. If cu_transquant_bypass_flag is true, the tool is disabled at the CU level in Figure 10B step 408 (this applies to, for example, SBT, MTS, LFNTS). If cu_transquant_bypass_flag is false, the tool can be enabled at the CU level in Figure 10B step 409. Once step 408 or 409 is applied, the processing ends.
[0103] · If the tool is designed to be lossless, then the test is performed in step 401 of Figure 10A to check if it is lossless.
[0104]
[0105] ο If the tool is lossless, there is no special application, and the process proceeds to the end of the processing.
[0106] ο If the tool is not lossless, then in step 403 of Figure 10A it is checked if the tool is near lossless.
[0107] ■ If the tool is not near lossless, since it is not designed to be lossy either, this means that it can be made lossless (step 406 of Figure 10A ). This can be applied, for example, to an MTS that only uses lossless transforms in the MTS transform set.
[0108] ■ If the tool is near lossless, then additional (secondary) lossless encoding / decoding steps are applied (step 407 of Figure 10A ). For example, lossless transforms without quantization can be applied.
[0109] ο Once step 306 or 307 is applied, the processing ends.
[0110] In the following of this disclosure, the application of this processing to specific tools of the VVC specification is described.
[0111] Disable tools that are not compatible with lossless encoding / decoding
[0112] The first case corresponds to Figure 10A steps 410 (picture level) and 408 (CU level) of . The first element is related to the zeroing transform with large CUs. In the embodiment, in the case of a CU with lossless encoding / decoding or if transform skip is used, the zeroing transform cannot be used. The syntax specification in Table 4 illustrates this, which highlights the proposed changes compared to the current VVC syntax in italic text. If the block is larger than max_actual_transf_size x max_actual_transf_siz (which is actually 32x32 in the current VVC version) and is losslessly encoded / decoded, then all coefficients are encoded / decoded, contrary to the case of a lossy encoded / decoded block where only a part of the coefficients are encoded / decoded.
[0113] In fact, compared to the current VVC specification, the transform size limit is only applied when transform_skip_flag is false and cu_transquant_bypass_flag is false.
[0114] Table 4: Modified syntax for not using the identity transform if transform skip or lossless coding is used for the current CU
[0115]
[0116] In a variant, if the picture-level transquant_bypass_enabled_flag is enabled and if the block is larger than max_actual_transf_size x max_actual_transf_size (e.g., 32x32 in the current VVC), then quadtree partitioning is inferred or forced. In other words, when lossless coding is desired, blocks larger than 64x64 are systematically partitioned so that the block size becomes 32x32 and thus does not undergo the identity transform. Table 5 shows the modified syntax.
[0117] Table 5: Syntax for quadtree partitioning inference when lossless is enabled at the picture level
[0118]
[0119]
[0120] In another variant, the cu_transquant_bypass_flag is coded only if the current CU is less than or equal to max_actual_transf_size x max_actual_transf_size (e.g., 32x32 in the current VVC). This means that larger blocks cannot be losslessly coded. Table 6 shows the associated syntax.
[0121] Table 6: Proposed syntax for coding cu_transquant_bypass only for blocks less than or equal to 32x32
[0122]
[0123] The second element is related to sub-block transform (SBT). Since SBT uses transform tree subdivision where one of the transform units is inferred to have no residuals, this tool cannot guarantee the reconstruction of the CU in a lossless case. In an embodiment, when the lossless mode is activated (i.e., when transquant_bypass_enabled_flag is true), SBT is disabled at the CU level. The corresponding syntax changes are shown in Table 7. It is only possible to decode the SBT-related syntax and activate SBT when transquant_bypass_enabled_flag is false.
[0124] Table 7: Coding unit syntax when SBT is disabled for lossless coding of the CU
[0125]
[0126]
[0127] In an alternative implementation, the SBT can only be used for TU tiling. In this embodiment, the residuals can be decoded for each sub-block, while in the initial design, some blocks were forced to have coefficients with a value of 0. The horizontal and vertical transforms are inferred as transform skips. The corresponding syntax changes are shown in Table 8.
[0128] Table 8: Coding unit syntax when residuals are inferred for two CUs in a CU for which the SBT is not coded
[0129]
[0130]
[0131] Figure 11 The application of the forward then reverse function for luminance shaping is shown. In fact, the third element is related to luminance shaping (LMCS). Shaping is a lossy transform because applying the forward then reverse function does not guarantee giving the original signal due to rounding operations. In LMCS, the reference pictures are stored in the original domain. In the intra case, the prediction process is implemented in the "shaped" domain, and once the samples are reconstructed, the inverse shaping is applied before the loop filtering step. In the inter case, after motion compensation, the predicted signal is forward shaped. Then, once the samples are reconstructed, the inverse shaping is applied before the loop filtering step.
[0132] In an embodiment, if lossless coding is allowed in the PPS, then LMCS is disabled at the slice level. The corresponding syntax changes are shown in Table 9.
[0133] Table 9: Slice header syntax modification for disabling lmcs if transquant bypass is enabled in the PPS
[0134]
[0135]
[0136] The fourth element is related to Multiple Transform Selection (MTS) and transform skip, and embodiments involve inferring that transform skip is true. Even when the quantization step size is equal to 1, DCT and DST transforms are lossy because rounding errors cause minor losses. In the first embodiment, transform skip can only be used if the CU is losslessly coded (checked by the value of the CU-level flag cu_transquant_bypass_flag). The transform_skip_flag is inferred to be 1, and tu_mts_idx is not coded. The corresponding syntax changes are shown in Table 10. When the transform_skip_flag is inferred to be 1, residual coding for transform skip is used for lossless blocks.
[0137] Table 10: Transform unit syntax for coding to disable MTS and transform skip when CU lossless is enabled
[0138]
[0139] In VTM-5.0, two different residual coding processes can be used. The first is for non-transform skip residual coding and is effective for coding the residuals of natural content blocks. The second is for transform skip residual coding, which is effective for coding the residuals of screen content blocks. The residual coding syntax selection is shown in Table 11.
[0140] Table 11: Residual coding syntax in VTM-5.0
[0141]
[0142] Figure 12 An example flowchart of an embodiment showing the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual coding when transform skip is inferred to be true for lossless coded blocks is shown. In the first step 800, cu_transquant_bypass_flag is parsed. If the flag is true, then in step 801, transform_skip_flag is inferred to be true. Otherwise, in step 802, transform_skip_flag is parsed. If transform_skip_flag is equal to false, then in step 803, regular residuals are parsed. Otherwise, if the flag is true, then in step 804, transform skip residuals are parsed.
[0143] Figure 13FIG. 0 shows an example flowchart of an embodiment including the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual decoding when transform skip is inferred to be false for lossless coding blocks. In a first step 900, cu_transquant_bypass_flag is parsed. If the flag is true, then in step 901 transform_skip_flag is inferred to be false, and in step 904 conventional residual decoding is used. Otherwise, in step 902 transform_skip_flag is parsed. If transform_skip_flag is equal to false, then in step 904 conventional residuals are parsed. Otherwise, if the flag is true, then in step 903 transform skip residuals are parsed.
[0144] Since the residual decoding for transform skip is designed to decode the residuals from screen content coding blocks, it may be less efficient for coding natural content for lossless blocks. Thus, in another embodiment, if the CU is losslessly coded, then transform_skip_flag is inferred to be 0 and tu_mts_idx is not coded. The residual decoding for conventional transform is used for lossless coding blocks. Thus, in this special mode (i.e., when cu_transquant_bypass_flag is true), even when transform_skip_flag is true, conventional coding will be used, whereas normally when transform_skip_flag is true, transform skip residual decoding should be used.
[0145] In other words, Figure 13 the embodiment described in FIG. 8 proposes that, in cases where information indicates the use of transform skip residual decoding, the type of residual decoding is determined as follows: obtain a flag indicating a special mode, and when the flag is true, select conventional residual decoding for residual decoding instead of the transform skip residual decoding that should be used. The flag indicating the special mode can be derived from other information, such as an indication that the coding is lossless (e.g., cu_transquant_bypass_flag), or an indication that quantization, transform processing, and loop filtering are bypassed (not used), or an indication that residual decoding is forced to be conventional residual decoding.
[0146] Figure 14 FIG. 12 shows an example flowchart of an embodiment including the parsing of cu_transquant_bypass_flag, transform_skip_flag, and residual decoding when transform skip is always parsed for lossless coding blocks.
[0147] In such an embodiment, to keep the design of the lossless codec block close to that of the lossy codec block, conventional residual coding / decoding and transform skip coding / decoding can be used. In this case, the transform_skip_flag is coded / decoded for the lossless codec block. This allows competition between two different ways of coding / decoding the residual coefficients and makes the residual coding / decoding better adapt to the content.
[0148] In the first step 1400, the cu_transquant_bypass_flag is parsed, and then in step 1401, the transform_skip_flag is parsed. If the transform_skip_flag is equal to false, the conventional residual is parsed in step 1403, otherwise if the flag is true, the transform skip residual is parsed in step 1402.
[0149] In fact, the transform skip residual coding / decoding method is optimized to code / decoded computer-generated content, where the correlation between the values of spatially adjacent coefficients is strong. Generally, in lossy coding / decoding, when the content is computer-generated and effective for such content, the transform skip coefficient coding / decoding is most commonly used.
[0150] However, in this embodiment, it is proposed to use a flag to select between the conventional coefficient coding / decoding method and the transform skip coefficient coding / decoding method. In fact, the transform skip coefficient coding / decoding can be used for computer-generated content, and the conventional coefficient coding / decoding method can be used for conventional content. This allows competition between two different ways of coding / decoding the residual coefficients and makes the residual coding / decoding method better adapt to the content. The flag only changes the coefficient coding / decoding method used to code / decoded the residual of the block.
[0151] In VVC, two methods are used to code / decoded the residual:
[0152] - When using transform and quantization, the conventional coefficient coding / decoding method is used. This method inherits from HEVC with some improvements and has been designed to code / decoded the residual after transform, where the energy is compressed in the low-frequency part of the signal.
[0153] - When using only quantization, the transform skip coefficient coding / decoding method is used. This method is designed for computer-generated content, where, due to a large amount of spatial correlation between the coefficients, the transform is often skipped. The main characteristics of the transform skip residual coding / decoding method are that they do not signal the last non-zero coefficient, start from the upper left corner of the block, use a reduced template of the previous coefficients to code / decoded the current coefficient, and the sign of the coefficient is CABAC context coded / decoded.
[0154] Figure 12 、 13As shown in FIGS. 13 and 14, different from the current VVC specification, the residual encoding and decoding method can be decoupled from the fact whether the transform has been skipped. In fact, in the current VVC specification, the transform_skip_flag syntax element being equal to true means that no transform is applied to the residual of the block, and the transform skip coefficient encoding and decoding method is used. In the proposed method, if the current block is losslessly encoded (transform and quantization are skipped), the regular coefficient encoding and decoding method is used, by inferring transform_skip_flag as false, and a better coefficient encoding and decoding method is used in the case of encoding and decoding natural content. In a variant, even if cu_transquant_bypass_flag is true (in this case, transform and quantization are skipped), a flag (here transform_skip_flag is reused for this purpose) is encoded to only indicate which residual encoding and decoding method is used.
[0155] One embodiment relates to the Low Frequency Non-Separable Transform (LFNTS). In this embodiment, if the current CU is losslessly encoded, the LFNTS is disabled. The corresponding syntax changes are shown in Table 12.
[0156] Table 12: Coding Unit Syntax for Disabling LFNTS When the CU is Losslessly Encoded
[0157]
[0158] One embodiment relates to joint Cb-Cr encoding and decoding. The joint encoding and decoding process of Cb and Cr residuals is irreversible. By using Equation 1 and Equation 2, where resCb = resCr, we get resjoint = 0, and recCb = predCb, recCr = predCr. In one embodiment, if cu_transquant_bypass_flag is true, joint Cb-Cr encoding and decoding is disabled at the CU level. The corresponding syntax changes are shown in Table 13.
[0159] Table 13: Residual Encoding and Decoding Syntax for Disabling Joint Cb-Cr Encoding and Decoding if cu_transquant_bypass_flag is True
[0160]
[0161] When Equation 1 separates resjoint_Cb from res Joint_Cr, as shown in Equation 3.
[0162] Equation 3: Modification to the Joint Residual Calculation Proposed in N0347
[0163]
[0164] With the proposed modification, the process is reversible but lossy due to rounding errors, as shown in Equation 4.
[0165] Equation 4: Modification for reconstructing the Cb and Cr signals from the joint residual codec
[0166]
[0167] This variant can be used for lossless codec when performing the quadratic lossless codec process as described below.
[0168] Adaptation of VVC Tools for Lossless Codec
[0169] Reference Figure 10A , and this case corresponds to step 406.
[0170] In one embodiment, a lossless transform is added to the MTS transform set. Those transforms should be selected for lossless codec. Examples of lossless transforms are the non-normalized Walsh-Hadamard transform and the non-normalized Haar transform. Examples of near-lossless transforms are the normalized Walsh-Hadamard transform and the normalized Haar transform.
[0171] The lossless transform can be obtained by the Walsh-Hadamard transform or the Haar transform. The non-normalized transforms of Walsh-Hadamard and Haar consist of ±1 and zeros, as shown in Equation 5 and Equation 6. For example, the non-normalized 4×4 Walsh-Hadamard matrix is:
[0172] Equation 5: Non-normalized 4x4 Walsh-Hadamard matrix
[0173]
[0174] And the non-normalized 4x4 Haar matrix is:
[0175] Equation 6: Non-normalized 4x4 Haar matrix
[0176]
[0177] Lossless reconstruction can be achieved using the Walsh-Hadamard transform or the Haar transform. To explain this, consider the example of a 2x2 transform that takes the residual samples (r0 and r1). The transform matrices for Haar and Walsh-Hadamard are:
[0178] Equation 7: Non-normalized 2x2 Walsh-Hadamard or Haar matrix
[0179]
[0180] The transform coefficients (c0 and c1) are obtained by:
[0181] Equation 8: Transform coefficients calculated using a 2x2 Walsh-Hadamard or Haar unnormalized transform
[0182]
[0183] The inverse transform is performed in this way:
[0184] Equation 9: Inverse transform using a 2x2 Walsh transform or Haar unnormalized transform
[0185]
[0186] In another example of a 4x4 Walsh-Hadamard transform, to transform 4 residual samples (r0, r1, r2, r3), the matrix is multiplied by a coefficient vector to obtain transform coefficients c0, c1, c2, and c3:
[0187] Equation 10: Transform coefficients calculated using a 4x4 Walsh-Hadamard unnormalized transform
[0188]
[0189] The inverse transform is performed in this way
[0190] Equation 11: Inverse transform using a 4x4 Walsh transform unnormalized transform
[0191]
[0192] Finally, to use the Haar transform to transform 4 residual samples, use the following equation Equation 12: Transform coefficients calculated using a 4x4 Haar transform
[0193]
[0194] The inverse transform is performed in this way
[0195] Equation 13: Inverse transform using a 4x4 Haar unnormalized transform
[0196]
[0197] The unnormalized Haar or Walsh-Hadamard transform can sometimes increase the bit rate because the dynamic range of the coefficients increases. For example, when using a 4x4 unnormalized Walsh-Hadamard matrix, the coefficient energy can be up to 2 times the energy of the residual samples. Therefore, the normalized matrix is achieved by dividing the unnormalized matrix by 2. However, doing so results in a loss of integer representation. Consider the same example of using a "normalized" Walsh-Hadamard transform to transform 4 residual samples:
[0198] Equation 14: Transformation coefficients calculated using the 4x4 Walsh-Hadamard normalization transform
[0199]
[0200] Since integer representation is used for the transformation coefficients, rounding will result in an error of +-1 / 2. Therefore, additional encoding and decoding steps are required to encode and decode the hint after dividing by 2. This can be done as follows:
[0201] First, calculate the hint for the division of each coefficient:
[0202] Equation 15: Hint for the division of each coefficient calculated using the 4x4 Walsh-Hadamard normalization transform
[0203]
[0204] where the hint (rem_c 0 , rem_c 1 , rem_c 2 and rem_c 3 ) can take values (-1, 0 or 1).
[0205] Then the hint is encoded into the bitstream. First, the valid hint bit is encoded, which indicates whether the hint is non-zero, and then, if valid, the sign of the hint is encoded (negative as 0 and positive as 1). The syntax is shown in Table 15. Differential pulse code modulation (DPCM) can also be used to exploit the correlation between hints. Once the hint is decoded, the decoder can calculate the inverse transform in this way:
[0206] Equation 16: Inverse transform using the 4x4 Walsh-Hadamard normalization transform
[0207]
[0208] This technique can be used for other transform sizes, but normalization cannot always be achieved. For example, the 2x2 Walsh-Hadamard transform matrix needs to be divided by the square root of 2 for normalization, which will always result in a loss because the hint part is not just +-0.5. However, dividing by 2 as in the 4x4 case is beneficial because it reduces the increased dynamic range of the transform coefficients.
[0209] This technique also applies to the Haar transform, although due to the non-uniform norm of the Haar transform matrix, this technique does not result in a normalized transform. For illustration, the 4x4 Haar transform is explained.
[0210] Equation 17: Transformation coefficients calculated using the 4x4 Haar normalization transform
[0211]
[0212] The hint is calculated in this way:
[0213] Equation 18: Hint for the division of each coefficient calculated using the 4x4 Haar normalization transform
[0214]
[0215] where the hint takes values (0, -1, and 1). On the decoder side, the decoded hint is used for the inverse transform process in this way:
[0216] Equation 19: Inverse transform using the 4x4 Haar normalization transform
[0217]
[0218] With Haar, Walsh-Hadamard, and no transform tools, we can have multiple ways to transform the residual data in a lossless manner. Generally speaking, for the horizontal and vertical transforms of the two-dimensional residual signal, we have 9 options. The following table provides their details:
[0219] Table 14
[0220]
[0221]
[0222] A direct way to utilize these transforms is to let the encoder select the best transform pair for rate reduction and encode the index of the pair used so that the decoder can deduce the pair used and perform the inverse transform.
[0223] In the VTM, the current transform pair is either the core DCT2 transform for horizontal and vertical transforms or an additional set of DST7 and DCT8 transforms called Multiple Transform Selection (MTS). MTS can be turned on / off by a high-level flag. To conform to this design, the default lossless transform for both horizontal and vertical is no transform, and the "multiple transform" is the Walsh-Hadamard transform and the Haar transform. They can also be controlled by the high-level flag "Lossless_MTS_Flag". Therefore, the transform selection table can be modified as:
[0224] Table 15
[0225]
[0226] Generally, two non-normalized transforms fit well for residuals with small dimensions. For example, they can use sizes up to 8x8 or 4x4.
[0227] Block-level with quadratic lossless encoding and decoding
[0228] Reference Figure 10A This situation corresponds to step 407.
[0229] If the CU is lossy coded / decoded but the error is limited below a given threshold (usually if the absolute value of the error is less than or equal to 1 as shown in Equation 20), an additional coding / decoding stage can be introduced to allow lossless coding / decoding. The error is measured as the difference between the original pixels and the reconstructed samples.
[0230] Equation 20: Limited error of a lossy coding / decoding unit to which quadratic lossless coding can be applied
[0231] -1 ≤ error ≤ 1
[0232] For example, if the DCT or DST transform is used where the quantization step size is equal to 1, quadratic lossless coding can be applied after lossy coding / decoding.
[0233] Figure 15 An example of a simplified block diagram of quadratic lossless coding is shown. The input to the process is a bitstream. In step 700, data is decoded from the bitstream. In particular, this provides the quantized transform coefficients, the transform type, and the newly introduced quadratic coefficients. In step 701, inverse quantization is applied to the quantized transform coefficients. In step 702, the inverse transform is applied to the resulting transform coefficients. The resulting residual is added to the predicted signal from intra or inter prediction from step 704 in step 703. The result is a reconstructed block. When the I_PCM mode is selected in step 710, the sample values are decoded directly without entropy decoding. When the lossless coding / decoding mode is selected, steps 701 and 702 are skipped. When the transform skip mode is selected, step 702 is skipped. When the block level with quadratic transform mode is selected, step 701 is skipped, the inverse transform is applied in step 720 to obtain a first residual, and the quadratic lossless residual is added to the first residual. Then, in step 703, the resulting residual is added to the predicted signal.
[0234] Since the error is limited below a given threshold, the syntax for coding / decoding this small residual is very simple. Basically, if the threshold is 1, only a valid flag and a sign flag may be needed to code / decode the quadratic residual, and the syntax is shown in Table 16.
[0235] Table 16: Syntax for quadratic residuals with a threshold of 1
[0236]
[0237] This quadratic lossless coding can also be applied on a region basis rather than on a block basis.
[0238] Region-level signaling
[0239] In the case of lossless encoding and decoding, it is very likely to encode and decode the entire region losslessly and not mix some lossless encoding and decoding blocks with lossy encoding and decoding blocks. To handle this situation, the actual syntax needs to encode and decode the cu_transquant_bypass_flag for each CU, which can be costly.
[0240] In the first embodiment, we propose to move the flag and encode and decode it before the split_cu_flag so that all child CUs from the parent CU can share the region_transquant_bypass_flag. The associated syntax is shown in Table 17.
[0241] Table 17: Proposed syntax for region transquant bypass flag
[0242]
[0243] Figure 16 An example flowchart showing the parsing process of the region_transquant_bypass_flag and split_cu_flag when using region-level signaling is presented. When parsing the syntax of the coding tree unit, the first step (500) is to initialize two variables to false, where isRegionTransquantCoded indicates whether the flag indicating lossless encoding and decoding has been encoded for the current region, and isRegionTransquant indicates whether the current region is losslessly encoded. Then step 501 checks whether lossless encoding and decoding are enabled at the picture level (Transquant_bypass_enabled_flag is true) and the region_transquant_bypass_flag has not been encoded for the current region (isRegionTransquantCoded is false). If these conditions are true, then in step 502, the flag region_transquant_bypass_flag is parsed. Otherwise, step 504 parses the split_cu_flag. After step 502, step 503 checks whether the region_transquant_bypass_flag is true. If the region_transquant_bypass_flag is true, then in step 504, the variable isRegionTransquant is set to true. Then in step 504, the split_cu_flag is parsed. If the split_cu_flag is true, the process returns to step 500, otherwise the tree is no longer divided and the process ends (step 505). The syntax of the current coding unit can be parsed.
[0244] Lossless encoding and decoding in the dual-tree case
[0245] Figure 17 Shows different partitions of luminance and chrominance in the dual-tree case. In VVC, an encoding and decoding structure called the dual-tree allows separate encoding of the luminance tree and chrominance tree for intra slices.
[0246] In one embodiment, the cu_transquant_bypass_flag is parsed separately for each tree, which allows for more flexible signaling of lossless encoding and decoding.
[0247] In a variant embodiment, the cu_transquant_bypass_flag is not encoded and decoded for the chrominance tree, but is inferred from the luminance tree. For a given chrominance CU, if one of the co-located luminance CUs is losslessly encoded, then the current chrominance CU is also losslessly encoded. The figure shows different partitions of the luminance tree and chrominance tree. For example, if luminance CU L4 is losslessly encoded, then chrominance CUs C3 and C4 are losslessly encoded. In another example, if luminance CU L3 is losslessly encoded, then chrominance CU C2 is also losslessly encoded, even if the co-located luminance CU L2 is not losslessly encoded.
[0248] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and are generally described in a way that may sound restrictive at least to show their respective characteristics. However, this is for the purpose of clear description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with aspects described in earlier documents.
[0249] The aspects described and contemplated in this application can be implemented in many different forms. Figure 1A 、 1B Figures 1 and 2 provide some embodiments, but other embodiments can be expected, and the discussion of these figures does not limit the breadth of the implementation. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any one of the methods, and / or computer-readable storage media storing a bitstream generated according to any one of the methods.
[0250] This document describes various methods, and each of these methods includes one or more steps or actions for implementing the method. Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.
[0251] The various methods and other aspects described in this application can be used to modify the modules of video encoder 100 and decoder 200 as shown in Figure 1A and Figure 1B such as motion compensation and motion estimation modules (170, 175, 275). Additionally, this aspect is not limited to VVC or HEVC, and can be applied to, for example, other standards and recommendations, whether existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, the aspects described in this application can be used alone or in combination.
[0252] Various numerical values are used in this application. These specific values are for illustrative purposes, and the aspects described are not limited to these specific values.
[0253] Various implementations relate to decoding. As used in this application, "decoding" can include, for example, all or part of the processing performed on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processing includes one or more of the processing typically performed by a decoder. In various embodiments, such processing also includes or alternatively includes the processing performed by the decoders of the various implementations described in this application.
[0254] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a more general decoding process will be clear and is considered to be fully understood by those skilled in the art.
[0255] Various implementations relate to encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processing includes one or more of the processing typically performed by an encoder. In various embodiments, such processing also includes or alternatively includes the processing performed by the encoders of the various implementations described in this application.
[0256] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or to a more general encoding process, and is considered to be fully understood by those skilled in the art.
[0257] Note that the grammatical elements used herein are descriptive terms. Thus, they do not exclude the use of other grammatical element names.
[0258] When a figure is presented as a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.
[0259] Various embodiments relate to rate distortion optimization. In particular, during the encoding process, a balance or trade-off between rate and distortion is typically considered, usually given a limit on computational complexity. Rate distortion optimization is typically formulated as minimizing a rate distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate distortion optimization problem. For example, these methods can be based on extensive testing of all encoding options, including all considered modes or codec parameter values, and a complete evaluation of their encoding and decoding costs and the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to save encoding complexity, especially by calculating an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed signal. A hybrid of these two methods can also be used, such as by using approximate distortion for only some of the possible encoding options and complete distortion for other encoding options. Other methods only evaluate a subset of the possible encoding options. More generally, many methods employ any of a variety of techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of the encoding and decoding costs and the associated distortion.
[0260] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described and are typically described in a way that may sound restrictive at least for the purpose of showing their respective characteristics. However, this is for the purpose of clarity of description and does not limit the application or scope of those aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with aspects described in earlier documents.
[0261] The implementations and aspects described herein may be implemented, for example, as a method or process, apparatus, software program, data stream, or signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the features discussed may be in other forms (e.g., apparatus or program). The apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a tablet computer, a smart phone, a mobile phone, a portable / personal digital assistant, and other devices that facilitate information communication between end users.
[0262] References to "one embodiment" or "an embodiment" or "an implementation" or "implementations", and other variations thereof, mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in implementations", and any other variations thereof, throughout this application do not necessarily all refer to the same embodiment.
[0263] Furthermore, this application may refer to "determining" various pieces of information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0264] Furthermore, this application may refer to "accessing" various pieces of information. Accessing information may include, for example, one or more of receiving information, (e.g., retrieving information from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0265] Furthermore, this application may refer to "receiving" various pieces of information. Like "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information, or (e.g., retrieving information from a memory). In addition, "receiving" is generally involved in one way or another during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0266] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", "frame", "strip", and "tile" may be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.
[0267] It should be understood that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following " / ", "and / or", and "at least one of..." is intended to include only the selection of the first-listed option (A), or only the selection of the second-listed option (B), or the selection of both options (A and B). As another example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include only the selection of the first-listed option (A), or only the selection of the second-listed option (B), or only the selection of the third-listed option (C), or only the selection of the first and second-listed options (A and B), or only the selection of the first and third-listed options (A and C), or only the selection of the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0268] In addition, as used herein, the term "signal" among other things refers to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a particular one of the illumination compensation parameters. In this way, in an embodiment, the same parameters are used on both the encoder side and the decoder side. Thus, for example, the encoder can send (explicit signaling) a particular parameter to the decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without sending (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be done in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the term "signal" has been mentioned previously, the term "signal" can also be used as a noun herein.
[0269] It will be apparent to those of ordinary skill in the art that the various implementations can generate various signals that are formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data generated by one of the implementations. For example, a signal can be formatted to carry the bitstream of the embodiment. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, signals can be transmitted over various different wired or wireless links. Signals can be stored on a processor-readable medium.
Claims
1. A method, comprising: obtaining information representing an encoded video; and decoding the video based on the information.
2. An apparatus, comprising a processor configured to: obtain information representing an encoded video; and decode the video based on the information.
3. A method, comprising: determining encoding parameters of a video; and encoding the video based on the encoding parameters.
4. An apparatus, comprising a processor configured to: determine encoding parameters of a video; and encode the video based on the encoding parameters.
5. A non - transitory computer - readable medium, comprising program code instructions that, when executed by a processor, are used to implement the steps of the method according to claim 1.
6. A non - transitory computer - readable medium, comprising program code instructions that, when executed by a processor, are used to implement the steps of the method according to claim 3.