Reconstruction by mixing prediction and residual
A complex mixing process in video encoding and decoding addresses quantization noise by scaling and offsetting prediction residuals, improving video quality and compression efficiency.
Patent Information
- Application Number
- JP2024575357
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-22
- Publication Date
- 2025-07-30
AI Technical Summary
Existing video encoding and decoding methods fail to effectively address quantization noise in the reconstruction process, leading to suboptimal video quality due to the simple addition of prediction and residual samples without correcting for quantization errors.
A more complex mixing process is introduced in the reconstruction step, involving operations such as scaling and offset adjustments of the decoded prediction residual and prediction samples, with parameters determined to minimize loss functions and constrained by quantization parameters.
The proposed mixing process significantly reduces quantization noise, enhancing video quality by accurately reconstructing samples and improving compression efficiency.
Smart Images

Figure 2025524454000001_ABST
Abstract
Description
Technical Field
[0001] This embodiment generally relates to methods and apparatuses for reconstruction in video encoding or decoding.
Background Art
[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction and transformation to exploit spatial and temporal redundancies within video content. Generally, intra prediction or inter prediction is used to utilize intra-picture or inter-picture correlations, and then the difference between the original block and the predicted block, often called the prediction error or prediction residue, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by inverse processes corresponding to entropy coding, quantization, transformation, and prediction.
Summary of the Invention
[0003] According to one embodiment, a method of video decoding is presented. The method includes obtaining a first data set corresponding to a predicted block of a block of a picture, obtaining a second data set corresponding to a decoded prediction residue for the block of the picture, adjusting the second data set to form an adjusted second data set based at least on mixing parameters, and combining the first data set and the adjusted second data set to form a decoded version of the block of the picture.
[0004] According to another embodiment, a method of video encoding is presented, the method including: obtaining a first dataset corresponding to a predicted block of a block of a picture; obtaining a second dataset corresponding to a reconstructed prediction residual for the block of the picture; adjusting the second dataset to form an adjusted second dataset based at least on mixing parameters; and combining the first dataset and the adjusted second dataset to form a reconstructed version of the block of the picture.
[0005] According to another embodiment, a video decoding apparatus is provided, including one or more processors and at least one memory coupled to the one or more processors, the one or more processors being configured to obtain a first dataset corresponding to a predicted block of a block of a picture, obtain a second dataset corresponding to a decoded prediction residual for the block of the picture, adjust the second dataset to form an adjusted second dataset based at least on mixing parameters, and combine the first dataset and the adjusted second dataset to form a decoded version of the block of the picture.
[0006] According to another embodiment, a video encoding apparatus is provided, including one or more processors and at least one memory coupled to the one or more processors, the one or more processors being configured to obtain a first dataset corresponding to a predicted block of a block of a picture, obtain a second dataset corresponding to a reconstructed prediction residual for the block of the picture, adjust the second dataset to form an adjusted second dataset based at least on mixing parameters, and combine the first dataset and the adjusted second dataset to form a reconstructed version of the block of the picture.
[0007] One or more embodiments further provide a computer program including instructions that, when executed by one or more processors, cause the one or more processors to execute an encoding method or a decoding method according to any of the embodiments described herein. One or more of the embodiments of the present invention also provide a computer-readable storage medium storing instructions for encoding or decoding video according to the methods described herein.
[0008] One or more embodiments further provide a computer-readable storage medium storing video data generated by the methods described herein. One or more embodiments further provide a method and an apparatus for transmitting or receiving video data generated by the methods described herein.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Best Mode for Carrying Out the Invention
[0010] FIG. 1 shows a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including various components described below and is configured to execute one or more of the aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. The elements of system 100 may be embodied as a single integrated circuit, multiple ICs, and / or individual components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed across multiple ICs and / or separate components. In various embodiments, system 100 is communicatively coupled to other systems or other electronic devices, for example, through a communication bus or dedicated input ports and / or output ports. In various embodiments, system 100 is configured to implement one or more of the aspects described in this application.
[0011] System 100 includes at least one processor 110 configured to execute instructions loaded therein, for example, to implement various aspects described in this application. The processor 110 may include a built-in memory, an input / output interface, and various other circuits known in the art. System 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include, by way of non-limiting example, an internal storage device, a removable storage device, and / or a network-accessible storage device.
[0012] System 100 includes, for example, an encoder / decoder module 130 configured to process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included within a device to perform encoding and / or decoding functions. As is known, the device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 130 may be implemented as a separate element of System 100 or may be incorporated within the processor 110 as a combination of hardware and software known to those skilled in the art.
[0013] To execute the various aspects described in this application, the program code loaded onto processor 110 or encoder / decoder 130 may be stored in storage device 140 and then loaded onto memory 120 for execution by processor 110. According to various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of the various items during execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or a portion of the decoded video, bitstream, matrix, variable, and intermediate or final results from processing of equations, expressions, operations, and operation logic.
[0014] In some embodiments, the internal memory of processor 110 and / or encoder / decoder module 130 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other embodiments, an external memory of the processing device (e.g., the processing device may be either processor 110 or encoder / decoder module 130) is used for one or more of these functions. The external memory may be memory 120 and / or storage device 140, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, an external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a high-speed external dynamic volatile memory such as RAM is used as a working memory for video coding and decoding operations such as MPEG-2, HEVC, or VVC.
[0015] Inputs to the elements of system 100 may be provided through various input devices, as shown in block 105. Such input devices may include, but are not limited to, (i) an RF section that receives RF signals wirelessly transmitted, for example, by a broadcasting station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0016] In various embodiments, the input device of block 105 has respective input processing elements known in the art. For example, the RF section may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) in certain embodiments, band-limiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel for example, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments may include one or more elements for performing these functions, such as a frequency selector, signal selector, band limiter, channel selector, filter, down-converter, demodulator, error corrector, and demultiplexer. The RF section may include a tuner for performing these various functions, including for example down-converting a received signal to a lower frequency (such as an intermediate frequency or a near baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving an RF signal transmitted over a wired (e.g., cable) medium, filtering, down-converting, and filtering again to a desired frequency band. In various embodiments, the order of the above (and other) elements is rearranged, some of these elements are omitted, and / or other elements performing similar or different functions are added. Adding elements may include inserting elements between existing elements, for example inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0017] In addition, the USB and / or HDMI terminals may each include an interface processor for connecting the system 100 to other electronic devices across the USB and / or HDMI connection. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor 110 as needed. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 110. Demodulation, error correction, and the demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder 130 that combines memory and storage elements to perform operations, to process the data streams necessary to present to the output devices.
[0018] The various elements of the system 100 may be provided within an integrated housing, in which the various elements are interconnected using an internal bus known in the art, including a suitable connection configuration 115, such as an I2C bus, wiring, and a printed circuit board, and can transmit data to each other.
[0019] The system 100 includes a communication interface 150 that enables communication with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented, for example, within a wired medium and / or a wireless medium.
[0020] In various embodiments, data is streamed to system 100 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of such embodiments are received via communication channel 190 and communication interface 150 adapted for Wi-Fi communication. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to enable streaming applications and other over-the-top communications. In other embodiments, a set-top box that distributes data via the HDMI connection of input block 105 is used to provide the streamed data to system 100. In still other embodiments, the RF connection of input block 105 is used to provide the streamed data to system 100.
[0021] System 100 may provide the output signal to various output devices, including display 165, speaker 175, and other peripheral devices 185. Other peripheral devices 185 may include, in various examples of the embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 100. In various embodiments, the control signal is communicated between system 100 and display 165, speaker 175, or other peripheral devices 185 using signaling such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using communication channel 190 via communication interface 150. Display 165 and speaker 175 may be integrated into a single unit with other components of system 100 in an electronic device, such as a television, for example. In various embodiments, display interface 160 includes a display driver, such as a timing controller (T Con) chip.
[0022] Alternatively, display 165 and speaker 175 may be separated from one or more of the other components, for example, if the RF section of input 105 is part of a separate set-top box. In various embodiments where display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0023] FIG. 2 shows an example of a video encoder 200, such as a Versatile Video Coding (VVC) encoder. FIG. 2 may further show an encoder with improvements to the VVC standard, or an encoder that employs technology similar to VVC.
[0024] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "encoded" and "coded" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Usually, although not necessarily, the term "reconstructed" is used on the encoder side and the term "decoded" is used on the decoder side.
[0025] Before encoding, the video sequence may undergo pre-encoding processing (201), such as applying a color conversion to the input color picture (e.g., conversion from RGB4:4:4 to YCbCr4:2:0), or performing remapping of the input picture components (e.g., using histogram equalization of one of the color components) to obtain a more flexible signal distribution for compression. Metadata can be associated with the pre-processing and attached to the bitstream.
[0026] In encoder 200, as follows, the picture is encoded by encoder elements. The picture to be encoded is divided (202) into units, such as coding units (CUs), and processed. Each unit is encoded using, for example, either an intra mode or an inter mode. When a unit is encoded in the intra mode, intra prediction (260) is performed. In the inter mode, motion estimation (275) and motion compensation (270) are performed. The encoder determines (205) which of the intra mode or the inter mode should be used to encode the unit, and indicates the intra or inter decision, for example, by a prediction mode flag. After prediction, prediction enhancement (285) is applied to the prediction block. The prediction residual is calculated, for example, by subtracting the predicted block from the original image block (210).
[0027] Next, the prediction residual is transformed (225) and quantized (230). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy coded (245) to output a bitstream. The encoder can skip the transformation and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transformation and quantization, i.e., the residual is directly coded without applying the transformation process or the quantization process.
[0028] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (240), inverse transformed (250), and the prediction residual is decoded. The decoded prediction residual and the predicted block are combined (255) to reconstruct an image block. An in-loop filter (265) is applied to the reconstructed picture, e.g., to perform deblocking / sample adaptive offset (SAO) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (280).
[0029] FIG. 3 shows a block diagram of an exemplary video decoder 300. As described below, in decoder 300, the bitstream is decoded by decoder elements. Video decoder 300 generally performs a decoding path that is inverse to the encoding path described in FIG. 2. Further, encoder 200 generally performs video decoding as part of video data encoding.
[0030] In particular, the input to the decoder includes a video bitstream, which can be generated by the video encoder 200. The bitstream is first entropy decoded (330) to obtain transform coefficients, motion vectors, and other coded information. The picture partitioning information indicates how the picture is partitioned. Thus, the decoder can split the picture according to the decoded picture partitioning information (335). The transform coefficients are inverse quantized (340), inverse transformed (350), and the prediction residual is decoded. The decoded prediction residual is combined with the predicted block (355) to reconstruct the picture block. The predicted block can be obtained from intra prediction (360) or motion compensation prediction (i.e., inter prediction) (375) (370). After prediction, prediction enhancement (390) is applied to the predicted block. An in-loop filter (365) is applied to the reconstructed picture. The filtered picture is stored in the reference picture buffer (380).
[0031] The decoded picture can further undergo post-decoding processing (385), such as inverse color conversion (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4), or inverse remapping that performs the inverse of the remapping process executed in the pre-encoding processing (201). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0032] The proposed method relates to a reconstruction step corresponding to an addition step (255 on the encoder side in FIG. 2, 355 on the decoder side in FIG. 3). In normal video coding standards or consortium solutions such as AVC, HEVC, VVC, and AV1, this reconstruction step, for the samples of a given block B, adds the decoded prediction residual samples Res’(p) (also called the residual), obtained from the inverse transform step (250, 350), to the prediction samples Pred(p) obtained from the prediction step (205, 370) or the prediction enhancement step (285, 390) in the absence of prediction enhancement, to obtain the reconstructed samples Rec(p) for each sample position p within the block. Rec(p)=Pred(p)+Res’(p) Equation 1
[0033] The residual samples Res(p) of the transform block are obtained on the encoder side by calculating the difference between the original samples Orig(p) and the prediction samples Pred(p). The resulting difference samples are then processed by a transform and then quantized. To reconstruct the block, the encoder then performs inverse quantization and inverse transform. The output is a block of decoded prediction residual samples (Res’(p)). This represents the sample difference corrupted by quantization noise Err(p). Res’(p)=Orig(p)-Pred(p)+Err(p) Equation 2 The reconstruction step does not take into account or correct this error (Err(p)). A further in-loop filter processing step is intended to solve the problem, but is also related to attempting to reduce the error in the reconstruction step itself.
[0034] This document proposes replacing the simple addition of the prediction sample block and the decoded prediction residual sample block in the reconstruction step with a mixing process. The mixing process is considered as a function f(Pred(p), Res’(p) and involves operations more complex than simple addition.
[0035] Replacement of addition by a mixing process for deriving a reconstructed sample FIG. 4 shows a modified encoder scheme (400) according to an embodiment. In particular, the addition step (255) is replaced by a more complex mixing step (295) than a simple addition of a predicted residual sample and a decoded predicted residual sample.
[0036] FIG. 5 shows a modified decoder scheme (500) according to an embodiment. In particular, the addition step (355) is replaced by a more complex mixing step (395) than a simple addition of a predicted residual sample and a decoded predicted residual sample.
[0037] The mixing steps on the encoder side and the decoder side (295, 395) are executed similarly. The mixing process of the prediction block and the residual block generates a reconstructed sample based on a mixing function. For all positions p within the block, Rec(p)=f(Pred(p),Res’(p) Equation 3 where f(a,b) is different from a simple addition. The function may further depend on the sample position p. Different implementations of the function f(.) are described below.
[0038] Scaling correction of the decoded predicted residual sample In the first embodiment, the mixing process performs the following operation on the samples of the block. Rec(p)=Pred(p)+β(p) * Res’(p) Equation 4 where β(p) is a weight coefficient (scaling coefficient) applied to the decoded predicted residual sample and enables correction of the quantization error of the residual. The value of β(p) is logically constrained to be close to 1 in the interval I = [1 - d, 1 + d], for example d = 0.1, because Res’(p) is expected to be close to the actual (unquantized) difference value Res(p) (Orig(p) - Pred(p)). More generally, the constraint interval can be I = [1 - d1, 1 + d2], where d1 and d2 can take the same value or different values.
[0039] In one example, the range of the interval depends on the quantization parameter QP used to quantize the block residual: I = [1 - d(QP), 1 + d(QP)], where 2 * d(QP) is the range of the interval.
[0040] Equation 4 is written in floating-point format. However, in practice, fixed-point arithmetic is used in the actual implementation. For example, Equation 4 can be implemented as follows. Rec(p) = Pred(p) + (β(p) * Res’(p) + rnd_offset) >> K Equation 5 where β(p) is a fixed-point value, K is a predefined fixed-point value, the operator >> corresponds to a right shift that moves the bit values of a binary value, and rnd_offset is a fixed rounding offset, for example 2 K-1 equal to.
[0041] For the sake of notation simplification, in the following equations, the sample position dependency “(p)” is removed when it is not necessary.
[0042] Derivation of the mixing parameter from adjacent samples The inference of β(p) (in the encoder and decoder) can be based on the adjacent reconstructed samples Rec, the predicted samples, and the decoded prediction residual samples, for example, the N rows above and the left column of the adjacent samples of the current block, where N is equal to 4 for example.
[0043] For example, the values of β for different reconstructed sample values can be estimated from adjacent samples (within the adjacent region ν related to the current block), knowing that the reconstructed adjacent samples follow Equation 4.
[0044] The sample range [0~1023] for a 10-bit signal is P non-overlapping intervals I1...I Pcan be divided. The intervals can have equal lengths (1024 / P for a 10-bit signal). Examples of P values are 16 (with an interval length of 64) or 8 (with an interval length of 128) to reduce complexity. For each interval I k for k = 1...P, the following can be applied. For all positions p within the adjacent region ν, β k is derived from all p such that Pred(p) is within I k for example, to minimize the mean squared error (MSE):
[0045] [Number] β k The value of is then used within the current block for all samples for which the predicted sample value Pred(p) is within I k For all p in the current block for which the predicted sample value Pred(p) is within I k it is as follows. Rec(p) = Pred(p) + β k (p) * Res’(p) Equation 7
[0046] Thus, each position within the current block (denoted as Blk) has a value β(p) corresponding to β k (p), and k is the index of the interval to which Pred(p) belongs.
[0047] In another example, to simplify the mixed parameter derivation, instead of several piecewise (interval-dependent) values, a single value β is applied to all samples in the current block (P = 1). Similar to the above, β is derived from all positions p within the adjacent region ν, for example, to minimize the MSE:
[0048] [Number]
[0049] In another example, to further simplify the mixed parameter derivation, a single value βQP It is applied to all samples quantized with the same QP value, rather than some values for different blocks. All previous reconstructed samples (or can also be restricted in some adjacent regions ν) are quantized with the same QP, and β QP is used to derive, for example, minimizing the MSE:
[0050]
Number
[0051] Therefore, each position within the current block (denoted as Blk) has a value β QP (p) corresponding to β(p), and QP is the quantization parameter applied to the block containing the sample at position p.
[0052] In the above, the mixing parameters (β, β k , β QP ) are derived with the loss function MSE (L2 norm). Other loss functions such as the mean absolute error (MAE, L1-norm), or the Huber loss function can also be used (it behaves like MSE for small errors, but behaves like MAE adjusted by a hyperparameter that can be associated with the QP value for large errors).
[0053] Use of regularization term In one example, a regularization term is used (added) to constrain β k (p), which belongs to the neighborhood V(p) of the pixel position p and is close to those in its neighborhood.
[0054] For example, β k is calculated to minimize the following function.
[0055]
Number
[0056] [Number]
[0057] As in the previous case, since the value β(q) is considered fixed in this equation, this may require several iterations. For example, the first iteration is applied without the regularization term (λ = 0). This yields a first estimate of β(p) for any p within the block. The next iteration is applied to the regularization term. The process stops after a given number of iterations (e.g., 3) or when the variation of β(p) becomes small.
[0058]
[0059] Alternatively, regularization is applied after the first estimation step based on Equation 6 or Equation 9. After this first estimation step, within the block, the estimated β(p) is available at each position p within the block. Next, a second step is applied to normalize (smooth) the β(p) values, for example, by simply applying a low-pass filter to the block of β(p) values.The regularization term can also be associated with the prior probability of the signal Proba(rec=r). For example, this probability can be calculated from previously encoded frames. In an embodiment, it is calculated from previously encoded frames with the same temporal ID. Alternatively, it can be encoded once per scene cut, per intra picture, per group of pictures (GOP), or per intra slice. The probability value can be provided in the form of a histogram H(x) covering the signal range (e.g., x is 0, ..., 1023 for a 10-bit signal). Alternatively, it can be modeled by a parametric function (e.g., polynomial) or a piecewise parametric function.
[0060] Referring to Equation 8 or Equation 10, the regularization term can consist of adding a term proportional to. -log(Proba(Rec(p)) This is proportional to the following equation. -log(Proba(Pred(p)+β(p) * Res’(p)))。 This regularization term constrains β(p) so as to tend to maximize this probability.
[0061] Signaling mixing parameter Additional syntax elements (e.g., flags) can be sent for each block (TU, CU, or CTU) to indicate whether the proposed mixing should be used for that particular block rather than simple addition. The syntax elements can be binary or non-binary when several mixing models are used. For example, a 1-bit syntax element is used when a single-piecewise model (P = 1) is used, an 8-bit syntax element is used when an 8-piecewise model (P = 8) is used, and a 16-bit syntax element is used when a 16-piecewise model (P = 16) is used. In a variant, the mixing parameters can be explicitly encoded and sent for each block to which they apply.
[0062] As described above, the value of β should be logically constrained to be close to 1, for example, in the interval I = [1 - d, 1 + d]. To signal the mixing parameters, many bits may be required to code these mixing parameters. Furthermore, the complexity of the calculations in the encoder can become quite high. To limit the bit cost and the complexity of the calculations, in one example, the mixing parameter β can only be selected from a predefined set having M values, for example M = 4 or 8. In this case, only the corresponding index needs to be inferred or coded and transmitted. The number of possible mixing parameters in the predefined set M can also depend on the slice type. With respect to these possible mixing parameter values, they can be varied and adapted to different QPs or content via some offline learning.
[0063] When the value of β is signaled, it is assumed to be estimated on the encoder side before being signaled. This can be done by rate-distortion optimization, where the distortion is estimated, for example, as follows.
[0064]
Equation
[0065] In the above, the decoded prediction residual is scaled and then added to the predicted sample in the mixing process. Other forms of mixing functions can also be used.
[0066] In one example, the decoded prediction residual Res’(p) is adjusted by an offset: Rec(p)=Pred(p)+(Res’(p)+b(p)) Equation 11
[0067] In another example, the decoded prediction residual Res’(p) is scaled by a scaling factor and adjusted by an offset: Rec(p)=Pred(p)+a(p) *Res’(p) + b(p), Equation 12
[0068] In yet another example, the prediction sample Pred(p) is also scaled by a scaling factor a(p). Rec(p) = a(p) * Pred(p) + b(p) * Res’(p) + c(p), Equation 13
[0069] Mixing in the transform domain In another embodiment, the mixing is performed in the transform domain. In the encoder (600), as shown in FIG. 6, a new transform step (296) applied to the prediction signal is inserted after the prediction enhancement step (285). The mixing step (295) is applied after the inverse quantization (240) and uses the output of the new transform step (296). The output of the mixing step (295) is processed by the inverse transform step (297) and returned to the pixel domain.
[0070] In VVC, the transform process consists of two steps: MTS (Multiple Transform Selection) followed by LFNST (Low Frequency Non-Separable Transform). MTS is composed of a set of multiple transforms, and for a transform unit (TU), one transform is selected (signaled or inferred) from the set. On the encoder side, when the transform from MTS is applied to the prediction residual, the secondary transform LFNST can be applied. On the decoder side, the inverse transform is the reverse process, i.e., applying the inverse LFNST followed by the inverse transform from MTS.
[0071] The MTS transform matrix used in 296 must be the same as that used in step 225 (same MTS transform size, same MTS transform matrix coefficients). In an embodiment, in 296, both the MTS transform and the LFNST are applied (when the LFNST is activated for the TU). In another embodiment, only the MTS transform is applied in step 296. In this case, an additional inverse LFNST step should be applied during the inverse quantization (240) before the mixing step (295) because they are in the same transform domain.
[0072] In the decoder (700), as shown in FIG. 7, a new transformation step (396) applied to the prediction signal is inserted after the prediction enhancement step (390). The mixing step (395) is applied after the inverse quantization (340) and uses the output of the new transformation step (396). The output of the mixing step (395) is processed by the inverse transformation step (397) and returns to the pixel region. Here, it should be noted that the mixing is applied to the inverse quantized transform coefficients before the inverse transformation, that is, the decoded prediction residual samples in the transform region. Similar to the mixing in the sample region (see FIGS. 4 and 5), the inverse quantized transform coefficients can be scaled and / or shifted by an offset when combined with the (scaled) transformed prediction samples.
[0073] The MTS transform matrix 396 must be the same as that used in steps 296 and 225. The inverse MTS transform matrix 397 must be the same as that used in step 297. In an embodiment, in 396, both MTS transform and LFNST are applied (when LFNST is activated for the TU). In another embodiment, only the MTS transform is applied in step 396. In this case, since they are in the same transform region, the inverse LFNST step should be applied between the inverse quantization (340) and the mixing step (395).
[0074] When the mixing process is applied in the transform region, the adjustment (scaling or offset) of the inverse quantized prediction residual can be regarded as an improvement of the inverse quantized prediction residual, or the adjustment in the mixing process can be regarded as an additional step of the inverse quantization process. The mixing process can reduce the distortion caused by quantization. That is, instead of passively performing inverse quantization based on the parameters designed for quantization, here, the inverse quantization functions more actively to reduce distortion.
[0075] The advantage of operating in the transform domain is that the quantization interval is fully known for each quantization coefficient. This interval is defined from the quantization parameters that control the quantization step. The quantization step depends linearly on the scaling coefficient "ls[x][y]", as derived, for example, from equations (1141) and (1142) of the VVC specification (ITU-T H.266, SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, Versatile Video Coding, 08 / 2020). The inverse quantization of the decoding coefficient dz[x][y] to obtain the inverse quantization coefficient dnc[x][y] is achieved by equation (1145) of the VVC specification. dnc[x][y]=(dz[x][y] * ls[x][y]+bdOffset)>>bdShift Equation 14 where bdOffset and bdShift are parameters that depend on the bit depth of the signal.
[0076] Therefore, the original coefficient coef[x][y] is expected to be within the interval d. Iorig=[dnc[x][y]-Delta / 2;dnc[x][y]+Delta / 2] Equation 15 Comprising the following Delta=(ls[x][y]>>bdShift) Equation 16
[0077] This property can be used to constrain the hybrid processing. Usually, this means that the original coefficient (in the transform domain) is within this interval and is equally probable for all values within this interval.
[0078]
Number
[0079] Clipping in the transform domain The constraints of Equation 15 can be used to perform clipping in the conversion domain. When the mixing process is executed in the conversion domain (Steps 295 and 395), the inputs to the mixing process are the conversion coefficients of the predicted signal (Cpred[x][y]) and the inverse quantized residual coefficients (Cres’[x][y]), the output from the mixing process is cf = Crec’[x][y], and cf is assumed to belong to the interval [Crec[x][y] - Delta / 2, Crec[x][y] + Delta / 2], where Crec[x][y] = Cpred[x][y] + Cres’[x][y], which means that cf can be clipped according to the following equation. cf = Clip3(Crec[x][y] - Delta / 2, Crec[x][y] + Delta / 2, cf), comprising the following.
[0080]
Number
[0081] As shown in FIG. 4 or FIG. 5, when the mixing process is executed in the sample domain (instead of the conversion domain), the clipping process shall follow the following steps as shown in FIG. 8. - Apply a conversion to the input reconstruction sample block Rec’ obtained after the mixing process to obtain the conversion coefficient block Crec’, and apply a conversion to the prediction sample block Pred to obtain the conversion coefficient block Cpred (Step 810). The conversion must be the same as that used in Steps 225, 296, and 396. Crec’ = transform(Rec’); Cpred = transform(Pred) - Clip each coefficient cf = Crec’[x][y] according to the following equation (Step 820). cf = Clip3(Crec[x][y] - Delta / 2, Crec[x][y] + Delta / 2, cf) where Crec[x][y] = Cpred[x][y] + Cres’[x][y], and Cres’[x][y] is the inverse quantized residual coefficient Next, Crec’[x][y] is modified as cf:Crec’[x][y]=cf. - To obtain the modified reconstructed sample block Rec, apply the inverse transform to the modified conversion coefficient block Crec’ (830). The inverse transform must be the same as that used in steps 250, 350, 297, and 397. Rec’ = inverse_transform(Crec’).
[0082] In the case of transform skip When the transform is skipped (from the encoder decision), the residual of each sample is within the quantization interval defined by [-Delta / 2;Delta / 2]. This property can be used to constrain the mixing process. Usually, this means that the original sample (in the sample domain) is within this interval and is equally probable for all values within this interval.
[0083] [Number]
[0084] Alternatively, when the transform skip is applied, the clipping process can be applied directly in the sample domain after step 255 and 355, or any step following 295 and 395. If the input to the mixing process is the predicted signal (Pred[x][y]) and the inverse quantized residual coefficient Res’[x][y], and the output from the mixing process is R = Rec’[x][y], then R is assumed to belong to the interval [Rec[x][y]-Delta / 2, Rec[x][y]+Delta / 2], where Rec[x][y]=Pred[x][y]+Res’[x][y], which means that R can be clipped according to the following equation. R = Clip3(Rec[x][y]-Delta / 2, Rec[x][y]+Delta / 2, R).
[0085] FIG. 9 shows the mixing process in the sample area for a block to be encoded or decoded according to an embodiment. Whether or not there is prediction refinement, an encoder or a decoder can obtain a prediction block from intra prediction or inter prediction (910). The decoded prediction residual (Res’) is also obtained, for example, after inverse quantization and inverse transformation of the transform coefficients (910). The mixing parameter of the block is obtained at 920. As described above, the mixing parameter can be a scaling coefficient for scaling the prediction residual (930), or an offset for adjusting the prediction residual (930). On the encoder side, in order to minimize the loss function, the mixing parameter can be obtained, for example, as described in Equations 6-10. On the decoder side, for example, a set of parameters can be predefined or decoded, and a specific mixing parameter is selected for the block from the set of parameters. Then, the prediction block and the decoded prediction residual can be combined, for example, as described in Equations 11-13 (940).
[0086] FIG. 10 shows a method for obtaining a mixing parameter according to an embodiment. In this embodiment, in step 1010, a set of mixing parameters β(i) (i = 1,..., M) is obtained, for example, decoded from a bitstream or predefined in a decoder for a slice or a picture. In step 1020, the decoder decodes a syntax element defining an index k into the set of mixing parameters for the current block. In step 1030, the mixing parameter is set as β(k).
[0087] FIG. 11 shows a method for determining a mixed parameter according to an embodiment. In this embodiment, in step 1110, a set of mixed parameters β(i) (i = 1,..., M) is obtained. For example, it is decoded from a bit stream or predefined in a decoder for a slice or a picture. Each mixed parameter in the set corresponds to a different QP. In step 1120, the decoder obtains the quantization parameter QP for the current block. In step 1130, the mixed parameter is set as the mixed parameter corresponding to QP (e.g., β(QP)).
[0088] FIG. 12 shows a method for determining a mixed parameter according to an embodiment. In this embodiment, in step 1210, a set of mixed parameters β(i) (i = 1,..., P) is obtained. For example, it is decoded from a bit stream or predefined in a decoder for a slice or a picture. Each mixed parameter in the set corresponds to a sample value interval. The decoder loops through all samples in the current block. In step 1220, the decoder obtains the sample value Pred(p) for the current sample. In step 1230, the decoder determines the interval k of Pred(p) (i.e., the sample value Pred(p) belongs to the interval k). In step 1240, the mixed parameter is set as the mixed parameter β(k).
[0089] Various methods are described herein, and each of the methods includes one or more steps or acts for achieving the described method. The order of the steps or acts, and / or the use of specific steps and / or acts, can be modified or combined, provided that a particular order of steps or acts is not required for proper operation of the method. Additionally, terms such as "first", "second", etc. can be used in various embodiments to modify elements, components, steps, operations, etc., for example, as "first decoding" and "second decoding". The use of such terms does not imply an ordering with respect to the modified operations, unless specifically required. Thus, in this example, the first decoding need not be performed before the second decoding and can occur, for example, before, during, or overlapping with the second decoding.
[0090] The various methods and other aspects described in this application can be used to modify modules of video encoders 200 and decoders 300, such as reconstruction modules (255, 355), as shown in FIGS. 2 and 3. Further, this aspect is not limited to VVC or HEVC and can be applied, for example, to other standards and recommendations, and extensions of any such standards and recommendations. Unless otherwise indicated or technically excluded, the aspects described in this application can be used individually or in combination.
[0091] In this application, various numerical values are used. The specific values are for illustrative purposes only, and the described aspects are not limited to these specific values.
[0092] Various implementations involve decoding. As used in this application, "decoding" may include, for example, all or a portion of the processing performed on an encoded sequence received to produce a final output suitable for display. In various embodiments, such processing may include, for example, one or more of the processing typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or to a more general decoding process as a whole will become apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.
[0093] Various implementations involve encoding. Similar to the above considerations regarding "decoding", "encoding" as used in this application may include, for example, all or a portion of the processing performed on an input video sequence to produce an encoded bitstream.
[0094] Note that the syntactic elements used in this specification are for illustrative purposes only. Thus, they do not preclude the use of other syntactic element names.
[0095] The implementations and aspects described in this specification may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if considered only in the context of a single form of implementation (e.g., considered only as a method), the implementation of the features considered may also be implemented in other forms (e.g., an apparatus or a program). The apparatus may be implemented, for example, with appropriate hardware, software, and firmware. The method may be implemented, for example, in an apparatus such as a processor that refers to general processing devices including a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor may further include, for example, a communication device such as a computer, a mobile phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate the communication of information between end users.
[0096] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation", and other variations thereof, mean that the particular features, structures, characteristics, etc. described in connection with that embodiment are included in at least one embodiment. Thus, the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and other variations that appear at various places throughout this application are not necessarily all referring to the same embodiment.
[0097] In addition, this application may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0098] Furthermore, this application may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0099] In addition, this application may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Furthermore, "receiving" typically involves in some way during operations such as storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0100] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", it should be understood that any use of the following, namely " / ", "and / or", and "at least one of", is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art in the relevant technical fields, this can be extended for as many listed items as there are.
[0101] Furthermore, as used herein, the term "signaling" means, among other things, indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals a quantization matrix for dequantization. Thus, in embodiments, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send a particular parameter to the decoder (explicit signaling) so that the decoder can use the same particular parameter. In contrast, if the decoder already has other parameters along with that particular parameter, signaling (implicit signaling) can be used that does not perform the transmission but simply enables the decoder to know and select that particular parameter. By avoiding the transmission of any actual functionality, bit savings are achieved in various embodiments. It will be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the term "signal", although the term "signal" may also be used as a noun herein.
[0102] As will be apparent to those of ordinary skill in the art, implementations can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. Signals can be transmitted over various different wired or wireless links, as is known. Signals can be stored on a processor-readable medium.
Claims
1. A method for video decoding, comprising: obtaining a first data set corresponding to a prediction block for a block of a picture; obtaining a second data set corresponding to a decoded prediction residual for the block of the picture; adjusting the second data set based on at least mixing parameters to form an adjusted second data set; and combining the first data set and the adjusted second data set to form a decoded version of the block of the picture.
2. A method for video encoding, comprising: obtaining a first data set corresponding to a prediction block for a block of a picture; obtaining a second data set corresponding to a reconstructed prediction residual for the block of the picture; adjusting the second data set based on at least mixing parameters to form an adjusted second data set; and combining the first data set and the adjusted second data set to form a reconstructed version of the block of the picture.
3. The method according to claim 1 or claim 2, wherein the second data set is scaled to form the adjusted second data set.
4. The method according to any one of claims 1 to 3, wherein the second data set is adjusted by at least an offset to form the adjusted second data set.
5. The method according to any one of claims 1 to 4, wherein the first data set is scaled when combined with the adjusted second data set.
6. The obtaining of the first data set includes inverse quantizing transform coefficients of the block to form inverse quantized transform coefficients, and inverse transforming the inverse quantized transform coefficients to reconstruct a prediction residual of the block, wherein the reconstructed prediction residual of the block is the second data set, and the prediction block is the first data set. The method according to any one of claims 1 to 5.
7. Said obtaining the first data set includes obtaining the prediction block of said block and converting the prediction block to form the first data set; said obtaining the second data set includes obtaining the transform coefficients of said block of said picture and inverse quantizing the transform coefficients of said block to form inverse quantized transform coefficients, and the inverse quantized transform coefficients are the second data set. The method according to any one of claims 1 to 5.
8. Said at least mixing parameter for adjusting the second data set is restricted to an interval between 1 - d and 1 + d, where d depends on a quantization parameter for inverse quantizing the transform coefficients of said block. The method according to any one of claims 1 to 7.
9. The method according to any one of claims 1 to 8, wherein the value of the mixing parameter for the prediction residual of the samples in said block depends on the value of the prediction for said samples.
10. The method according to any one of claims 1 to 8, wherein the same mixing parameter is applied to all prediction residuals of said block.
11. The method according to any one of claims 1 to 9, wherein the same mixing parameter is applied to blocks having the same quantization parameter.
12. The method according to any one of claims 1 and 3 to 11, further comprising decoding said at least mixing parameter for said block.
13. Said decoding comprises obtaining a plurality of mixing parameters, decoding the index of said block, and selecting, from said plurality of mixing parameters, one mixing parameter corresponding to said index for adjusting the second data set. The method according to claim 12.
14. In the transform domain, clipping the values of the samples in the reconstructed version of said block to be between a lower limit and an upper limit, where the lower limit and the upper limit are based on a parameter indicating the quantization step of said block and the sum of the inverse quantized transform coefficient of said sample and the transformed prediction. The method according to any one of claims 1 to 13.
15. An apparatus for video decoding comprising one or more processors and at least one memory, wherein said one or more processors Obtain a first dataset corresponding to a prediction block for a block of a picture, obtain a second dataset corresponding to the decoded prediction residual for the block of the picture, adjust the second dataset based on at least mixing parameters to form an adjusted second dataset, An apparatus for video decoding, configured to combine the first dataset and the adjusted second dataset to form a decoded version of the block of the picture.
16. An apparatus for video encoding, comprising one or more processors and at least one memory, wherein the one or more processors obtain a first dataset corresponding to a prediction block for a block of a picture, obtain a second dataset corresponding to the reconstructed prediction residual for the block of the picture, adjust the second dataset based on at least mixing parameters to form an adjusted second dataset, An apparatus for video encoding, configured to combine the first dataset and the adjusted second dataset to form a reconstructed version of the block of the picture.
17. The apparatus according to claim 15 or claim 16, wherein the second dataset is scaled to form the adjusted second dataset.
18. The apparatus according to any one of claims 15 to 17, wherein the second dataset is adjusted by at least an offset to form the adjusted second dataset.
19. The apparatus according to any one of claims 15 to 18, wherein the first dataset is scaled when combined with the adjusted second dataset.
20. The obtaining of the first dataset includes inverse quantizing the transform coefficients of the block to form inverse quantized transform coefficients, and inverse transforming the inverse quantized transform coefficients to reconstruct the prediction residual of the block, wherein the reconstructed prediction residual of the block is the second dataset, and the prediction block is the first dataset. The apparatus according to any one of claims 15 to 19.
21. Said obtaining the first data set includes obtaining the prediction block for said block and converting the prediction block to form the first data set. Said obtaining the second data set includes obtaining the conversion coefficients for said block of said picture and inverse quantizing the conversion coefficients for said block to form inverse quantized conversion coefficients, and the inverse quantized conversion coefficients are the second data set. The apparatus according to any one of claims 15 to 20.
22. The at least one mixing parameter for adjusting the second data set is restricted to an interval between 1 - d and 1 + d, where d depends on the quantization parameter for inverse quantizing the conversion coefficients of said block. The apparatus according to any one of claims 15 to 21.
23. The value of the mixing parameter for the prediction residual of the samples in said block depends on the value of the prediction for said samples. The apparatus according to any one of claims 15 to 22.
24. A signal including a bitstream formed by performing the method according to any one of claims 1 to 14.
25. A computer-readable storage medium storing instructions for encoding or decoding video according to the method according to any one of claims 1 to 14.