Method and apparatus for video encoding or decoding
By using a linear model of local lighting compensation (LIC) parameters in video encoding and decoding, the computational complexity is reduced and discontinuity is prevented, and the inefficiency and visual artifact problems caused by local lighting changes are solved, and more efficient video processing is achieved.
Patent Information
- Application Number
- CN202510189955.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-06
- Filing Date
- 2020-03-05
- Publication Date
- 2025-08-08
AI Technical Summary
The existing video encoding and decoding technologies are highly complex in computing when dealing with local lighting changes, resulting in inefficient efficiency and prone to visual artifacts.
A linear model of local lighting compensation (LIC) parameters is used to reduce the computational complexity through the derivation process, and a regularization process is used to prevent discontinuity problems, thereby improving the estimation accuracy.
It reduces the computational complexity of video encoding and decoding, improves encoding and decoding efficiency, and reduces the emergence of visual artifacts.
Smart Images

Figure CN120455658A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application date of March 5, 2020, application number 202080017644.5 (international application number PCT / US2020 / 021098), and invention name “Local illumination compensation for video encoding or decoding”. Technical Field
[0002] At least one of the present embodiments generally relates to local illumination compensation for video encoding or decoding. Background Art
[0003] To achieve high compression efficiency, image and video codecs typically employ prediction and transforms to exploit spatial and temporal redundancy in video content. Typically, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations. The difference between the original and predicted blocks (often expressed as a prediction error or prediction residual) is then transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded using the inverse of entropy coding, quantization, transform, and prediction. Summary of the Invention
[0004] One or more of the present embodiments provide a method and apparatus for encoding / decoding a picture using Local Illumination Compensation (LIC) parameters that are calculated using a derivation that provides reduced complexity and thus improved performance of LIC related processes.
[0005] According to a first aspect of at least one embodiment, a video encoding method includes: predicting picture data of at least one block in a picture, wherein the prediction includes performing motion compensation and local illumination compensation based on a reference block, and the local illumination compensation includes applying a linear model based on the sum of absolute differences of adjacent reconstructed samples of the reference block and corresponding reference samples, wherein the adjacent reconstructed samples and the corresponding reference samples of the reference block are co-located according to an L-shape that is substantially adjacent to a block to be predicted, the L-shape including a row of pixels located on a top side of the prediction block and a column of pixels located on a left side of the prediction block, and the co-location is determined based on a motion vector of the prediction block.
[0006] According to a second aspect of at least one embodiment, a video decoding method includes: predicting picture data of at least one block in a picture, wherein the prediction includes performing motion compensation and local illumination compensation based on a reference block, and the local illumination compensation includes applying a linear model based on the sum of absolute differences of adjacent reconstructed samples of the reference block and corresponding reference samples, wherein the adjacent reconstructed samples and corresponding reference samples of the reference block are co-located according to an L-shape that is substantially adjacent to the block to be predicted, and the L-shape includes a row of pixels located on the top side of the prediction block and a column of pixels located on the left side of the prediction block, and the co-location is determined based on the motion vector of the prediction block.
[0007] According to a third aspect of at least one embodiment, a device includes: an encoder for encoding picture data of at least one block in a picture or video, wherein the encoder is configured to predict the picture data of at least one block in the picture, wherein the prediction includes performing motion compensation and local illumination compensation based on a reference block, and the local illumination compensation includes applying a linear model based on the sum of absolute differences between adjacent reconstructed samples of the reference block and corresponding reference samples, wherein the adjacent reconstructed samples and the corresponding reference samples of the reference block are co-located according to an L-shape that is substantially adjacent to the block to be predicted, and the L-shape includes a row of pixels located on the top side of the prediction block and a column of pixels located on the left side of the prediction block, and the co-location is determined based on the motion vector of the prediction block.
[0008] According to a fourth aspect of at least one embodiment, a device includes: a decoder for decoding picture data of at least one block in a picture or video, wherein the decoder is configured to predict picture data of at least one block in the picture, wherein the prediction performs motion compensation and local illumination compensation based on a reference block, and the local illumination compensation includes applying a linear model based on the sum of absolute differences between adjacent reconstructed samples of the reference block and corresponding reference samples, wherein the adjacent reconstructed samples and the corresponding reference samples of the reference block are co-located according to an L-shape that is substantially adjacent to the block to be predicted, and the L-shape includes a row of pixels located on the top side of the prediction block and a column of pixels located on the left side of the prediction block, and the co-location is determined based on the motion vector of the prediction block.
[0009] According to variant aspects of the first, second, third and fourth aspects, the parameters of the linear model are calculated as follows:
[0010]
[0011] where cur(r) is the adjacent reconstructed sample in the current picture, ref(s) is the reference sample constructed with motion compensation converted by the motion vector mv from the reference picture, and s=r+mv.
[0012] According to variant aspects of the first, second, third and fourth aspects, the parameter "a" is derived with an additional simple regularization term "corr" and is determined in the following manner:
[0013]
[0014] in And for example reg_shift=7.
[0015] According to a fifth aspect of at least one embodiment, a computer program comprising program code instructions executable by a processor is proposed, which computer program implements the steps of the method according to at least the first or second aspect.
[0016] According to a sixth aspect of at least one embodiment, a computer program product is proposed, which is stored on a non-transitory computer-readable medium and includes program code instructions executable by a processor, and implements the steps of the method according to at least the first aspect or the second aspect.
[0017] According to an aspect of at least one embodiment, a method for decoding picture data of at least one block in a picture is proposed, the method comprising: determining a first set of LIC parameters for a first local illumination compensation (LIC) model based on a first set of LIC parameters in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters for a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, determining a predictor for the at least one block in response to the determining; and reconstructing the at least one block based on the predictor.
[0018] According to an aspect of at least one embodiment, an apparatus for decoding picture data of at least one block in a picture is provided, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: determine a first set of LIC parameters for a first local illumination compensation (LIC) model based on a first set of LIC parameters in an L-shape associated with the at least one block, and determine a second set of local illumination compensation parameters for a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, determine a predictor for the at least one block in response to the determining; and reconstruct the at least one block based on the predictor.
[0019] According to an aspect of at least one embodiment, a method for encoding picture data of at least one block in a picture is proposed, the method comprising: determining a first set of LIC parameters for a first local illumination compensation (LIC) model based on a first set of LIC parameters in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters for a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, determining a predictor for the at least one block in response to the determining; and encoding the at least one block based on the predictor.
[0020] According to an aspect of at least one embodiment, an apparatus for encoding picture data of at least one block in a picture is proposed, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to: determine a first set of LIC parameters for a first local illumination compensation (LIC) model based on a first set of LIC parameters in an L-shape associated with the at least one block, and determine a second set of local illumination compensation parameters for a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, determine a predictor for the at least one block in response to the determining; and encode the at least one block based on the predictor. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A block diagram of an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder, is shown.
[0022] Figure 2 A block diagram of an example of a video decoder 200 , such as a HEVC decoder, is shown.
[0023] Figure 3 An example of a codec tree unit and a codec tree in the compressed domain is shown.
[0024] Figure 4 An example of partitioning a CTU into coding units, prediction units, and transform units is shown.
[0025] Figure 5 An example of an "L-shape" for local illumination compensation is shown.
[0026] Figure 6 A block diagram illustrating an example of a system in which various aspects and embodiments may be implemented is shown.
[0027] Figure 7 An example embodiment of a multi-model LIC is shown in which the models are split by a threshold.
[0028] Figure 8A An example is shown in which a single model is determined based on extreme values.
[0029] Figure 8B Another example is shown in which a single model is determined based on the average extreme value.
[0030] Figure 9A The prediction method in the case of bidirectional prediction including the first method for deriving LIC parameters is shown.
[0031] Figure 9B The prediction method in the case of bidirectional prediction including a second method for deriving LIC parameters is shown.
[0032] Figure 10A An example of a multi-model discontinuity problem is shown.
[0033] Figure 10B A first example of a technique for solving the multi-model discontinuity problem is shown where the models are constructed as lines passing through the (minimum; maximum) points of each subset.
[0034] Figure 11 An example embodiment of a LIC including a regularization step is shown.
[0035] Figure 12 A regularization function is shown in which parameters are adjusted when they are outside a given range.
[0036] Figure 13 The temporal depth associated with the layered picture coding principle is shown.
[0037] Figure 14 An example embodiment of regularization for handling multi-model LIC is shown.
[0038] Figure 15A A method for determining the discontinuity between two LIC models is shown.
[0039] Figure 15B A method for calibrating LIC models that ensures multi-model continuity is presented in accordance with at least one embodiment.
[0040] Figure 16 An example of a block diagram showing multi-model LIC parameter correction.
[0041] Figure 17 An example embodiment for partitioning two LIC models is shown.
[0042] Figure 18 A method for deriving LIC parameters using two LIC models is shown.
[0043] Figure 19 An example embodiment for determining single model parameters when a discontinuity problem occurs is shown.
[0044] Figure 20 An example embodiment of a decision-making process is shown. DETAILED DESCRIPTION
[0045] In at least one embodiment, video encoding or decoding uses LIC and simplifies the calculation of LIC parameters by reducing its complexity, thereby improving video encoding or decoding performance.
[0046] In at least one embodiment, video encoding or decoding uses a LIC that uses at least one linear model whose parameters are determined using a derivation process that includes a regularization process to improve the estimation of local illumination variations and a correction process to prevent discontinuity problems when using multiple linear models. Thus, visual artifacts are prevented.
[0047] Figure 1 A block diagram of an example of a video encoder 100, such as a High Efficiency Video Coding (HEVC) encoder, is shown. Figure 1 An encoder in which the HEVC standard is improved or an encoder adopting a technique similar to HEVC, such as the JEM (Joint Exploration Model) encoder developed by JVET (Joint Video Exploration Team), may also be shown.
[0048] Before being encoded, a video sequence may undergo a pre-encoding process (101). This is performed, for example, by applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0) or performing a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.
[0049] In HEVC, to encode a video sequence having one or more pictures, the pictures are partitioned (102) into one or more slices, where each slice may include one or more slice segments. Slice segments are organized into codec units, prediction units, and transform units. The HEVC specification distinguishes between "blocks" and "units," where a "block" addresses a specific region in a sample array (e.g., luma, Y), and a "unit" includes all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data associated with the block (e.g., motion vectors).
[0050] For the codec in HEVC, the picture is partitioned into codec tree blocks (CTBs) of square shape with configurable size, and consecutive sets of codec tree blocks are grouped into slices. A codec tree unit (CTU) contains the CTBs of the coded color components. The CTB is the root of a quadtree partitioned into codec blocks (CBs), and a codec block can be partitioned into one or more prediction blocks (PBs) and forms the root of a quadtree partitioned into transform blocks (TBs). Corresponding to the codec blocks, prediction blocks and transform blocks, a codec unit (CU) includes a tree-structured set of prediction units (PUs) and transform units (TUs), the PU includes prediction information for all color components, and the TU includes a residual codec syntax structure for each color component. The sizes of the CB, PB and TB of the luminance component apply to the corresponding CU, PU and TU. In this application, the term "block" may be used to refer to any of, for example, CTU, CU, PU, TU, CB, PB and TB. Additionally, "block" may also be used to refer to macroblocks and partitions specified in H.264 / AVC or other video codec standards, and more generally to data arrays of various sizes.
[0051] In the example of encoder 100, a picture is encoded by an encoder element as described below. The picture to be encoded is processed in units of CUs. Each CU is encoded using intra or inter mode. When encoding a CU in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of intra mode or inter mode to use to encode the CU and indicates the intra / inter decision via a prediction mode flag. The prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0052] A CU in intra mode is predicted from reconstructed neighboring samples within the same slice. HEVC offers a set of 35 intra prediction modes, including DC, planar, and 33 angular prediction modes. The intra prediction reference is reconstructed from rows and columns adjacent to the current block. The reference extends over twice the block size in both horizontal and vertical directions, using available samples from previously reconstructed blocks. When an angular prediction mode is used for intra prediction, the reference samples may be copied along the direction indicated by the angular prediction mode.
[0053] Two different options can be used to encode the applicable luma intra prediction mode for the current block. If the applicable mode is included in a constructed list of three most probable modes (MPMs), the mode is signaled by its index in the MPM list. Otherwise, the mode is signaled by a fixed-length binarization of the mode index. The three most probable modes are derived from the intra prediction modes of the top and left neighboring blocks.
[0054] For inter CUs, the corresponding codec block is further divided into one or more prediction blocks. Inter prediction is performed at the PB level, and the corresponding PU contains information about how to perform inter prediction. Motion information (e.g., motion vector and reference picture index) can be signaled in two ways: "merge mode" and "advanced motion vector prediction (AMVP)".
[0055] In merge mode, the video encoder or decoder assembles a candidate list based on already coded blocks, and the video encoder signals the index of one of the candidates in the candidate list. At the decoder side, the motion vector (MV) and reference picture index are reconstructed based on the signaled candidate.
[0056] In AMVP, a video encoder or decoder assembles a candidate list based on motion vectors determined from already coded blocks. The video encoder then signals an index into the candidate list to identify the motion vector predictor (MVP) and signals the motion vector difference (MVD). On the decoder side, the motion vector (MV) is reconstructed as MVP + MVD. The applicable reference picture index is also explicitly coded in the PU syntax for AMVP.
[0057] The prediction residual is then transformed (125) and quantized (130), including at least one embodiment for adapting the chroma quantization parameters described below. The transform is typically based on a separable transform. For example, a DCT transform is applied first in the horizontal direction and then in the vertical direction. In recent codecs such as JEM, the transforms used in the two directions may be different (e.g., DCT in one direction and DST in the other), which results in a wide variety of 2D transforms, whereas in previous codecs the variety of 2D transforms available for a given block size was typically limited.
[0058] The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also bypass both the transform and quantization, i.e., encode and decode the residual directly without applying the transform or quantization process. In direct PCM encoding, no prediction is applied, and the codec unit samples are encoded and decoded directly into the bitstream.
[0059] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed picture, for example, to perform deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored in a reference picture buffer (180).
[0060] Figure 2 A block diagram of an example of a video decoder 200, such as an HEVC decoder, is shown. In the example of the decoder 200, a bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs the same operations as described above. Figure 1 The encoding pass described in
[0045] is reciprocated by the decoding pass, which performs video decoding as part of encoding the video data. Figure 2 Decoders in which improvements are made to the HEVC standard or decoders employing techniques similar to HEVC, such as the JEM decoder, may also be shown.
[0061] Specifically, the input to the decoder comprises a video bitstream that may be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, picture partitioning information, and other codec information. The picture partitioning information indicates the size of the CTU and the manner in which the CTU is partitioned into CUs and, where applicable, into PUs. The decoder may therefore partition (235) the picture into CTUs and each CTU into CUs based on the decoded picture partitioning information. The transform coefficients are dequantized (240), including for adapting at least one embodiment of the chroma quantization parameters described below, and inverse transformed (250) to decode the prediction residual.
[0062] The decoded prediction residual and the prediction block are combined (255) to reconstruct the image block. The prediction block (270) can be obtained from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275). As described above, AMVP and merge mode techniques can be used to derive motion vectors for motion compensation, which can use interpolation filters to calculate interpolated values of sub-integer samples of a reference block. An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored at a reference picture buffer (280).
[0063] The decoded picture may further undergo post-decoding processing (285), such as an inverse color transform (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.
[0064] Figure 3 An example of a codec tree unit and codec tree in the compressed domain is shown. In the HEVC video compression standard, a picture is partitioned into so-called codec tree units (CTUs), which are typically 64×64, 128×128, or 256×256 pixels in size. Each CTU is represented by a codec tree in the compressed domain. This is a quadtree of CTUs, where each leaf is called a codec unit (CU).
[0065] Figure 4 An example of partitioning a CTU into codec units, prediction units, and transform units is shown. Each CU is then given some intra or inter prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), each of which is assigned some prediction information. The intra or inter codec mode is assigned at the CU level.
[0066] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "encoding" or "codec" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side. The terms "block" or "picture block" can be used to refer to any of CTU, CU, PU, TU, CB, PB, and TB. In addition, the terms "block" or "picture block" can be used to refer to macroblocks, partitions, and subblocks specified in H.264 / AVC or other video codec standards, and more generally to sample arrays of many sizes.
[0067] Figure 5 An example of an "L-shape" for local illumination compensation is shown. In practice, emerging video compression tools as studied in the Joint Exploration Model (JEM) and the general video codec reference software developed by the JVET (Joint Video Exploration Team) group [2] use some additional temporal prediction tools such as Local Illumination Compensation (LIC) with related parameters determined at the decoder side.
[0068] Basically, the purpose of LIC is to compensate for illumination changes that may occur between a prediction block and its reference block used for temporal prediction via motion compensation. The use of LIC is usually signaled on CU level via a flag (LIC flag) associated with each codec unit (CU) coded in inter mode, or inferred from previously decoded blocks in case the current CU is coded in merge mode, for example. When this tool is activated, the decoder calculates some prediction parameters ( Figure 5 ). In the considered prior art codec (JEM), the use of LIC for a given block depends on a flag associated with this block, called the LIC flag.
[0069] In the following, we will refer to the “L-shape” associated with the current block as the set consisting of samples on rows above the current block and samples on columns to the left of the current block, as Figure 5 In variations, more than one row (or column) may be used.
[0070] The first implementation of LIC uses a LIC model based on a simple linear correction of Equation 1 applied to the regular current block prediction:
[0071] Ycorr(x)=a.Ypred(x)+b (Equation 1)
[0072] where Ypred(x) is the predicted sample value at position x, Ycorr(x) is the illumination-compensated predicted sample value at position x, and (a, b) are the LIC parameters.
[0073] The LIC parameters (a, b) are weights and biases based on minimizing the error between the current sample and the linearly modified reference sample, which are defined in Equation 2 as follows:
[0074] dist=∑ r∈Vcur,s∈Vref (cur(r)-a.ref(s)-b) 2 (Equation 2)
[0075] in:
[0076] cur(r) is the current image ( Figure 5 The adjacent reconstructed samples in the right side of
[0077] ref(s) is the image obtained from the reference image ( Figure 5 ), where s=r+mv, cur(r) and ref(r) are the co-located samples in the reconstructed L-shape and the reference L-shape, respectively.
[0078] The value of (a, b) is obtained using least squares minimization (LSM) as shown in the following equation 3:
[0079]
[0080] Note that the value of N can be further adjusted (incrementally reduced) to keep the sum term in Equation 3 below the maximum allowed integer storage value (e.g., sum term < 2 16 ). In addition, for large blocks, the subsampling of the top and left sample sets can be increased.
[0081] In case of additional conditions for selecting reconstructed samples, N may be equal to "numValid", or non-valid samples may be replaced with replicated valid samples.
[0082] The value of (a, b) can be obtained, for example, using a simpler calculation than least squares minimization, such as using extrema.
[0083] Once the encoder or decoder obtains the LIC parameters (a, b) of the current CU, the prediction of the current CU includes the following (unidirectional prediction case):
[0084] pred(current_block)=a×ref_block+b (Equation 4)
[0085] Where current_block is the current block to be predicted, pred(current_block) is the prediction of the current block, and ref_block is a reference block constructed using the conventional motion compensation (MC) process for temporal prediction of the current block. In a variant, in the case of bidirectional prediction, ref_block is a weighted sum of two reference blocks.
[0086] It should be noted that the adjacent reconstruction sample set and the reference sample set (see Figure 5 The gray samples in the image have the same number and pattern as the gray samples in the image. In the following, we refer to the neighboring reconstructed sample set (or reference sample set) located to the left of the current block as "left samples", and the neighboring reconstructed sample set (or reference sample set) located to the top of the current block as "top samples". We refer to the combination of the "left sample" and "top sample" sets as "sample set".
[0087] Table 1 provides an estimate of the complexity of the LIC parameter derivation according to Equation 3. Complexity is measured in this paper as the number of operations required to derive the LIC parameters. In this table, N = 2 kcorresponds to the number of reconstructed and reference samples with a bit depth equal to "d". The first column identifies the required operations, the second column measures the number of bits required in memory, the third to sixth columns count the number of required sum, multiplication, shift (divide by 2), and integer division operations respectively, and the last row provides the total number of required operations.
[0088]
[0089] Table 1
[0090] In at least a first embodiment, the LIC parameter calculation process uses the sum of absolute differences (SAD) as shown in Equation 6:
[0091]
[0092] Where cur(r) is the current image ( Figure 5 The adjacent reconstructed samples in the right side of the image are obtained by motion compensation from the reference image ( Figure 5 The adjacent reconstructed samples (cur(r)) and reference samples (ref(s)) of the current block are co-located with respect to the L-shape by the relationship “s=r+mv”, as Figure 5 shown.
[0093] Table 2 provides an estimate of the complexity of the LIC parameter derivation according to this first embodiment.
[0094]
[0095]
[0096] Table 2
[0097] The use of the sum of absolute differences for the LIC parameter calculation results in a significantly more efficient calculation. First, the memory requirements are reduced (column 2). The required additional additions are largely offset by a dramatic reduction in multiplications (column 4). Thus, this technique greatly improves the efficiency of LIC-related calculations and, more generally, improves the efficiency of encoding or decoding.
[0098] In a variant embodiment, the parameter a of Equation 6 is determined using the regularization term "corr" as shown in Equation 7:
[0099]
[0100] In at least one embodiment, the regularization term is defined as shown in Equation 8:
[0101]
[0102] Here, reg_shift takes a value of 7, for example.
[0103] Figure 6 A block diagram of an example of a system in which various aspects and embodiments are implemented is shown. System 1000 can be embodied as a device including the various components described below, and is configured to perform one or more aspects of the various aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, encoders, transcoders, and servers. The elements of system 1000 can be embodied individually or in combination in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects of the various aspects described in this document.
[0104] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive, and / or optical disk drive. As non-limiting examples, the storage device 1040 may include an internal storage device, an attached storage device, and / or a network accessible storage device.
[0105] System 1000 includes an encoder / decoder module 1030, which is configured to process data to provide encoded video or decoded video, for example, and may include its own processor and memory. Encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both encoding and decoding modules. In addition, encoder / decoder module 1030 may be implemented as a separate element of system 1000, or may be incorporated into processor 1010 as a combination of hardware and software known to those skilled in the art.
[0106] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0107] In several embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, such as dynamic volatile memory and / or non-volatile memory. In several embodiments, external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, fast external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as for MPEG-2, HEVC, or VVC (Versatile Video Codec).
[0108] Input to the elements of system 1000 may be provided through various input devices, as indicated at block 1130. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, for example, by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0109] In various embodiments, the input device of block 1130 has associated corresponding input processing elements as are known in the art. For example, the RF section may be associated with the elements required to: (i) select a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band), (ii) downconvert the selected signal, (iii) again band-limit to a narrower frequency band to select a signal frequency band (which may be referred to as a channel in some embodiments), (iv) demodulate the downconverted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select the desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted on wired (for example, cable) medium, and by filtering, down-conversion and filtering to the desired frequency band again to perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some in these elements, and / or add other elements of execution similar or different functions.Adding element can include and insert element between existing element, such as, for example, insert amplifier and analog to digital converter.In various embodiments, the RF part comprises antenna.
[0110] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within the processor 1010, as desired. Similarly, various aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within the processor 1010, as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in combination with memory and storage elements to process the data streams as needed for presentation on an output device.
[0111] The various elements of system 1000 may be disposed within an integrated housing in which the various elements may be interconnected and data transferred therebetween using a suitable connection arrangement (eg, an internal bus known in the art, including an I2C bus, wiring, and printed circuit boards).
[0112] System 1000 includes a communication interface 1050 capable of communicating with other devices via a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data through communication channel 1060. Communication interface 1050 may include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0113] In various embodiments, data is streamed to the system 1000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received via a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that delivers data via an HDMI connection of the input block 1130 to provide streaming data to the system 1000. Other embodiments use an RF connection of the input block 1130 to provide streaming data to the system 1000.
[0114] System 1000 can provide output signals to various output devices, including display 1100, speakers 1110, and other peripherals 1120. In various examples of embodiments, other peripherals 1120 include one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripherals 1120 using signaling using protocols such as AV.Link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections via respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 via communication interface 1050 using communication channel 1060. Display 1100 and speakers 1110 can be integrated into a single unit with other components of system 1000 in an electronic device such as, for example, a television. In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (TCon) chip.
[0115] For example, if the RF portion of input 1130 is part of a separate set-top box, the display 1100 and speaker 1110 can alternatively be separated from one or more of the other components. In various embodiments in which the display 1100 and speaker 1110 are external components, an output signal can be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output). The implementation described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of an implementation (for example, discussed only as a method), the implementation of the features discussed can also be implemented in other forms (for example, a device or program). The device can be implemented with, for example, appropriate hardware, software, and firmware. The method can be implemented in, for example, a device such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as, for example, a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.
[0116] Figure 7 An example embodiment of a multi-model LIC in which the models are divided by a threshold is shown. In such an embodiment, the two linear models use different LIC parameters ((a0, b0) and (a1, b1)), such as Figure 7 As shown, and determined, for example, using Equation 3 above or using other methods. The first model and first LIC parameters are used for a first subset of samples, and the second model and second LIC parameters are used for a second subset of samples. Hereinafter, this mode is referred to as multi-model LIC (MM-LIC). This embodiment uses two linear models (a0, b0) and (a1, b1) and a threshold value on the reconstructed luminance sample values to determine which model to use. The threshold value is determined, for example, as the average value of the reconstructed luminance sample values.
[0117] Figure 8A An example is shown in which a single model is determined based on extreme values. In this example, two reference samples (XA, XB) with a minimum value (Min) and a maximum value (Max) and associated reconstructed samples (YA, YB) are used to obtain the value of "a" (Equation 3b):
[0118]
[0119] Figure 8BAnother example is shown in which a single model is determined based on the average extreme value. In this example, the minimum and maximum values are replaced by the average of two or more minimum values or two or more maximum values, respectively. In the example shown in the figure, Min1 and Min2 are two minimum values. These elements are averaged to determine the values of XA and YA. Max1 and Max2 are two maximum values. These elements are averaged to determine the values of XB and YB. The value of "a" is then determined based on these corrected values of XA, XB, YA, and YB. This allows for increased robustness and is particularly effective in the case of outliers (e.g., due to noise). In the example shown, Max1 is clearly a contour line, and using only the values of Min1 and Max1 to determine the linear model will result in a model that is not very accurate.
[0120] Figure 9A The prediction method is shown in the case of bidirectional prediction, which includes the first method for deriving LIC parameters. In this case, the prediction is based on two references, called reference 0 and reference 1, combined together. The LIC process is applied twice, first to the reference 0 prediction (LIC-0) and second to the reference 1 prediction (LIC_1). The two predictions are then combined as usual using either the default weighting (P = (P0 + P1 + 1) >> 1) or the bidirectional prediction weighted average (BPWA): P = (g0.P0 + g1.P1 + (1 << (s-1))) >> s).
[0121] Figure 9B A prediction method is shown in the case of bidirectional prediction including a second method for deriving LIC parameters. In this variant, referred to as method B, conventional predictions are first combined in step 360. Then, a single LIC process is applied in step 330, and a prediction is calculated in step 340 based on the determined LIC parameters.
[0122] Another method for parameter estimation for intra prediction is named Cross Component Linear Model (CCLM). The prediction of the chroma sample block (Ycorr) can be constructed from the (scaled down) reconstructed luma sample block (Ypred) corrected with the linear model. In this case, the same model as Equation 1 is used, where:
[0123] Ypred(x) is the reconstructed luminance sample value at position x,
[0124] Ycorr(x) is the predicted value of the chrominance sample at position x.
[0125] The value of (a, b) is estimated based on the minimum and maximum values (Minluma, Maxluma) of the reconstructed luma samples and the minimum and maximum values Ycurr (Minchroma, Maxchroma) of the reconstructed chroma samples of the L-shape, as shown in Equation 5:
[0126] a=((〖Max〗_chroma-〖Min〗_chroma)) / ((〖Max〗_luma-〖Min〗_luma))
[0127] b = [Min]_chroma - a. [Min]_luma (Equation 5)
[0128] like Figure 7 As shown, in order to improve LIC, two different LIC models can be used, with e.g. Figure 8A 8B . The first model and the first LIC parameters are used for a first subset of samples, and the second model and the second LIC parameters are used for a second subset of samples. The division between the two sets is performed with respect to a threshold.
[0129] Figure 10A An example of the multi-model discontinuity problem is shown. This is a case of model discontinuity, in other words, when two lines generated by two models ((a0; b0) and (a1; b1)) do not intersect at a threshold. This discontinuity problem can cause visual artifacts and / or reduce coding efficiency and should be prevented.
[0130] Figure 10B A first example of a technique for resolving the multi-model discontinuity problem is shown, where the models are constructed as lines passing through the (minimum; maximum) points of each subset. Correction is accomplished by replacing Max0 and Min1 with the average value between the previous values (point T in the figure) and using these corrected values to determine the LIC parameters for both models. This prevents discontinuity because the maximum value of the first model is equal to the minimum value of the second model (equal to T in the figure).
[0131] At least one issue facing LIC is that the derivation of LIC parameters is based on a relatively small number of samples compared to the current block size. Furthermore, these samples are not co-located with the current block, making them potentially poor estimates of local illumination variations. Furthermore, LIC parameter estimation also depends on the method used to derive (a, b). For example, in the case of Equation 3, if the denominator is low, some inconsistency in the value of "a" may occur. In another example, if the numerator is low, inconsistency in the value of "a" may occur. This is typically the case, for example, when the distribution of ref(s) or cur(r) is narrow. Therefore, in some cases, the derivation of LIC parameters includes some uncertainty. To reduce this uncertainty, a regularization process is inserted into the prediction stage so that potential problems are detected and the LIC parameters are corrected accordingly to prevent visual artifacts and / or improve coding efficiency.
[0132] Figure 11 An example embodiment of a LIC including a regularization step is shown. In such an embodiment, the overall principles including the LIC parameter derivation (330) remain unchanged, but once the LIC parameters have been calculated, a regularization function is applied in step 350 to potentially correct the previously determined LIC parameters. Such regularization or correction is particularly desirable when the derivation of the LIC parameters is uncertain, in other words, when the confidence level of the determined LIC parameters is too low.
[0133] Figure 12 A regularization function is shown in which parameters are adjusted when they are outside a given range. In this variant embodiment, the values (a, b) are adjusted if "a" and / or "b" exceed some predefined thresholds (th_a, th_b) around the default (1, 0) values. For example: for a 10-bit sample, th_a = 0.25 and th_b = 100. In at least one embodiment, the predefined thresholds th_a and th_b are evenly distributed around the default value equal to 1.0.
[0134] In at least one embodiment:
[0135] If (a<1-th_a) or (a>1+th_a), adjust "a".
[0136] If (b<-th_b) or (b>th_b), adjust "b".
[0137] In a variation, the thresholds are unevenly distributed around the default (1,0) value:
[0138] If (a<1-th_a1) or (a>1+th_a2), adjust "a".
[0139] If (b<-th_b1) or (b>th_b2), adjust "b".
[0140] In at least one embodiment, the regularization function provides for adjusting the parameter when the parameter is outside a given range. In such an embodiment, parameter a is adjusted first, and b is determined according to Equation 1. In fact, with respect to Equation 1, the derivation of the LIC parameter is based on the property r1:
[0141] DCrec=a.DCref+b(r1)
[0142] Where DCrec is the average value of the reconstructed samples and DCref is the average value of the reference samples. Therefore, in this embodiment:
[0143] If ((b>th_b)&&(a<1-th_a)), then
[0144] "a" is set to "1-th_a", and
[0145] Calculate "b" using (r1).
[0146] If ((b<-th_b)&&(a>1+th_a)), then
[0147] "a" is set to "1+th_a", and
[0148] "b" is calculated using (r1), and thus b = DCrec - a.DCref.
[0149] In at least another embodiment, the property (r1) is also verified using the adjusted value (a+da, b+db):
[0150] DCrec=(a+da).DCref+(b+db)(r2)
[0151] Given (r1), (r2) becomes:
[0152] Da=-db / DCref(r3)
[0153] According to this embodiment, a is adjusted to (a+da) and b is adjusted to (b+db), the adjustment "da" is calculated explicitly with (r3), and "db" is determined as:
[0154] If (b<-th_b), then db = -th_b-b
[0155] If (b>th_b), then db=th_b-b
[0156] These regularization adjustments ensure that the LIC parameter conforms to a certain range of values and, therefore, the prediction step behaves correctly under good conditions.
[0157] In at least one embodiment, the values of "th_a" and / or "th_b" are encoded in the bitstream. For example, they can be encoded in a sequence header, a picture header, a slice header, or a title header. In another embodiment, they are fixed or associated with a profile and / or level.
[0158] In at least another embodiment, the values of "th_a" and / or "th_b" are functions of at least a parameter (or set of parameters) P and can be obtained from the bitstream using a decoding process. For example, P may include the picture order count (POC) distance between the current picture (POCcur) and the reference picture (POCref).
[0159] For example: th_a = 0.25x(1+0.25x abs(POCcur–POCref))
[0160] Or: th_a = 0.25 x (1 + 0.25 x min(4; abs(POCcur – POCref)))
[0161] In another example, parameter P includes the temporal depth of the current picture. Figure 13 The temporal depth related to the principle of hierarchical picture coding and decoding is shown.
[0162] In at least one embodiment, parameter P includes the number of samples of the L shape actually used for estimating the LIC model (e.g., numValid). In this case, the values of "th_a" and / or "th_b" are functions of P.
[0163] For example: if (numValid < Nc), then
[0164] th_a = th_a1 and th_b = th_b1
[0165] Otherwise
[0166] th_a = th_a2 and th_b = th_b2
[0167] Where, for example, th_a1 = 0.3, th_a2 = 0.4, th_b1 = 0, th_b2 = 80, and Nc is a confidence threshold. This threshold may be different in terms of luminance and chrominance due to different numbers of samples. In one example, the threshold 32 is used for luminance and the threshold 16 is used for chrominance. In a variant, NC is a function of the current block size, e.g., NC = 0.5 x (blockWidth + blockHeight).
[0168] In this embodiment, a low value of numValid indicates a reduced number of samples and thus implies a low confidence in the validity of the LIC model.
[0169] In the case of multi-model LIC, multiple regularization processes are required: one regularization process for each model ( Figure 14 element 350 in), because each individual model of the multi-model LIC is independent of other individual models.
[0170] Figure 14 An example embodiment for processing the regularization of multi-model LIC is shown. As introduced above, two different LIC models can be used and can lead to discontinuity problems. In step 380, a correction process is added to ensure the continuity of multiple models. This process can be implemented directly after the regularization process.
[0171] In at least one embodiment, the correction process (380) is applied only in merge mode. In practice, when the encoder detects a multi-model discontinuity problem, it can disable LIC for a block. This is performed by encoding the LIC flag as false for this block. However, in merge mode, the LIC flag is inherited from another neighboring block, and the encoder has less flexibility to avoid this problem unless it recursively re-encodes the previous block.
[0172] In at least one embodiment, if the discontinuity size is above a discontinuity threshold (DT), a correction process (380) is applied.
[0173] In at least one embodiment, if the discontinuity size is above DT, the correction process (380) is applied only in merge mode.
[0174] The discontinuity size (DS) can be calculated by the decoder in different ways:
[0175] as the distance between the values at the intersection between each linear model and the threshold line, e.g., Figure 15A shown.
[0176] As the difference between the maximum value of the first model M0 and the minimum value of the second model M1. DS = max0 - min1
[0177] The value of the discontinuity threshold DT may be implicitly known by the decoder, encoded in the bitstream, calculated from other decoded parameters, or obtained by other means.
[0178] In at least one embodiment and as Figure 15B As shown, the correction step 380 of FIG8 operates as follows:
[0179] For Yref="threshold", C0 and C1 are calculated as the values given by models M0 and M1 respectively,
[0180] Calculate T as the average between these two values: T = (C0 + C1) / 2,
[0181] M0 and M1 are corrected when the line passes through (Avg0; T) and (T; Avg1), where Avg0 and Avg1 are the averages of the sample values below and above the threshold, respectively.
[0182] In a variant embodiment of the previous embodiment, M0 and M1 are corrected to lines passing through (Min0; T) and (T; Max1), where Min0 is the minimum value of the samples associated with the first model M0 and Max1 is the maximum value of the samples associated with the second model M1.
[0183] Figure 16An example of a block diagram for multi-model LIC parameter correction is shown. This element corresponds to block 380 of Figure 8. First, in step 381, the discontinuity size DS between the two models is calculated. Then, in step 382, this discontinuity size DS is compared with a discontinuity threshold DT. When DS>DT (branch "yes"), a discontinuity problem has been detected. In this case, a third step 383 includes correcting the LIC parameters of the two models using one of the above-described embodiments. When no discontinuity is detected (branch "no"), the LIC parameters are not corrected. Then, in step 340, a prediction is performed using the LIC parameters of the two models.
[0184] Figure 17 An example embodiment for partitioning two LIC models is shown. In such an embodiment, the two linear models use different LIC parameters ((a0, b0) and (a1, b1)), such as Figure 7 As shown, and determined, for example, using Equation 3 or using other methods. The first model and first LIC parameters are used for a first subset of samples, and the second model and second LIC parameters are used for a second subset of samples. Hereinafter, this scheme is referred to as multi-model LIC (MM-LIC).
[0185] The division between the two sets is done with respect to a threshold. Figure 6 In , NM0 and NM1 represent the number of reconstructed neighboring samples with values below and above the threshold, respectively.
[0186] Determining the threshold for dividing the two models may be performed according to various embodiments.
[0187] In at least one embodiment, the threshold for dividing the two models is determined, for example, as the average value of the reconstructed samples ("cur(s)").
[0188] In at least one embodiment, the threshold is determined so that the number of samples used in each model (NM0 and NM1) is substantially the same. This increases the effectiveness of the LIC model. In fact, when the number of samples of one of the models is too small (NM0 is much larger than NM1 or NM0 is much smaller than NM1), the effectiveness of the corresponding LIC model (NM1 or NM0, respectively) is uncertain, and the LIC parameters derived using Equation 3 may be unreliable because they are based on too few samples. This can be achieved using Figure 6 The histogram shown is done by counting the number of samples. In a variant, only if the size of the histogram (2 比特-深度-s (2 bit-depth–s This embodiment should only be used when )) is less than N=NM0+NM1.
[0189] In an alternative embodiment, a single LIC model is used when the number of samples of the models is not well balanced. This can be determined by dividing the number of samples of the model with the highest number of samples by the number of samples of the other model. If the ratio is greater than a threshold, the models are considered not well balanced, and in this case, a single linear model is used. An example of a threshold is 10. In another example, this can be determined if one of the number of samples is below a predetermined value (e.g., NMi = 4 samples).
[0190] In at least one embodiment, by using a value less than the reconstruction sample range (e.g., 0...2 比特-深度 ) to achieve memory savings. This is done by using an appropriate scaling factor "s" corresponding to, for example, a right shift of the sample values. In this case, the histogram buffer size is reduced to 2 比特-深度-s , and the histogram step size is 2 s , thus saving memory.
[0191] In a variant embodiment, the lowest and highest values of the histogram are saturated. This is illustrated by the values MinRange and MaxRange in the figure. All sample values below the MinRange are set to the MinRange, and all sample values above the MaxRange are set to the MaxRange. When the potential range of thresholds is known, only a rough distribution around the possible thresholds is required, allowing some memory savings. For example, an estimate of the sample value range can be inferred from previously coded samples in the same or previously coded pictures, or the range value can be encoded in the bitstream.
[0192] In at least one embodiment, a single pass over the data is performed by aggregating the partial sums over a histogram: is calculated at the same intervals as the values of the histogram. Once the values such as The value of and the threshold are chosen, and the partial sums are accumulated using the histogram values so that there is no need to loop over the ref and cur values again. This corresponds to an approximation of the sum of absolute differences, namely:
[0193]
[0194] Among them, n cur (h) is the number of reconstructed cur samples with values ∈ [h2s; (h+1).2s], and H = {h0, h1...} is the number of histogram values. The same approximation can be applied to the reference samples. This single pass is particularly interesting when the LIC parameters a and b are obtained not using least squares minimization (LSM) as shown in Equation 3, but by using the sum of absolute differences (SAD) as shown in Equation 6:
[0195]
[0196] Where cur(r) is the current image ( Figure 5 ), ref(s) is the adjacent reconstructed sample from the reference image ( Figure 5 The left side of the block is a reference sample constructed by motion compensation (converted by motion vector mv), and s = r + mv. The adjacent reconstructed samples (cur (r)) and the reference samples (ref (s)) of the current block are co-located with respect to the L shape by the relationship "s = r + mv", as shown in Figure 5 shown.
[0197] This embodiment is also applicable when least squares minimization (LSM) is used to obtain the LIC parameters a and b. The sum terms of the two models are simply added to obtain the sum term of the single model.
[0198] Figure 18 A method for deriving LIC parameters using two LIC models is shown. In this method, once a threshold is determined from the histogram at step 370, the value of the reconstructed sample is compared to the threshold at step 330 to determine whether the reconstructed sample will be used to determine the LIC parameters for the first LIC model or the second LIC model. Then, at step 340, corresponding LIC parameters are determined for both models, and predictions are calculated using both models. These predictions are then combined (at step 360), for example using a simple average or a bi-prediction weighted average.
[0199] In at least one embodiment, the multi-model discontinuity problem is detected by the encoder, and the LIC feature can be disabled by the encoder for the block where the multi-model discontinuity occurs. This can be signaled by the encoder by encoding the LIC flag as false for that block. However, in merge mode, the LIC flag is inherited from another neighboring block, and the encoder has less flexibility to avoid this problem unless it recursively re-encodes the previous block.
[0200] In at least one embodiment, in case of multi-model discontinuity for a block coded in merge mode, the encoder does not use MM-LIC but uses a single LIC model.
[0201] In a variant embodiment, the discontinuity size (DS) is calculated by the decoder, and if the DS is above a discontinuity threshold, a single LIC model approach is selected, otherwise a multi-model LIC model is used. The discontinuity threshold is implicitly known to the decoder, either encoded in the bitstream, calculated from other decoding parameters, or obtained using other means.
[0202] Figure 19An example embodiment for determining the parameters of a single model when a discontinuity problem occurs is shown. In this embodiment, which corresponds to the case where a single LIC model approach is selected due to a discontinuity problem, the LIC parameters of the single model are derived directly from the average values (Avg0 and Avg1, respectively) of the samples below and above the threshold that separates the two models. This has the following advantages: the parameters of the single model are efficiently derived in a simple manner without the need to rescan the samples, and Equation 6 does not need to be recalculated because the average values Avg0 and Avg1 have already been calculated.
[0203] Figure 20 An example embodiment of a decision process according to one embodiment is shown. The process is intended to decide whether LIC should use a single linear model or multiple linear models. First, the LIC parameters are determined for the case where two models are used. Then, in step 335, a decision is made to decide whether LIC should be used and which model to use. The decision is made based on the different factors mentioned above (unbalanced number of samples, discontinuity issues). When a single model should be used, then in step 336, for example, a single linear model is used. Figure 9B The method shown determines the LIC parameters for a single model and performs block prediction at step 342. When the decision is to select multi-model LIC, then block prediction is performed at step 340 using the LIC parameters calculated for both models.
[0204] The decision of whether to use a single linear model or multiple linear models for LIC is valid for both the encoder and decoder, but the decision not to use LIC is only valid on the encoder side.
[0205] Reference to "one embodiment" or "an embodiment" or "an implementation" or "an implementation" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations thereof in various places throughout this specification are not necessarily all referring to the same embodiment. Additionally, the application or its claims may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory. Additionally, the application or its claims may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.
[0206] Additionally, this application or its claims may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information can include, for example, one or more of accessing information or retrieving information (e.g., from a memory or optical storage medium). Furthermore, during operations such as, for example, storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information, "receiving" is often involved in one way or another.
[0207] It should be appreciated that, for example, in the case of "A / B," "A and / or B," and "at least one of A and B," use of any of the following " / ," "and / or," and "at least one of" is intended to encompass selecting only the first listed option (A), or only the second listed option (B), or both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C," such phrases are intended to encompass selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended to as many of the listed items as will be apparent to one of ordinary skill in this and related arts.
[0208] It will be apparent to those skilled in the art that an implementation may generate various signals formatted to carry information that may, for example, be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave using the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted via a variety of different wired or wireless links. The signal may be stored on a processor-readable medium.
Claims
1. A method for decoding picture data of at least one block in a picture, the method comprising: determining a first set of local illumination compensation (LIC) parameters of a first LIC model based on a first set of reconstructed samples in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters of a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, In response to the determining, determining a predictor for the at least one block; and The at least one block is reconstructed based on the predictor.
2. The method according to claim 1, wherein In response to the determining, determining the predictor for the at least one block includes disabling local illumination compensation if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a second value.
3. The method of claim 1, wherein In response to the determining, determining the predictor for the at least one block includes determining the predictor based on a single set of local illumination compensation parameters if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a third value, or if a ratio between the number of samples of a reconstructed sample set having a highest number of samples and the number of samples of another reconstructed sample set is greater than a fourth value.
4. The method according to claim 3, wherein: Local illumination compensation parameters for the single set are determined based on an average of the first set of reconstruction samples and an average of the second set of reconstruction samples.
5. The method according to claim 1, wherein In a case where it is determined that the discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is greater than a second value, the method includes correcting the first set of local illumination compensation parameters and the second set of local illumination compensation parameters based on a value obtained for the first value by the first LIC model and a value obtained for the first value by the second LIC model.
6. The method according to claim 1, wherein The first value is equal to an average value of the reconstructed samples in the L-shape.
7. The method according to claim 1, wherein The first value is determined such that the number of reconstructed samples in the first set and the number of reconstructed samples in the second set are substantially the same.
8. An apparatus for decoding picture data of at least one block in a picture, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform: determining a first set of local illumination compensation (LIC) parameters of a first LIC model based on a first set of reconstructed samples in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters of a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, In response to the determining, determining a predictor for the at least one block; and The at least one block is reconstructed based on the predictor.
9. The device according to claim 8, wherein In response to the determining, determining the predictor for the at least one block includes disabling local illumination compensation if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a second value.
10. The apparatus of claim 8, wherein In response to the determining, determining the predictor for the at least one block includes determining the predictor based on a single set of local illumination compensation parameters if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a third value, or if a ratio between the number of samples of a reconstructed sample set having a highest number of samples and the number of samples of another reconstructed sample set is greater than a fourth value.
11. The device according to claim 10, wherein Local illumination compensation parameters for the single set are determined based on an average of the first set of reconstruction samples and an average of the second set of reconstruction samples.
12. The device according to claim 8, wherein Upon determining that the discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is greater than a second value, the one or more processors are configured to perform correction of the first set of local illumination compensation parameters and the second set of local illumination compensation parameters based on a value obtained by the first LIC model for the first value and a value obtained by the second LIC model for the first value.
13. The device according to claim 8, wherein The first value is equal to an average value of the reconstructed samples in the L-shape.
14. The device according to claim 8, wherein The first value is determined such that the number of reconstructed samples in the first set and the number of reconstructed samples in the second set are substantially the same.
15. A method for encoding picture data of at least one block in a picture, the method comprising: determining a first set of local illumination compensation (LIC) parameters of a first LIC model based on a first set of reconstructed samples in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters of a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, In response to the determining, determining a predictor for the at least one block; and The at least one block is encoded based on the predictor.
16. The method according to claim 15, wherein In response to the determining, determining the predictor for the at least one block includes disabling local illumination compensation if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a second value.
17. The method of claim 15, wherein In response to the determining, determining the predictor for the at least one block includes determining the predictor based on a single set of local illumination compensation parameters if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a third value, or if a ratio between the number of samples of a reconstructed sample set having a highest number of samples and the number of samples of another reconstructed sample set is greater than a fourth value.
18. The method according to claim 17, wherein: Local illumination compensation parameters for the single set are determined based on an average of the first set of reconstruction samples and an average of the second set of reconstruction samples.
19. The method according to claim 15, wherein In a case where it is determined that the discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is greater than a second value, the method includes correcting the first set of local illumination compensation parameters and the second set of local illumination compensation parameters based on a value obtained for the first value by the first LIC model and a value obtained for the first value by the second LIC model.
20. The method according to claim 15, wherein The first value is equal to an average value of the reconstructed samples in the L-shape.
21. The method according to claim 15, wherein The first value is determined such that the number of reconstructed samples in the first set and the number of reconstructed samples in the second set are substantially the same.
22. An apparatus for encoding picture data of at least one block in a picture, the apparatus comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform: determining a first set of local illumination compensation (LIC) parameters of a first LIC model based on a first set of reconstructed samples in an L-shape associated with the at least one block, and determining a second set of local illumination compensation parameters of a second LIC model based on a second set of reconstructed samples in the L-shape, wherein the first set of reconstructed samples includes reconstructed samples below a first value and the second set of reconstructed samples includes reconstructed samples above the first value, In response to the determining, determining a predictor for the at least one block; and The at least one block is encoded based on the predictor.
23. The device according to claim 22, wherein In response to the determining, determining the predictor for the at least one block includes disabling local illumination compensation if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a second value.
24. The apparatus of claim 22, wherein In response to the determining, determining the predictor for the at least one block includes determining the predictor based on a single set of local illumination compensation parameters if a discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is determined to be greater than a third value, or if a ratio between the number of samples of a reconstructed sample set having a highest number of samples and the number of samples of another reconstructed sample set is greater than a fourth value.
25. The apparatus according to claim 24, wherein Local illumination compensation parameters for the single set are determined based on an average of the first set of reconstruction samples and an average of the second set of reconstruction samples.
26. The apparatus of claim 22, wherein: Upon determining that the discontinuity between the first set of local illumination compensation parameters and the second set of local illumination compensation parameters is greater than a second value, the one or more processors are configured to perform correction of the first set of local illumination compensation parameters and the second set of local illumination compensation parameters based on a value obtained by the first LIC model for the first value and a value obtained by the second LIC model for the first value.
27. The apparatus of claim 22, wherein: The first value is equal to an average value of the reconstructed samples in the L-shape.
28. The apparatus according to claim 22, wherein The first value is determined such that the number of reconstructed samples in the first set and the number of reconstructed samples in the second set are substantially the same.