Video encoding and decoding methods and apparatuses using sub-block-based local illumination compensation

By determining the linear model parameters of local lighting compensation based on spatial proximity samples in video encoding and decoding, and dividing the block into sub-blocks in parallel processing, the local lighting compensation sample availability problem in the prior art is solved, and the efficiency of video encoding and decoding is improved.

CN113557730BActive Publication Date: 2025-06-17INTERDIGITAL VC HOLDINGS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080020338.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-03-12
Filing Date
2020-03-05
Publication Date
2025-06-17
Estimated Expiration
2040-03-05

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have sample availability problems in local lighting compensation, resulting in a reduced efficiency of pipeline decoding process.

Method used

Linear model parameters for local lighting compensation are determined based on the spatial proximity reconstruction sample and the corresponding reference sample for blocks in the video picture, and the blocks are divided into parallel processed sub-blocks for motion compensation.

Benefits of technology

The local lighting compensation efficiency during video encoding and decoding is improved, the pipeline decoding process is optimized, and the lighting compensation ability for blocks is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113557730B_ABST
    Figure CN113557730B_ABST
Patent Text Reader

Abstract

Describes different implementations, particularly an implementation of video encoding and decoding based on a linear model responsive to neighboring samples. Thus, for a block being encoded or decoded in a picture, refined linear model parameters are determined for a current sub-block in the block, and for encoding the block, local illumination compensation uses a linear model for the current sub-block based on the refined linear model. In a first embodiment, the number N of reconstructed samples increases with the available data for the sub-block. In a second embodiment, partial linear model parameters are determined for the sub-block, and refined linear model parameters are derived from a weighted sum of the partial linear model parameters. In a third embodiment, the sub-blocks are processed independently by LIC.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one in the present embodiment generally relates to, for example, a method or apparatus for video encoding or decoding, and more particularly, to a method or apparatus for the following operations: for a block being encoded or decoded, determining linear model parameters for local illumination compensation based on neighboring samples; and the block is divided into sub-blocks for parallel processing for motion compensation. Background Art

[0002] The field of technology of one or more implementations generally relates to video compression. At least some embodiments relate to improving compression efficiency compared to the following systems: existing video compression systems such as HEVC (HEVC refers to High Efficiency Video Coding, also known as "ITU-T H.265 of the ITU's Telecommunication Standardization Sector (10 / 2014), Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services - Coding of Moving Video, High Efficiency Video Coding, Recommendation ITU-T H.265" described in H.265 and MPEG-H Part 2), or video compression systems under development (such as VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team JVET)).

[0003] To achieve high compression efficiency, image and video coding and decoding schemes typically employ prediction (including motion vector prediction) and transformation to exploit spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image and the predicted image (usually represented as prediction error or prediction residual) is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded through an inverse process corresponding to entropy coding, quantization, transformation, and prediction.

[0004] Recent additions to the high compression technology include a prediction model based on linear modeling of the neighborhood of the block whose response is being processed. In particular, during the decoding process, some prediction parameters are calculated based on samples located in the spatial neighborhood of the block being processed. This spatial neighborhood includes the reconstructed picture samples and the corresponding samples in the reference pictures. Such a prediction model with prediction parameters determined based on the spatial neighborhood is implemented in Local Illumination Compensation (LIC). Additionally, other methods of the high compression technology include new tools for motion compensation, such as affine motion compensation, sub-block based Temporal Motion Vector Prediction (sbTMVP), Bi-Directional Optical Flow (BDOF), and Decoder-Side Motion Vector Refinement (DMVR). Some of these tools require processing blocks in multiple sub-blocks in several consecutive operations. The tools are applied in sequence during the decoding process. To cope with the constraints of real-time decoding, the decoding process is pipelined so as to process blocks and sub-blocks in parallel. This pipelined decoding process raises questions regarding the availability of samples in the spatial neighborhood of the blocks used in LIC. Therefore, it is necessary to optimize the decoding pipeline for local illumination compensation. Summary of the Invention

[0005] An object of the present invention is to overcome at least one drawback of the prior art. To this end, according to a general aspect of at least one embodiment, a method for video coding is proposed, including, for a block being encoded in a picture, determining linear model parameters for local illumination compensation based on spatially neighboring reconstructed samples and corresponding reference samples; and encoding the block using local illumination compensation based on the determined linear model parameters. Determining the linear model parameters for the block further includes determining refined linear model parameters for a current sub-block in the block, and for encoding the block, local illumination compensation uses a linear model for the current sub-block based on the refined linear model parameters.

[0006] According to another general aspect of at least one embodiment, a method for video decoding is proposed, including: for a block being decoded in a picture, determining linear model parameters for local illumination compensation based on spatially neighboring reconstructed samples and corresponding reference samples; and decoding the block using local illumination compensation based on the determined linear model parameters. Determining the linear model parameters for the block further includes determining refined linear model parameters for a current sub-block in the block, and for decoding the block, local illumination compensation uses a linear model for the current sub-block based on the refined linear model parameters.

[0007] According to another general aspect of at least one embodiment, a device for video coding is proposed, including components for implementing any one of the embodiments of the encoding method.

[0008] According to another general aspect of at least one embodiment, a device for video decoding is provided, including components for implementing any one of the embodiments of the decoding method.

[0009] According to another general aspect of at least one embodiment, a device for video encoding is provided, including one or more processors and at least one memory. The one or more processors are configured to implement any one of the embodiments of the encoding method.

[0010] According to another general aspect of at least one embodiment, a device for video decoding is provided, including one or more processors and at least one memory. The one or more processors are configured to implement any one of the embodiments of the decoding method.

[0011] According to another general aspect of at least one embodiment, determining the refined linear model parameters for a current sub-block includes: accessing the spatial neighboring reconstructed samples and corresponding reference samples of the current sub-block; and determining the refined linear model parameters based on the previously accessed spatial neighboring reconstructed samples and corresponding reference samples for the block. Advantageously, the data of the neighboring samples is used when the data of the neighboring samples becomes available and LIC is performed by the sub-block.

[0012] According to a variant of this embodiment, determining the refined linear model parameters for a current sub-block includes: determining the refined linear model parameters based on all the previously accessed spatial neighboring reconstructed samples and corresponding reference samples for the block.

[0013] According to another variant of this embodiment, determining the linear model parameters for a block includes: iteratively determining the refined linear model parameters for the sub-blocks in the block in raster scan order.

[0014] According to another variant of this embodiment, determining the refined linear model parameters for a current sub-block includes: determining the refined linear model parameters based on the previously accessed spatial neighboring reconstructed samples and corresponding reference samples of the samples closest to the current sub-block.

[0015] According to another variant of this embodiment, store the accessed spatial neighboring reconstructed samples and corresponding reference samples of the current sub-block into the buffer of the previously accessed spatial neighboring reconstructed samples and corresponding reference samples for the block; and determine the refined linear model parameters based on the stored samples.

[0016] According to another variant of this embodiment, process the partial sums of the accessed spatial neighboring reconstructed samples and corresponding reference samples of the current sub-block and store the partial sums into the buffer of the partial sums for the block; and determine the refined linear model parameters based on the stored partial sums.

[0017] According to another general aspect of at least one embodiment, determining refined linear model parameters for a current sub-block includes: determining partial linear model parameters based on spatially neighboring reconstruction samples and corresponding reference samples for the current sub-block; and determining the refined linear model parameters according to a weighted sum of previously determined partial linear model parameters for the sub-block.

[0018] According to another general aspect of at least one embodiment, refined linear model parameters are determined independently for sub-blocks of a block.

[0019] According to another general aspect of at least one embodiment, the reconstruction samples and the corresponding reference samples are co-located with respect to an L-shape that includes a row of samples on the block and a column of samples to the left of the block, and the co-location is determined according to motion compensation information for the block generated by motion compensation order processing.

[0020] According to another general aspect of at least one embodiment, the motion compensation information for a block includes motion predictors, and the motion predictors for the block are refined into motion compensation information in parallel for each sub-block; and the co-location is determined according to the motion predictors for the block rather than the motion compensation information for the block.

[0021] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that contains data content generated by the method or apparatus according to any of the foregoing descriptions.

[0022] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated by the method or apparatus according to any of the foregoing descriptions.

[0023] One or more of the present embodiments also provide a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the above methods. The present embodiments also provide a computer-readable storage medium having stored thereon a bitstream generated according to the above method. The present embodiments also provide a method and apparatus for transmitting a bitstream generated according to the above method. The present embodiments also provide a computer program product including instructions for performing any of the described methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Examples are shown for representing the concepts of coding tree units (CTUs) and coding trees (CTs) of compressed HEVC pictures.

[0025] Figure 2 Examples are shown of deriving LIC parameters from neighboring reconstruction samples and corresponding reference samples converted using motion vectors of square and rectangular blocks in the prior art.

[0026] Figure 3 and 4 shows an example of the derivation of LIC parameters and the compensation of local illuminance in the case of bidirectional prediction.

[0027] Figure 5 shows an example of the downsampling of L-shaped neighboring samples for a rectangular block.

[0028] Figure 6 、 7a Figures 7b and 8 respectively show examples of sub-block-based motion compensation prediction, affine motion compensation prediction, sub-block-based temporal vector prediction, and decoder-side motion vector refinement.

[0029] Figure 9 shows an example encoding or decoding method according to the prior art that includes using a linear model in sub-block-based motion compensation in a pipeline.

[0030] Figure 10 shows an example of an encoding or decoding method according to the general aspects of at least one embodiment.

[0031] Figure 11 shows an example encoding or decoding method according to the general aspects of at least one embodiment that includes using a linear model in sub-block-based motion compensation in a pipeline.

[0032] Figure 12 、 13 Figures 14 shows various examples of reference samples corresponding to the current sub-block LIC linear model according to the general aspects of at least one embodiment.

[0033] Figure 15 shows an example of an encoding or decoding method according to the general aspects of at least one embodiment.

[0034] Figure 16 shows a block diagram of an embodiment of a video encoder in which various aspects of the embodiment can be implemented.

[0035] Figure 17 shows a block diagram of an embodiment of a video encoder in which various aspects of the embodiment can be implemented.

[0036] Figure 18 shows a block diagram of an example apparatus in which various aspects of the embodiment can be implemented. Detailed Description

[0037] It should be understood that the drawings and the description have been simplified to illustrate elements relevant to a clear understanding of the principles, and at the same time, for the purpose of clarity, many other elements found in typical encoding and / or decoding devices have been removed. It will be understood that although the terms first and second may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another.

[0038] Various embodiments are described regarding the encoding / decoding of pictures. They can be used to encode / decoder a part of a picture, such as a slice or a tile, or an entire picture sequence.

[0039] Various methods are described above, and each method includes one or more steps or actions for implementing the method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0040] At least some embodiments relate to methods for exporting and applying LIC parameters for each sub-block in parallel processing in a pipeline architecture.

[0041] In Section 1, some limitations regarding the derivation of linear model parameters for light compensation are disclosed.

[0042] In Section 2, several embodiments of modified methods for deriving linear model parameters for light compensation compatible with a pipeline process are disclosed.

[0043] In Section 3, additional information and general embodiments are disclosed.

[0044] 1 Limitations on Deriving Linear Model Parameters for LIC

[0045] 1.1 Introduction to the Derivation of LIC Parameters

[0046] The tool local illumination compensation (LIC) based on a linear model is used to compensate for the illuminance change between the picture being encoded and its reference picture using a scaling factor a and an offset b. It is adaptively enabled or disabled for each inter-prediction coded coding unit (CU). When LIC is applied to a CU, the mean squared error (MSE) method is adopted to derive the parameters a and b by using the neighboring samples of the current CU and their corresponding reference samples. More specifically, by the motion information MV of the current block on the current CU in the reference picture, the neighboring samples of the current CU on the current block and the corresponding reference CU Figure 2 on the current block Figure 2 on the current block Figure 2Neighboring samples of the reference block on it are used. The LIC parameter minimizes the error between the neighboring samples of the current CU and the corresponding linearly modified reference samples. For example, the LIC parameter minimizes the mean squared error difference (MSE) between the top and left neighboring reconstructed samples rec_cur(r) (accessed Figure 2 of the neighboring reconstructed samples on the right) of the current CU and the top and left neighboring reconstructed samples rec_ref(s) (accessed Figure 2 of the additional reference samples on the left) of their corresponding reference samples determined by inter-frame prediction, where s = r + MV and MV is the motion vector from inter-frame prediction:

[0047] dist = ∑ r∈Vcur,s∈Vref (rec_cur(r) - a.rec_ref(s) - b) 2 (Equation 1)

[0048] The values of (a, b) are obtained by minimizing (Equation 2) using the least squares method:

[0049]

[0050] The enabling or disabling of LIC for the current CU depends on a flag associated with the current CU, called the LIC flag.

[0051] Once the encoder or decoder obtains the LIC parameters for the current CU, the prediction pred(current_block) of the current CU includes the following (unidirectional prediction case):

[0052] pred(current_block) = a × ref_block + b (Equation 3)

[0053] where current_block is the current block to be predicted, pred(current_block) is the prediction of the current block, and ref_block is the reference block constructed using the regular motion compensation (MV) process and used for the temporal prediction of the current block.

[0054] The value of the number N of reference samples used in the derivation is adjusted so that the summation terms in Equation 2 remain below the maximum integer storage value allowed (e.g., N < 2 16 ) or to handle rectangular blocks. Therefore, the reference samples are downsampled (horizontally and / or vertically using the downsampling steps of stepH or stepV) before being used to derive the LIC parameters (a, b), as Figure 5 shown.

[0055] The set of neighboring reconstructed samples and the set of reference samples (see Figure 3The gray samples) have the same quantity and the same pattern. Hereinafter, we use "left samples" to denote the neighboring reconstruction set (or reference sample set) located on the left side of the current block, and use "top samples" to denote the neighboring reconstruction set (or reference sample set) located on the top of the current block. We use "sample set" to denote one of the "left samples" and "top samples" sets. Preferably, the "sample set" belongs to the left or top neighboring line of the block. Generally, the term "L-shaped" denotes a set composed of samples on the row above the current block (top neighboring line) and samples on the column to the left of the current block (left neighboring line), as Figure 2 shown in gray in

[0056] In the case of bidirectional prediction, local illumination compensation is applied to both reference pictures. According to the first variant (referred to as method-a), the LIC process is applied twice, first for reference 0 prediction (LIST-0), and second for reference 1 prediction (LIST_1). Figure 3 shows the derivation of the LIC parameters according to the first variant and their application to each of the reference 0 prediction (LIST-0) and reference 1 prediction (LIST_1). Then, the two predictions are combined together as usual using default weighting (P = (P0 + P1 + 1) >> 1) or bidirectional prediction weighted average (BPWA): P = (g0.P0 + g1.P1 + (1 << (s - 1))) >> s).

[0057] According to the second variant (referred to as method b), in the case of bidirectional prediction, the conventional predictions are first combined, and then a single LIC process is applied. Figure 2 shows the derivation of the LIC parameters according to the second variant and their application to the combined prediction from LIST-0 and LIST_1.

[0058] According to another variant (referred to as method-c, based on method-b), in the case of bidirectional prediction, the conventional predictions are first combined, and then the LIC-0 and LIC-1 parameters are directly derived from the minimization of the error:

[0059] dist = ∑ r∈Vcur,s∈Vref (rec_cur(r) - a0.rec_ref0(s) - a1.rec_ref1(s) - b) 2 (Equation 2b)

[0060] 1.2 Pipeline sub-block processing in inter-frame prediction

[0061] In the latest developments of VVC, some prediction processes between CUs are performed on a per-sub-block basis, further dividing the CU into smaller prediction units and computing the transform on larger CUs. These tools increase data-dependency constraints because the sub-blocks are decoded in parallel to meet real-time constraints, and thus not all neighboring pixels of the sub-blocks are available. For example, neighboring pixels in the current picture are not available. They are being encoded / decoded. Neighboring pixels in the reference picture are available, but the motion vectors used to identify the reference blocks are not yet known (in the case of dmvr, sbTMVP). Additionally, in some implementations, memory access is a bottleneck that limits accessing neighboring pixels only during the process of a given sub-block. For completeness, some of these tools will be briefly introduced below.

[0062] 1.2.1 Affine Motion Compensation Prediction (4×4 Sub-blocks)

[0063] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In the latest developments of VVC, block-based affine transform motion compensation prediction is applied. The affine motion field of a block is described by the motion vectors of two control points (4 parameters) or three control points (6 parameters) (CPMV). Sub-block-based affine transform prediction is applied to each 4×4 luminance sub-block of the current 16×16 block, as Figure 6 shown.

[0064] 1.2.2 Sub-block-based Temporal Motion Vector Prediction (SbTMVP) (8×8 Sub-blocks)

[0065] The latest developments of VVC also support the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the CU in the current picture. SbTMVP predicts at the sub-CU level. Additionally, SbTMVP applies a motion shift before obtaining the temporal motion information from the co-located picture, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. Figure 7a shows the spatial neighboring blocks A0, A1, B0, B1 used by SbTMVP and Figure 7b shows the derivation of the sub-CU motion field by applying the motion shift from the spatial neighbor and scaling the motion information from the corresponding co-located sub-CU. The sub-CU size used in SbTMVP is fixed at 8×8, and like for the affine merge mode, the SbTMVP mode only applies to CUs whose width and height are both greater than or equal to 8.

[0066] 1.2.3 Bidirectional Optical Flow (BDOF) (4×4 Sub-blocks)

[0067] In the latest development of VVC, the Bi-Directional Optical Flow (BDOF) tool, previously known as BIO, is used to refine the bi-directional prediction of CUs at the 4×4 sub-block level. The BDOF mode is based on the optical flow concept, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 and L1 predicted samples. Then the motion refinement is used to adjust the bi-directional prediction samples in the 4×4 sub-block.

[0068] 1.2.4 DMVR (16×16 sub-block)

[0069] In the latest development of VVC, decoder-side motion vector refinement (DMVR) is a bi-directional prediction technique for merging blocks with two initially signaled motion vectors (MVs), which can be further refined by using bilateral matching prediction. In the bi-directional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. For each 16×16 (maximum size; if the block is smaller, the block contains only one sub-block) sub-block, the SAD between two reference blocks based on each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-directional prediction signal.

[0070] 1.2.5 Pipeline sub-block processing

[0071] As described above, some tools for inter-frame prediction may need to utilize several consecutive operations to process the current block in multiple sub-blocks. To meet the real-time constraints, the consecutive operations are pipelined so that multiple sub-blocks are processed in parallel. Figure 9Shows an example encoding or decoding method according to the prior art that includes using a linear model in pipelined sub - block motion compensation. A virtual pipeline data unit (VPDU) is defined as a non - overlapping unit in a picture. In a hardware decoder, consecutive VPDUs are processed simultaneously by multiple pipeline stages. In most pipeline stages, the VPDU size is roughly proportional to the buffer size, so it is important to keep the VPDU size small. A typical example of the VPDU size is 64×64 luma samples. To keep the VPDU size at 64×64 luma samples, a canonical partitioning limit (with syntax signaling modifications) is applied. In the first step of calculating the MV, the initial motion information (such as the motion vector of a block) of a block (VPDU unit) is determined. In the encoding method, the initial motion information is obtained from motion estimation. In the decoding method, the initial motion information is obtained from the encoded bitstream. Then, in the pipelined operation, the initial motion information is refined for each sub - block. For example, a 32×32 block is partitioned into 16×16 sub - blocks. The data of the current sub - block is accessed. Then sub - block processing is applied, for example, in step DMVR, to obtain the refined motion information for the current sub - block. In parallel ( Figure 9 the second line), the data of the subsequent sub - block is accessed using the same hardware resources (memory access) as previously used for the current sub - block. Then the prediction from motion compensation is determined in step MC. In parallel, sub - block processing is applied to obtain the refined motion information for the subsequent sub - block ( Figure 9 the second line).

[0072] Since the derivation of the LIC parameter uses the refined motion information to determine the reference samples for the sub - block, the pipeline imposes a limitation on the availability of the samples necessary for calculating the LIC parameter of the block.

[0073] Therefore, according to Figure 9 the prior - art solution shown, due to the unavailability of the data of the subsequent sub - block, the derivation of the LIC parameter cannot be processed in parallel in the pipeline, so the LIC derivation is postponed after motion compensation of all sub - blocks of the block. The parallel computation of the block is affected, and a delay is introduced in the pipeline.

[0074] According to another prior - art solution, the LIC is disabled in the case of parallel processing of sub - blocks for motion compensation. This solution presents performance issues in the encoding / decoding process.

[0075] Thus, once the motion information for the sub-blocks of a block is available, at least one embodiment improves the linear model process by iteratively deriving refined LIC parameters and applying the resulting linear model per sub-block. This is achieved by: continuously increasing the number of neighboring reconstruction samples and corresponding reference samples during the derivation process using the availability of data, deriving LIC parameters from partial LIC parameters derived from the neighboring samples of the sub-block, independently deriving LIC parameters for each sub-block, or using the initially determined motion information, as detailed in the next section.

[0076] 2 includes at least one embodiment of a method for refining and applying LIC parameters per sub-block

[0077] To address the limitations presented in Section 1, a general aspect of at least one embodiment aims to improve the accuracy of the linear model in a pipelined implementation by refining and applying LIC parameters per sub-block.

[0078] 2.1 includes a general aspect of at least one embodiment for refining and applying LIC parameters per sub-block.

[0079] Figure 10 An example of an encoding or decoding method according to a general aspect of at least one embodiment is shown. Figure 10 The method includes adapting the neighboring reconstruction samples and corresponding reference samples used in the linear model parameter derivation according to the availability of the required data issued from the pipelined parallel processing.

[0080] The encoding or decoding method 10 determines the linear model parameters based on the spatially neighboring reconstruction samples and corresponding reference samples being encoded or decoded. The linear model is then used in the encoding or decoding method. Such linear model parameters include the scaling factor a and offset b of the LIC model as defined in Equations 2 and 3. As Figure 11 illustrated, the block is divided into sub-blocks for parallel processing for inter-frame prediction by motion compensation in the pipeline. According to a non-limiting example, as Figures 12 to 14 shown, a 32×32 block is divided into 4 16×16 sub-blocks. The sub-blocks are sorted from 1 to 4 according to their processing order. According to a non-limiting example, the processing order is as Figures 12 to 14 shown in a raster scan order. According to a non-limiting example, the processing order is determined according to the position of the sub-blocks within the block. Of course, the principles will be readily derivable for other block sizes, sub-block divisions, or sub-block orders.

[0081] In a first step 11, the encoding or decoding method 10 determines the first sub-block for the current block ( Figure 12 the current CU on Figure 12The linear model parameters of 1) above. Based on the available data generated from sub - block - based motion compensation, access the spatially - adjacent reconstructed samples of the first sub - block and the spatially - adjacent samples of the reference sub - block (a reference block of 16×16(1)). According to the motion compensation information MV, the reference sub - block is the co - located sub - block of the first sub - block in the reference picture. For the sake of brevity, in this disclosure, the spatially - adjacent reconstructed samples of a sub - block / block and the spatially - adjacent samples of the reference sub - block / block can be referred to as the adjacent samples of the sub - block / block. The linear model parameters are determined, for example, according to Equation 2, where N is the number of spatially - adjacent reconstructed samples of the first sub - block. Advantageously, both the top samples and the left - hand samples of the first sub - block are available. In step 13, then apply the linear model based on the LIC parameters of the first sub - block to the reference sub - block, as in Equation 3, to obtain the LIC prediction for the first sub - block. Then use the prediction in an encoding or decoding method.

[0082] In parallel with the processing of the first sub - block, a second sub - block ( Figure 12 above 2) is processed. In step 12, determine the refined linear model parameters for the second sub - block based on the newly available data generated from the sub - block - based motion compensation of the first and second sub - blocks. As Figure 12 shown, the current sub - block and the corresponding reference sub - block (a reference block of 16×16(2)) can now access the top - right samples of the current block. As detailed in the examples in Sections 2.2.1 and 2.2.2 later, any combination of the spatially - adjacent reconstructed samples and the corresponding reference samples of the first sub - block or the second sub - block is used to determine the refined LIC parameters for the second sub - block. Then, repeat step 13 for the second sub - block: then apply the linear model based on the refined LIC parameters to the reference sub - block as in Equation 3 to obtain the LIC prediction for the second sub - block. Then use the LIC prediction in an encoding or decoding method.

[0083] Again, in parallel with the processing of the first sub - block and the second sub - block ( Figure 11 the third row of the pipeline), a third sub - block ( Figure 12 above 3) is processed. In the iterative step 12, determine the linear model parameters for the third sub - block based on the newly available data generated from the sub - block - based motion compensation of the first, second, and third sub - blocks. As Figure 12 shown, the current sub - block and the corresponding reference sub - block (a reference block of 16×16(3)) can now access the bottom - left samples of the current block. For a 4 - sub - block partition, the adjacent samples of the first, second, and third sub - blocks define the adjacent samples of the entire block. Again, any combination of the adjacent samples of the first, second, and third sub - blocks is used to determine the refined LIC parameters for the third sub - block. Then, repeat step 13 for the third sub - block: then apply the linear model based on the refined LIC parameters to the reference sub - block as in Equation 3 to obtain the LIC prediction for the third sub - block. Then use the LIC prediction in an encoding or decoding method.

[0084] Finally, in parallel with the processing of the previous sub-blocks, the fourth sub-block ( Figure 12 the 4 on) is processed. In iteration step 12, refined linear model parameters for the fourth sub-block are determined based on any combination of previously available neighboring samples. As Figure 12 shown, no additional neighboring samples are available at this step. Then, step 13 is repeated for the fourth sub-block: Then the linear model based on the refined LIC parameters is applied to the reference sub-block of the fourth sub-block to obtain an LIC prediction for the fourth sub-block. Then the LIC prediction is used in an encoding or decoding method.

[0085] Thus, the linear model parameters for the block are determined by determining refined linear model parameters for the current sub-block in the block and determining an illumination compensation prediction for the current sub-block based on the refined linear model parameters for the current sub-block. Advantageously, the LIC derivation and application according to the general aspects of at least one embodiment are easily compatible with pipelined sub-block based motion compensation.

[0086] Figure 11 An example encoding or decoding method is shown that includes using a linear model in pipelined sub-block based motion compensation according to the general aspects of at least one embodiment. As Figure 11 shown, the LIC (including linear model derivation and linear model application) is processed in parallel per sub-block.

[0087] A more detailed example embodiment is now described in detail.

[0088] 2.2 At least one embodiment including continuously refining LIC parameters with available data

[0089] Not all neighboring samples or motion information are available for calculating LIC parameters for the entire block at once, but are available for calculating per sub-block for sub-blocks.

[0090] In the following method, LIC parameters are calculated for the first sub-block and then refined for subsequent sub-blocks when the data is available, i.e., subsequent sub-blocks are processed in a pipeline for inter-frame prediction.

[0091] According to the first embodiment, neighboring samples are used for the current sub-block accessed in memory and stored in a buffer of LIC samples for the block when they become available. Then the stored samples are used to calculate the LIC parameters. Given the sub-block position, the available neighboring samples can be immediately derived.

[0092] The LIC parameters a and b, which are a scaling factor and an offset respectively, are calculated using equation 2 defined above, but with a subset of available neighboring samples, for sub-block i

[0093]

[0094]

[0095] 2.2.1 Raster scan order

[0096] According to a particular variant of the first embodiment, when calculating each sub - block, neighboring samples are available and all available samples can be used to calculate the LIC parameters. This means that, for example, for the sub - blocks in the second row, all the above - mentioned neighboring samples are available and used to determine the refined LIC parameters for the second, third, and fourth sub - blocks as shown in Figure 12 . In other words, the number of samples used in the LIC parameter derivation increases continuously with the available neighboring data of the block.

[0097] 2.2.2 Position - dependent method with partial model buffer

[0098] According to another particular variant of the first embodiment, when calculating each sub - block as described above, neighboring samples are available, but the LIC parameters are calculated using only the closest available samples. This means that, for example, for the sub - blocks in the second row, instead of using all the above - mentioned neighboring samples, only those directly above (e.g., excluding the upper - right pixel of the lower - left sub - block 3) are used, as shown in Figure 13 .

[0099] 2.2.3 Embodiment with at least partial sums

[0100] According to another particular variant of the first embodiment, instead of storing neighboring samples in the buffer, partial sums (sumX j,k ) of the model are stored in the buffer. Storing a reduced fixed number of values improves the memory footprint and eliminates the recalculation of partial sums for each sub - block LIC process.

[0101] The minimum partial sums can be stored, for example, only the top neighbor of the current sub - block, or only the left neighbor of the current sub - block, where j is the top or left neighborhood and k is the sub - block index:

[0102] sumC j,k = ∑ j,k cur(r)

[0103] sumR j,k = ∑ j,k ref(s)

[0104] sumRC j,k = ∑ j,k ref(s)×cur(r)

[0105] sumCC j,k = ∑ j,kcur(r) × cur(r)

[0106]

[0107]

[0108] According to this variant, store the partial sums with the smallest size corresponding to the sub-block width or sub-block height. In other words, the partial sums cannot be further divided, and each combination of partial sums by addition may determine the linear model parameters.

[0109] Alternatively, according to another specific variant of the first embodiment, if the previous partial sums have been aggregated, store fewer partial sums. For example, the partial sums of the first sub-block (top left) can be aggregated and reused for the second and third sub-blocks. In other words, these fewer partial sums can be divided, but it is useless to maintain more granularity, so they have been pre-combined.

[0110] In another specific variant of the first embodiment, store other partial data. This variant is particularly advantageous if the model does not originate from least squares optimization. The stored partial data is different because they depend on the model. For example, the partial data can be:

[0111] ∑ j,k cur(r)

[0112] ∑ j,k ref(s)

[0113] ∑ j,k |cur(r)|

[0114] 2.3 At least one embodiment including a combined sub-model

[0115] According to the second embodiment, instead of storing the neighboring samples or partial sums for deriving the LIC parameters, derive and store the partial LIC parameters (a i and b i ). For example, the parameters for the first sub-block - top left (1) - are stored as (a1 and b1). The parameters for the second sub-block - top right (2) - are calculated as (normal equations; integer division implemented with shifts):

[0116]

[0117]

[0118] where a 2t and b 2t are partial models calculated only from the top neighbor of the second sub-block. This avoids dividing by a number of samples N that may not be a power of 2.i Advantages

[0119] In addition, this gives more weight to samples closest to the current sub-block in LIC parameter refinement, thus improving prediction accuracy.

[0120] 2.4 At least one embodiment of determining a refinement linear model independently for each sub-block

[0121] According to the third embodiment, instead of using multiple partial sums or partial models, only the available neighboring samples for the current sub-block are used to determine the LIC parameters for the current sub-block, as Figure 14 shown. Similarly, the number of pixels is advantageously directly a power of 2, which simplifies subsequent division (division by a power of 2 is replaced by a simpler bit shift). For a 4-sub-block segmentation, the first sub-block LIC parameter is derived from the left sample of the neighboring samples in the row above and the top sample of the neighboring samples in the left column; the second sub-block LIC parameter is derived from the right sample of the neighboring samples in the row above, the third sub-block LIC parameter is derived from the bottom sample of the neighboring samples in the left column, and LIC is disabled for the fourth block.

[0122] 2.5 At least one embodiment including using motion information from an initial motion calculation

[0123] To prevent data dependency issues, LIC parameters can also be calculated before the sub-block refinement process, as Figure 9 shown in the pipeline process. Motion vector predictors (e.g., MV0 and MV1 in the case of DMVR) Figure 8 are used instead of actual motion vectors (e.g., MV0' and MV1' in Figure 8 ) to obtain neighboring reference pixels. Thus, LIC parameters can be calculated before the sub-block process, and LIC can be applied per sub-block, where the same parameters can be used for each sub-block, as Figure 15 shown.

[0124] 3 Additional embodiments and information

[0125] This application describes multiple aspects, including tools, features, embodiments, models, schemes, etc. Many of these aspects are specifically described and are typically described in a way that may sound restrictive at least for showing individual features. However, this is for the purpose of clear description and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide more aspects. In addition, these aspects can also be combined and interchanged with aspects described in earlier documents.

[0126] The aspects described and considered in this application can be implemented in many different forms. The following Figure 16 ,17 Examples 18 are provided, but other embodiments are also contemplated, and Figure 16 、 17 the discussion of 18 does not limit the breadth of implementation. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects may be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing a bitstream generated according to any of the described methods.

[0127] In this application, the terms "reconstruction" and "decoding" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "image", "picture", and "frame" may be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0128] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.

[0129] The various methods and other aspects described in this application can be used to modify modules, for example, such as Figure 16 and Figure 17 the motion compensation (170, 275), motion estimation (175), entropy coding / decoding, intra (160, 260), and / or decoding modules (145, 230) of the video encoder 100 and decoder 200 shown. Additionally, the proposed aspects are not limited to VVC or HEVC and can be applied to, for example, other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise stated or technically precluded, the aspects described in this application can be used alone or in combination.

[0130] For example, various numerical values are used in this application. Specific values are used for illustrative purposes and the described aspects are not limited to these specific values.

[0131] Figure 16 Encoder 100 is shown. Variants of this encoder 100 are envisioned, but for clarity purposes, encoder 100 is described below without describing all expected variants.

[0132] Before being encoded, the video sequence may undergo pre-encoding processing (101). For example, a color transformation is applied to the input color picture (e.g., a conversion from RGB 4:4:4 to YCbCr 4:2:0), or re-mapping of the input picture components is performed to obtain a signal distribution that is more resilient to compression (e.g., histogram equalization using one of the color components). Metadata may be associated with the pre-processing and appended to the bitstream.

[0133] In the encoder 100, the pictures are encoded by encoder elements as described below. The picture to be encoded is segmented (102) and processed in units such as CUs, for example. Each unit is encoded using an intra or inter mode. When a unit is encoded in the intra mode, it performs intra prediction (160). In the inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of the intra mode or inter mode to use to encode the unit and indicates the intra / inter decision by, for example, a prediction mode flag. For example, the prediction residual is calculated by subtracting (110) the prediction block from the original image block.

[0134] Then the prediction residual is transformed (125) and quantized (130). The quantized transform coefficients, along with the motion vectors and other syntax elements, are entropy coded and decoded (145) to output the bitstream. The encoder may skip the transformation and apply quantization directly to the untransformed residual signal. The encoder may bypass the transformation and quantization, i.e., directly code and decode the residual without applying the transformation or quantization process.

[0135] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155), and the image block is reconstructed. A loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce the encoding of artifacts. The filtered image is stored in the reference picture buffer (180).

[0136] Figure 17 A block diagram of the video decoder 200 is shown. As described below, in the decoder 200, the bitstream is decoded by decoder elements. The video decoder 200 generally performs a decoding channel that is opposite to the encoding channel described in Figure 17 The encoder 100 generally also performs video decoding as part of encoding the video data.

[0137] In particular, the input to the decoder includes a video bitstream, which may be generated by the video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other encoded information. The picture segmentation information indicates how the picture is segmented. Thus, the decoder can partition (235) the picture according to the decoded picture segmentation information. The transform coefficients are dequantized (240) and inverse transformed (250) to decode the prediction residuals. The decoded prediction residuals and the prediction blocks are combined (255), and the image block is reconstructed. The prediction block can be obtained (270) from intra prediction (260) or motion compensated prediction (i.e., inter prediction) (275). A loop filter (265) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (280).

[0138] The decoded picture may also undergo post-decoding processing (285), such as inverse color transformation (e.g., conversion from YCbCr 4:2:0 to RGB 4:4:4) or inverse remapping that performs the inverse process of the remapping process performed in the pre-encoding process (101). The post-decoding processing may use metadata derived in the pre-encoding process and signaled in the bitstream.

[0139] Figure 18 A block diagram illustrating an example of a system in which various aspects and embodiments are implemented is shown. The system 1000 may be embodied as a device including various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of the system 1000 may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components individually or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of the system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, the system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, the system 1000 is configured to implement one or more aspects described in this document.

[0140] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. Processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., volatile memory devices and / or non-volatile memory devices). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, storage device 1040 may include internal storage devices, additional storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.

[0141] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded video or decoded video, and encoder / decoder module 1030 may include its own processor and memory. Encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Additionally, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software known to those skilled in the art.

[0142] Program code to be loaded into processor 1010 or encoder / decoder 1030 to execute various aspects described herein may be stored in storage device 1040 and subsequently loaded into memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of various items during the execution of the processes described herein. Such stored items may include but are not limited to input video, decoded video or a portion of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operation logic.

[0143] In some embodiments, the memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 also refers to ISO / IEC 13818, and 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team JVET).

[0144] Input to the elements of the system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) section that receives an RF signal, for example, transmitted over the air by a broadcaster, (ii) component (COMP) input terminals (or a set of COMP input terminals), (iii) universal serial bus (USB) input terminals, and / or (iv) high definition multimedia interface (HDMI) input terminals. Other examples not shown in Figure 18 include composite video.

[0145] In various embodiments, the input device of block 1130 has associated therewith corresponding input processing elements known in the art. For example, the RF section may be associated with elements suitable for performing the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements to perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0146] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via the USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented within, for example, a separate input processing IC or within processor 1010 as needed. Similarly, various aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within processor 1010 as needed. The streams of demodulation, error correction, and demultiplexing are provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, which operate in combination with memory and storage elements to process the data stream as needed for presentation on an output device.

[0147] The various elements of system 1000 may be disposed within an integrated housing. Within the integrated housing, the various elements may be interconnected using suitable connection arrangements (such as internal buses known in the art, including inter-integrated circuit (I2C) buses, wiring, and printed circuit boards) and data may be transmitted therebetween.

[0148] System 1000 includes a communication interface 1050 capable of communicating with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0149] In various embodiments, a wireless network such as a Wi-Fi network, e.g., IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), is used to transmit or otherwise provide a data stream to system 1000. The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box that transmits data via an HDMI connection of an input box 1130 to provide streaming data to system 1000. Still other embodiments use an RF connection of the input box 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use a wireless network other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0150] System 1000 may provide output signals to various output devices, the various output devices including a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 may be used for a television, a tablet computer, a laptop computer, a mobile phone (cellular phone), or other devices. The display 1100 may also be integrated with other components (e.g., as in a smart phone) or separate (e.g., an external monitor for a laptop computer). In various examples of embodiments, the other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (the abbreviation DVR for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more of the peripheral devices 1120 that function based on the output of system 1000. For example, the disc player performs the function of playing the output of system 1000.

[0151] In various embodiments, control signals are communicated between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 may be integrated in a single unit with other components of system 1000 in an electronic device such as, for example, a television. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.

[0152] Display 1100 and speaker 1110 may alternatively be separate from one or more other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments where display 1100 and speaker 1110 are external components, the output signals may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0153] Embodiments may be implemented by computer software implemented by processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments may be implemented by one or more integrated circuits. Memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, which are non-limiting examples. Processor 1010 may be of any type suitable for the technical environment and may include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0154] Various implementations involve decoding. "Decoding" as used in this application may include, for example, all or part of the process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include or alternatively include processes performed by the decoders of the various implementations described in this application, such as determining local illumination compensation parameters and performing local illumination compensation per sub-block, where the sub-blocks are processed in parallel for motion compensation in a pipeline architecture.

[0155] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Based on the context of the specific description, it will be clear whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally to a broader decoding process, and is considered to be well understood by those skilled in the art.

[0156] All implementations involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can include, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by the encoders of the various implementations described in this application, for example, determining local illumination compensation parameters and performing local illumination compensation on a per-subblock basis to allow subblocks to be processed in parallel for motion compensation in a pipeline architecture.

[0157] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Based on the context of the specific description, it will be clear whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally to a broader encoding process, and is considered to be well understood by those skilled in the art.

[0158] Note that the syntactic elements used here, such as LIC flags, are descriptive terms. Therefore, they do not exclude the use of other syntactic element names.

[0159] When the drawings are presented in the form of a flowchart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when the drawings are presented as a block diagram, it should be understood that it also provides a flowchart of the corresponding method / process.

[0160] The implementations and aspects described herein can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even when discussed in the context of only a single implementation form (e.g., only discussed as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a program). The apparatus can be implemented, for example, with appropriate hardware, software, and firmware. The method can be implemented, for example, with a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0161] References to "an embodiment" or "embodiments" or "an implementation" or "implementations" and other variants thereof mean that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in an embodiment" or "in embodiments" or "in an implementation" or "in implementations" and any other variants thereof throughout this application do not necessarily all refer to the same embodiment.

[0162] In addition, this application may relate to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0163] Furthermore, this application may relate to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0164] In addition, this application may relate to "receiving" various information. Receiving, like "accessing", is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically involved in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0165] It should be understood that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely " / ", "and / or", and "at least one of", is intended to include only selecting the first-listed option (A), or only selecting the second-listed option (B), or selecting both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include only selecting the first-listed option (A), or only selecting the second-listed option (B), or only selecting the third-listed option (C), or only selecting the first and second-listed options (A and B), or only selecting the first and third-listed options (A and C), or only selecting the second and third-listed options (B and C), or selecting all three options (A and B and C). As will be clear to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.

[0166] In addition, as used herein, the term "signal" particularly refers to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a particular one of a plurality of parameters for region-based parameter selection for LIC. For example, enabling / disabling LIC may depend on the size of the region. Thus, in one embodiment, the same parameters are used on both the encoder side and the decoder side. Accordingly, for example, the encoder may transmit (explicitly signal) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as other parameters, signaling (implicitly signal) may be used without transmission to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the term "signal", "signal" can also be used as a noun herein.

[0167] It will be apparent to those of ordinary skill in the art that embodiments can generate various signals that are formatted to carry information such as can be stored or transmitted. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry the bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is well known, signals can be transmitted over a variety of different wired or wireless links. A signal can be stored on a processor-readable medium.

[0168] We have described a number of embodiments. The features of these embodiments can be provided individually or in any combination across various claim categories and types. In addition, embodiments can individually or in any combination across various claim categories and types include one or more of the following features, devices, or aspects:

[0169] · Modify the local illumination compensation used in the inter-frame prediction process applied in the decoder and / or encoder.

[0170] · Modify the derivation of the local illumination compensation parameters used in the inter-frame prediction process applied in the decoder and / or encoder.

[0171] · Adapt the samples used in local illumination compensation to a pipelined motion compensation sub-block architecture.

[0172] ·Iteratively determine refined linear model parameters for the current sub-block in a block based on available data; determine refined linear model parameters based on all previously accessed neighboring samples for the block.

[0173] ·Iteratively determine refined linear model parameters for sub-blocks in a block in raster scan order.

[0174] ·Determine refined linear model parameters based on previously accessed neighboring samples of samples closest to the current sub-block.

[0175] ·Store neighboring samples of the currently accessed sub-block into buffer samples for the block.

[0176] ·Process and store partial sums obtained from neighboring samples of the current sub-block; determine partial linear model parameters for the current sub-block and determine refined linear model parameters based on a weighted sum of previously determined partial linear model parameters for the sub-block.

[0177] ·Refine linear model parameters independently for sub-blocks of a block.

[0178] ·Enable or disable illumination compensation for a sub-block according to the position of the sub-block in the block.

[0179] ·Insert a syntax element in the signaling that enables a decoder to identify the illumination compensation method to be used.

[0180] ·A bitstream or signal including one or more of the described syntax elements or variants thereof.

[0181] ·A bitstream or signal including a syntax that conveys information generated according to any of the described embodiments.

[0182] ·Insert a syntax element in the signaling that enables a decoder to adapt the LIC in a manner corresponding to the way used by the encoder.

[0183] ·Create and / or send and / or receive and / or decode a bitstream or signal including one or more of the described syntax elements or variants thereof.

[0184] ·Create and / or send and / or receive and / or decode according to any of the described embodiments.

[0185] ·A method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the described embodiments.

[0186] ·A TV set-top box, cellular phone, tablet computer, or other electronic device that adapts LIC parameters according to any of the described embodiments.

[0187] · A television, set-top box, cellular phone, tablet computer, or other electronic device that performs adaptation of LIC parameters according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display).

[0188] · A television, set-top box, cellular phone, tablet computer, or other electronic device that selects a channel (e.g., using a tuner) to receive a signal including an encoded image and performs adaptation of LIC parameters according to any of the described embodiments.

[0189] · A television, set-top box, cellular phone, tablet computer, or other electronic device that receives a signal including an encoded image via over-the-air reception (e.g., using an antenna) and performs adaptation of LIC parameters according to any of the described embodiments.

Claims

1. A method for video encoding, comprising: For a block being encoded in a picture, model parameters for local illumination compensation are determined based on spatially neighboring reconstructed pixels and corresponding reference pixels of the block. The block is segmented into a plurality of sub - blocks, motion information of a current sub - block is used to identify the corresponding reference pixels of the current sub - block, and the motion information of the plurality of sub - blocks is successively refined. Determining the model parameters for the block includes: Accessing spatially neighboring reconstructed pixels of the current sub - block among the spatially neighboring reconstructed pixels of the block and corresponding reference pixels of the current sub - block among the corresponding reference pixels of the block; Accessing spatially neighboring reconstructed pixels of one or more sub - blocks among the spatially neighboring reconstructed pixels of the block and corresponding reference pixels of the one or more sub - blocks among the corresponding reference pixels of the block, the one or more sub - blocks being different from the current sub - block, and motion information of the one or more sub - blocks having been previously successively refined to motion information of the current sub - block; and Determining refined model parameters for the current sub - block in the block based on the spatially neighboring reconstructed pixels and the corresponding reference pixels of the one or more sub - blocks and the current sub - block; and Encoding the block using local illumination compensation based on the model parameters, where encoding the block includes processing the local illumination compensation sub - block by sub - block using a model based on the refined model parameters for the current sub - block.

2. The method according to claim 1, wherein, Determining the refined model parameters for the current sub - block includes: determining the refined model parameters based on spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of the one or more sub - blocks and the current sub - block.

3. The method according to claim 1, wherein, Determining the refined model parameters for the current sub - block includes: determining the refined model parameters based on spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of one or more sub - blocks closest to pixels of the current sub - block and the current sub - block.

4. The method according to claim 1, wherein, Determining the refined model parameters for the current sub - block includes: Processing partial sums of spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of the one or more sub - blocks and the current sub - block; Storing the partial sum for the current sub - block into a buffer for partial sums for the block; and Determining the refined model parameters based on the stored partial sums.

5. The method according to claim 1, wherein determining the refinement model parameters for the current sub - block comprises: Determining partial model parameters based on spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub - block; Determining refined model parameters according to a weighted sum of partial model parameters for all sub - blocks of the one or more sub - blocks.

6. A method for video decoding, comprising: For a block being decoded in a picture, model parameters for local illumination compensation are determined based on spatially neighboring reconstructed pixels and corresponding reference pixels of the block. The block is segmented into a plurality of sub - blocks, motion information of a current sub - block is used to identify the corresponding reference pixels of the current sub - block, and the motion information of the plurality of sub - blocks is successively refined. Determining the model parameters for the block includes: Accessing spatially neighboring reconstructed pixels of the current sub - block among the spatially neighboring reconstructed pixels of the block and corresponding reference pixels of the current sub - block among the corresponding reference pixels of the block; Access the spatially neighboring reconstructed pixels of one or more sub - blocks among the spatially neighboring reconstructed pixels of the block and the corresponding reference pixels of one or more sub - blocks among the corresponding reference pixels of the block, where the one or more sub - blocks are different from the current sub - block, and the motion information of the one or more sub - blocks has been continuously refined to the motion information of the current sub - block; and Determine refined model parameters for the current sub - block in the block based on the spatially neighboring reconstructed pixels and the corresponding reference pixels of the one or more sub - blocks and the current sub - block; and Decode the block using local illumination compensation based on the model parameters; where decoding the block includes processing the local illumination compensation sub - block by sub - block using a model based on the refined model parameters for the current sub - block.

7. The method according to claim 6, wherein, Determining the refined model parameters for the current sub - block includes: determining the refined model parameters based on the spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of the one or more sub - blocks and the current sub - block.

8. The method according to claim 6, wherein, Determining the refined model parameters for the current sub - block includes: determining the refined model parameters based on the spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of the one or more sub - blocks that are closest to the pixels of the current sub - block and the current sub - block.

9. The method according to claim 6, wherein, Determining the refined model parameters for the current sub - block includes Processing the partial sums of the spatially neighboring reconstructed pixels and corresponding reference pixels of all sub - blocks of the one or more sub - blocks and the current sub - block; Storing the partial sum for the current sub - block into a buffer for the partial sums of the block; And Determining the refined model parameters based on the stored partial sums.

10. The method according to claim 6, wherein, Determining the refined model parameters for the current sub - block includes: Determining partial model parameters based on the spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub - block; Determining the refined model parameters according to the weighted sum of the partial model parameters of all sub - blocks of the one or more sub - blocks.

11. An apparatus for video encoding, comprising one or more processors and at least one memory, and wherein the one or more processors are configured to: For a block being encoded in a picture, determine model parameters for local illumination compensation based on spatially neighboring reconstructed pixels of the block and corresponding reference pixels, the block being segmented into a plurality of sub - blocks, motion information of a current sub - block being used to identify the corresponding reference pixel of the current sub - block, and the motion information of the plurality of sub - blocks being continuously refined, wherein determining the model parameters for the block comprises: Access the spatially neighboring reconstructed pixels of the current sub - block among the spatially neighboring reconstructed pixels of the block and the corresponding reference pixels of the current sub - block among the corresponding reference pixels of the block; Access the spatially neighboring reconstructed pixels of one or more sub - blocks among the spatially neighboring reconstructed pixels of the block and the corresponding reference pixels of one or more sub - blocks among the corresponding reference pixels of the block, where the one or more sub - blocks are different from the current sub - block, and the motion information of the one or more sub - blocks has been continuously refined to the motion information of the current sub - block; And Determine refined model parameters for the current sub - block in the block based on the spatially neighboring reconstructed pixels and the corresponding reference pixels of the one or more sub - blocks and the current sub - block; And Encode the block using local illumination compensation based on the model parameters; where, for encoding the block, the one or more processors are configured to process the local illumination compensation sub - block by sub - block using a model based on the refined model parameters for the current sub - block.

12. The apparatus according to claim 11, wherein, The one or more processors are configured to: determine the refinement model parameters for the current sub-block based on all sub-blocks of the one or more sub-blocks and spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub-block.

13. The apparatus according to claim 11, wherein, The one or more processors are configured to: determine the refinement model parameters for the current sub-block based on all sub-blocks of the one or more sub-blocks closest to the pixels of the current sub-block and spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub-block.

14. The apparatus according to claim 11, wherein, The one or more processors are configured to: Process partial sums of spatially neighboring reconstructed pixels and corresponding reference pixels from all sub-blocks of the one or more sub-blocks and the current sub-block; Store the partial sum for the current sub-block into a buffer for partial sums of the block; And Determine the refinement model parameters based on the stored partial sums.

15. The apparatus according to claim 11, wherein, The one or more processors are configured to: Determine partial model parameters based on the spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub-block; Determine the refinement model parameters based on a weighted sum of the partial model parameters for all sub-blocks of the one or more sub-blocks.

16. An apparatus for video decoding, comprising one or more processors and at least one memory, and wherein the one or more processors are configured to: For a block being decoded in a picture, determine model parameters for local illumination compensation based on spatially neighboring reconstructed pixels of the block and corresponding reference pixels, the block being segmented into a plurality of sub - blocks, motion information of a current sub - block being used to identify the corresponding reference pixel of the current sub - block, and the motion information of the plurality of sub - blocks being continuously refined, wherein determining the model parameters for the block comprises: Access the spatially neighboring reconstructed pixels of the current sub-block among the spatially neighboring reconstructed pixels of the block and the corresponding reference pixels of the current sub-block among the corresponding reference pixels of the block; Access the spatially neighboring reconstructed pixels of one or more sub-blocks among the spatially neighboring reconstructed pixels of the block and the corresponding reference pixels of the one or more sub-blocks among the corresponding reference pixels of the block, where the one or more sub-blocks are different from the current sub-block, and the motion information of the one or more sub-blocks has been continuously refined to the motion information of the current sub-block; And Determine the refinement model parameters for the current sub-block in the block based on the spatially neighboring reconstructed pixels and the corresponding reference pixels of the one or more sub-blocks and the current sub-block; And Decode the block using the local illumination compensation based on the model parameters, where, in order to decode the block, the one or more processors are configured to use a model based on the refinement model parameters for the current sub-block and process the local illumination compensation sub-block by sub-block.

17. The device according to claim 16, wherein, The one or more processors are configured to: determine the refinement model parameters based on all sub-blocks of the one or more sub-blocks and spatially neighboring reconstructed pixels and corresponding reference pixels.

18. The device according to claim 16, wherein, The one or more processors are configured to: determine the refinement model parameters for the current sub-block based on all sub-blocks of the one or more sub-blocks closest to the pixels of the current sub-block and spatially neighboring reconstructed pixels and corresponding reference pixels of the current sub-block.

19. The device according to claim 16, wherein, The one or more processors are configured to: Process partial sums of spatially neighboring reconstructed pixels and corresponding reference pixels from all sub-blocks of the one or more sub-blocks and the current sub-block; Store the partial sum for the current sub-block into a buffer for partial sums of the block; And Determine the refinement model parameters based on the stored partial sums.

20. The device according to claim 16, wherein, The one or more processors are configured to: Determine partial model parameters based on the spatially adjacent reconstructed pixels and corresponding reference pixels of the current sub-block; Determine refined model parameters according to the weighted sum of the partial model parameters of all sub-blocks for the one or more sub-blocks.

Citation Information

Patent Citations

  • Reference Processing Using Advanced Motion Models for Video Coding

    US20130121416A1

  • Systems and methods of determining illumination compensation parameters for video coding

    US20160366415A1

  • Rolling intra prediction for image and video coding

    US20170230673A1