Improved local lighting compensation

LIC models using templates and spatial domain filters address local illumination variations in video coding, enhancing efficiency and quality in video compression and transmission.

JP2025535311AActive Publication Date: 2025-10-24TENCENT AMERICA LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025522101
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-31
Filing Date
2023-09-05
Publication Date
2025-10-24
Estimated Expiration
2043-09-05

AI Technical Summary

Technical Problem

Existing video coding technologies struggle to effectively compensate for local illumination variations in video encoding and decoding, leading to inefficiencies in compressing and transmitting video data.

Method used

Implementing local illumination compensation (LIC) models that utilize templates of reconstructed neighboring samples and co-located samples to derive parameters for improving video block prediction, including nonlinear terms and n-tap spatial domain filters, to generate compensated samples.

Benefits of technology

Enhances video coding efficiency by accurately modeling local illumination changes, reducing data volume, and improving the quality of compressed video transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025535311000001_ABST
    Figure 2025535311000001_ABST
Patent Text Reader

Abstract

A processing circuit for video decoding receives coding information for a current block in a current picture from a coding video bitstream, the coding information indicating applying local illumination compensation (LIC) to the current block in the current picture. The processing circuit derives parameters of an LIC model according to a first template for the current block and a second template for a reference block in the reference picture. The reference block is indicated based on a motion vector of the current block. The first template includes a subset of reconstructed neighboring samples above and to the left of the current block, and the second template includes co-located samples relative to the subset of reconstructed neighboring samples. The processing circuit applies the LIC model to the current block according to the reference block to generate compensated samples for the current block.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. patent application Ser. No. 18 / 240,950, entitled "IMPROVEMENT OF LOCAL ILLUMINATION COMPENSATION," filed Aug. 31, 2023, which claims the benefit of priority to U.S. provisional application Ser. No. 63 / 417,923, entitled "Improvement of Local Illumination Compensation," filed Oct. 20, 2022. The entire disclosure of the prior application is incorporated by reference.

[0002] [Technical field] This disclosure describes embodiments that relate generally to video coding. [Background technology]

[0003] The background art discussion provided herein is intended to generally present the context for the present disclosure, and the inventors' work, to the extent described in this background art section, as well as aspects of the description that are not admitted as prior art at the time of filing, are not admitted expressly or implicitly as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video files across different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation is commonly represented by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure provide methods and apparatuses for video encoding / decoding. In some examples, the apparatus for video decoding includes a receiving circuit and a processing circuit. The processing circuit receives coding information for a current block in a current picture from a coding video bitstream, where the coding information indicates applying local illumination compensation (LIC) to the current block in the current picture. The processing circuit derives parameters of an LIC model according to a first template (also referred to as a first subset template) of the current block and a second template (also referred to as a second subset template) of a reference block in the reference picture. The reference block is indicated based on a motion vector of the current block. The first template includes a subset of reconstructed neighboring samples above and to the left of the current block, and the second template includes co-located samples relative to the subset of reconstructed neighboring samples. The processing circuit applies the LIC model to the current block according to the reference block to generate compensated samples for the current block.

[0006] In some examples, the first template includes a row of reconstructed neighboring samples immediately above the current block. In some examples, the first template includes a column of reconstructed neighboring samples immediately to the left of the current block. In some examples, the first template includes one or more rows of reconstructed neighboring samples above the current block. In some examples, the first template includes one or more columns of reconstructed neighboring samples to the left of the current block.

[0007] In some examples, the processing circuit decodes syntax indicating a first template for selection from a plurality of candidate templates, and in some examples, the processing circuit determines to use the first template to derive parameters of the LIC model according to at least one of a size of the current block, a shape of the current block, an aspect ratio of the current block, or reconstructed neighboring samples.

[0008] According to one aspect of the present disclosure, a processing circuit receives coding information for a current block in a current picture from a coding video bitstream, the coding information indicating application of local illumination compensation (LIC). The processing circuit determines a reference block in a reference picture based on a motion vector of the current block, and classifies samples in a first block into at least a first class and a second class according to a classification criterion. The first block is one of the current block and the reference block, samples in the second block are classified according to co-located samples in the first block, and the second block is another one of the current block and the reference block. The processing circuit classifies template samples of the first block into at least a first class and a second class according to the classification criterion, and template samples of the second block are classified according to corresponding co-located template samples of the first block. The processing circuit derives first parameters of a first LIC model according to the first class template samples of the first block and the first class template samples of the second block, derives second parameters of a second LIC model according to the second class template samples of the first block and the second class template samples of the second block, and applies the first LIC model to the first class samples in the current block and the second LIC model to the second class samples in the current block to generate compensated samples for the current block.

[0009] In some examples, the first LIC model and the second LIC model have at least one different parameter value.

[0010] In some examples, the processing circuit determines an amplitude threshold based on an average of the sample values ​​in the first block and classifies the sample into a first class or a second class based on a comparison of the sample to the amplitude threshold.

[0011] In some examples, the processing circuit determines a gradient threshold based on gradient values ​​of samples in the first block and classifies the samples into a first class or a second class based on a comparison of the gradient values ​​of the samples with the gradient threshold.

[0012] In some examples, the processing circuit derives the first parameter of the first LIC model and the second parameter of the second LIC model according to at least one of an autocorrelation matrix using a least mean squares operation and an LDL decomposition operation.

[0013] In some examples, the processing circuitry decodes syntax indicating the application of more than one LIC model to generate the compensated samples for the current block.

[0014] According to one aspect of the present disclosure, a processing circuit receives coding information for a current block in a current picture from a coding video bitstream, where the coding information indicates applying local illumination compensation (LIC). The processing circuit derives parameters of an LIC model according to a first template for the current block and a second template for a reference block in the reference picture. The reference block is indicated based on a motion vector of the current block, and the LIC model is different from a linear model based on the amplitude of a single reference sample. The processing circuit applies the LIC model to the current block according to the reference block to generate compensation samples for the current block.

[0015] In some instances, the LIC model includes nonlinear terms.

[0016] In some examples, the LIC model includes an n-tap spatial domain filter, where n is greater than 1. The n-tap spatial domain filter has at least one filter shape of a cross, a diamond, and a square. In one example, the processing circuit performs mean square error minimization according to a first template of the current block and a second template of the reference block to calculate filter coefficients of the n-tap spatial domain filter.

[0017] In some examples, the LIC model includes a slope term that is linearly based on the slope of the reference sample.

[0018] In some examples, the processing circuitry decodes syntax indicating a LIC model for selection from a plurality of candidate LIC models.

[0019] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding. [Brief explanation of the drawings]

[0020] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates the locations of spatial merge candidates according to one embodiment of the present disclosure. [Figure 5] 1 illustrates candidate pairs considered for redundancy checking of spatial merge candidates, according to one embodiment of the present disclosure. [Figure 6] 10 illustrates exemplary motion vector scaling for temporal merge candidates. [Figure 7] 10 illustrates exemplary candidate positions for temporal merge candidates for the current CU. [Figure 8] 10A-10C show diagrams of adjacent sample locations for parameter calculations for cross-component linear models in some examples. [Figure 9] 10 shows the location of the luma sample positions that are input to the spatial 5-tap components in one example. [Figure 10]10 shows a diagram illustrating a reference region containing six lines of chroma samples above and to the left of a prediction unit. [Figure 11] 1 shows four Sobel-based gradient filter patterns in some examples. [Figure 12A] An example of a block template pattern is shown below. [Figure 12B] An example of a block template pattern is shown below. [Figure 12C] An example of a block template pattern is shown below. [Figure 12D] An example of a block template pattern is shown below. [Figure 12E] An example of a block template pattern is shown below. [Figure 12F] An example of a block template pattern is shown below. [Figure 13A] 1 illustrates example filter shapes in some embodiments. [Figure 13B] 1 illustrates example filter shapes in some embodiments. [Figure 13C] 1 illustrates example filter shapes in some embodiments. [Figure 14] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 15] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 16] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 17] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 18] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure. [Figure 19] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure. [Figure 20] FIG. 1 is a schematic diagram of a computer system according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0021] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, a video encoder and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media (including CDs, DVDs, memory sticks, etc.), etc.

[0022] The video processing system (100) includes a video source (101) and a capture subsystem (113) that can include, for example, a digital camera, creating a stream of uncompressed video pictures (102). In one example, the video picture stream (102) includes samples captured by the digital camera. The video picture stream (102), shown with a thick line to emphasize its high data volume compared to the coded video data (104) (or coded video bitstream), can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The coded video data (104) (or coded video bitstream), shown with a thin line to emphasize its low data volume compared to the video picture stream (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as client subsystems 106 and 108 in FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include, for example, a video decoder 110 within an electronic device 130. The video decoder 110 decodes an input copy 107 of the encoded video data and creates an output stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) can be encoded according to several video coding / compression standards. Examples of these standards include ITU-T Recommendation H.265. In one example, a developing video coding standard is informally known as Versatile Video Coding (VVC).The disclosed subject matter may be used in the context of VVC.

[0023] It is noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may include a video encoder (not shown).

[0024] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.

[0025] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210). In one embodiment, one coded video sequence may be received at a time, with the decoding of each coded video sequence being independent of the decoding of the other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device that stores the coded video data. The receiver (231) may receive the coded video data along with other data (e.g., coded audio data and / or auxiliary data streams), which may be forwarded to respective using entities (not shown). The receiver (231) may separate the coded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the "parser" (220)). In certain applications, the buffer memory (215) is part of the video decoder (210). In other cases, it may be external to the video decoder (not shown). In still other cases, there may be a buffer memory (not shown) external to the video decoder 210, for example, to prevent network jitter, and there may be yet another buffer memory 215 within the video decoder 210, for example, to handle playback timing. If the receiver 231 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory 215 may not be needed or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be needed and may be relatively large, advantageously adaptively sized, and may be implemented at least in part in an operating system and similar elements (not shown) external to the video decoder 210.

[0026] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and potentially include information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but can be coupled to the electronic device (230), as shown in FIG. 2. The rendering device control information may be in the form of a Supplementary Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a set of subgroup parameters for at least one subgroup of pixels in a video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.

[0027] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).

[0028] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or portion thereof (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors. Which units are involved and how they are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.

[0029] In addition to the functional blocks described above, the video decoder (210) can be conceptually subdivided into multiple functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate:

[0030] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220), along with control information (including which transform to use, block size, quantization coefficients, quantization scaling matrix, etc.) The scalar / inverse transform unit (251) can output blocks containing sample values ​​that can be input to an aggregator (255).

[0031] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks may relate to blocks that do not use prediction information from a previously reconstructed picture but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates blocks of the same size and shape as the block being reconstructed using surrounding, already reconstructed information retrieved from the current picture buffer (258). The current picture buffer (258), for example, buffers the partially reconstructed and / or fully reconstructed current picture. In some cases, the aggregator (255) adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).

[0032] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to retrieve samples used for prediction. After motion-compensating the retrieved samples according to the symbols (221) associated with the block, these samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) retrieves prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221) that may have X, Y, and reference picture components. Motion compensation may also include interpolation of sample values ​​retrieved from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.

[0033] The output samples of the aggregator 255 may be subjected to various loop filtering techniques in a loop filter unit 256. The video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and made available to the loop filter unit 256 as symbols 221 from the parser 220. The video compression may also be responsive to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, as well as to previously reconstructed loop-filtered sample values.

[0034] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and stored in a reference picture memory (257) for use in future inter-picture prediction.

[0035] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.

[0036] The video decoder (210) may perform decoding operations according to a given video compression technology or standard, such as ITU-T Rec. H.265. A coded video sequence may conform to the syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence conforms to both the syntax and profile of the video compression technology or standard as documented in the video compression technology or standard. Specifically, a profile can select certain tools from all tools available in the video compression technology or standard as the only tools for use under that profile. Compliance may also require the complexity of the coded video sequence to fall within a range defined by a level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further constrained through a hypothetical reference decoder (HRD) specification and metadata about HRD buffer management conveyed in the coded video sequence.

[0037] In one embodiment, the receiver (231) may receive additional (redundant) data along with the coded video. The additional data may be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0038] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.

[0039] The video encoder (303) may receive video samples from a video source (301) (not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).

[0040] The video source (301) may provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 Y CrCB, RGB, etc.), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media presentation system, the video source (301) may be a storage device that stores pre-prepared video. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. Video data may be provided as multiple individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, each of which may contain one or more samples, depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. The following discussion will focus on samples.

[0041] According to one embodiment, the video encoder (303) may encode and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other required time constraints. Achieving an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) can control and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller (350) can include rate control-related parameters (e.g., picture skip, quantization, lambda value for rate-distortion optimization techniques), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other appropriate functionality associated with the video encoder (303) optimized for a particular system design.

[0042] In some embodiments, the video encoder (303) is configured to operate in an encoding loop. As a very simplified description, in one example, the encoding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data similar to that generated by the (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-for-bit accurate results independent of the location of the decoder (local or remote), the contents in the reference picture memory (334) are also bit-for-bit accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values ​​as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (including the resulting drift when synchronization cannot be maintained due to, for example, channel errors) is used in several related technologies as well.

[0043] The operation of the "local" decoder (333) can be the same as the "remote" decoder (210), such as the video decoder (210), already described in detail above in connection with Figure 2. However, briefly referring to Figure 2, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder (345) and parser (220) can be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (433).

[0044] In one embodiment, decoder technology, excluding analysis / entropy decoding, present in the decoder is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the subject matter of the disclosure focuses on decoder operation. A description of the encoder technology can be omitted, as it is the reverse of the decoder technology, which is described generically. In certain areas, more detailed descriptions are provided below.

[0045] During operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.

[0046] The local video decoder (333) may decode the coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may typically be a replica of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures and store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have common content as reconstructed reference pictures (without transmission errors) obtained by the far-end video decoder.

[0047] The predictor (335) may perform a predictive search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or specific metadata (reference picture motion vectors, block shapes, etc.), which may serve as suitable prediction references for the new picture. The predictor (335) may operate sample block-by-pixel block to find suitable prediction references. In some cases, the input picture determined by the search results obtained by the predictor (335) may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).

[0048] The controller (350) may manage the encoding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.

[0049] The output of all the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0050] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which may be a hardware or software link to a storage device that stores the coded video data. The transmitter (340) may merge the coded video data from the video encoder (330) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (not shown).

[0051] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign each coded picture a particular coding picture type. The coding picture type may affect the coding technique that may be applied to each picture. For example, a picture may be assigned as one of the following picture types:

[0052] An intra picture (I-picture) may be one that can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art will recognize these variations of I-pictures and their respective uses and characteristics.

[0053] A predicted picture (P picture) may be one that can be coded and decoded using intra- or inter-prediction, which uses at most one motion vector and reference index to predict the sample values ​​of each block.

[0054] Bidirectionally predicted pictures (B-pictures) may be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multi-predicted pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0055] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I-picture may be non-predictively coded or predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.

[0056] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Rec. H.265. In doing so, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancy in the input video sequence. Thus, the coded video data may conform to a syntax specified by the video coding technique or standard being used.

[0057] In one embodiment, the transmitter (340) may transmit additional data along with the coded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other types of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0058] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or other) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. When a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. A motion vector points to a reference block within the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0059] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. Bi-prediction techniques use two reference pictures, such as a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may also be past and future, respectively, in display order) in a video. A block in the current picture can be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. A block can be predicted by a combination of the first and second reference blocks.

[0060] Furthermore, to improve coding efficiency, merge mode techniques can be used in inter-picture prediction.

[0061] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be partitioned into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine the CU's prediction type, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In one embodiment, the prediction operation of coding (encoding / decoding) is performed on a prediction block basis. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values ​​(e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.

[0062] It is noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.

[0063] Aspects of the present disclosure provide a further technique that can be used in inter-prediction techniques, called local illumination compensation (LIC), to improve coding performance.

[0064] Various inter-prediction modes can be used in video coding. For example, in VVC, for an inter-predicted CU, motion parameters can include an MV, one or more reference picture indexes, a reference picture list usage index, and further information about specific coding features to be used for inter-predicted sample generation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU and cannot have significant residual coefficients, coded motion vector or MV differences (e.g., MVDs), or reference picture indices. A merge mode can be specified, in which motion parameters for the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, and optionally further information as introduced in VVC. The merge mode can be applied not only to skip mode but also to inter-predicted CUs. In one example, an alternative to the merge mode is explicit transmission of motion parameters, in which the MV, the corresponding reference picture index for each reference picture list, and a reference picture list usage flag, as well as other information, are explicitly signaled for each CU.

[0065] In one VVC-like embodiment, the VVC Test model (VTM) reference software supports enhanced merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), and geometric partitioning mode (GPM). Inter-prediction and related methods are described in more detail below.

[0066] In some examples, enhanced merge prediction can be used. In one example, such as VTM4, a merge candidate list is constructed by including, in order, five types of candidates: spatial motion vector predictor (MVP) from spatially neighboring CUs, temporal MVP from co-located CUs, history-based MVP (HMVP) from a first-in-first-out (FIFO) table, pairwise average MVP, and zero MV.

[0067] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowed size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the best merge candidate (e.g., merge index) can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0068] Some examples of the generation process for each category of merge candidates are provided below. In one embodiment, spatial candidates are derived as follows: The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates in the positions shown in Figure 4.

[0069] 4 shows the positions of spatial merge candidates according to an embodiment of the present invention. Referring to FIG. 4, the derivation order is B1, A1, B0, A0, and B2. Position B2 is only considered when none of the CUs in positions A0, B0, B1, and A1 are available (e.g., because the CUs belong to another slice or another tile) or are intra-coded. After the candidate in position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0070] To reduce computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs connected by arrows in Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used in the redundancy check do not have the same motion information.

[0071] 5 illustrates candidate pairs considered for spatial merge candidate redundancy checking according to an embodiment of the present disclosure. Referring to FIG. 5, the pairs connected by arrows are A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and candidates at positions B0 and / or B2 can be compared with the candidate at position B1.

[0072] In one embodiment, temporal candidates are derived as follows: In one example, only one temporal merge candidate is added to the candidate list. Figure 6 shows exemplary motion vector scaling for temporal merge candidates. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) (e.g., indicated by a dotted line in Figure 6) can be derived based on a co-located CU (612) belonging to a co-located reference picture (604). The reference picture list used to derive the co-located CU (612) can be explicitly signaled in the slice header. The scaled MV (621) for the temporal merge candidate can be obtained as indicated by a dotted line in Figure 6. The scaled MV (621) can be scaled from the MV of the co-located CU (612) using picture order count (POC) distances tb and td. The POC distance tb may be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td may be defined as the POC difference between the co-located reference picture (604) of the co-located picture (603) and the co-located picture (603). The reference picture index of the temporal merge candidate may be set to zero.

[0073] 7 shows exemplary candidate positions (e.g., C0 and C1) for temporal merge candidates for the current CU. The position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is located at the bottom right corner of the co-located CU (710) of the current CU. Candidate position C1 is located at the center of the co-located CU (710) of the current CU. If the CU at candidate position C0 is unavailable, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and located in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0074] In some examples, local illumination compensation (LIC) is used as an inter-prediction technique to model local illumination variations between a current block and its predicted block (also called a reference block) by using a linear function. The predicted block is in a reference picture and may be pointed to by a motion vector (MV). Parameters of the linear function may include a scale α and an offset β, and the linear function may be represented by α×p[x,y]+β to compensate for illumination changes, where p[x,y] indicates a reference sample at position [x,y] within the reference block (also called a predicted block), and the reference block is pointed to by the MV. In some examples, the scale α and the offset β may be derived based on a template of the current block and a corresponding reference template of the reference block by using a least-squares method; therefore, no signaling overhead is required, except that a LIC flag may be signaled to indicate the use of LIC.

[0075] In some examples, LIC is used for uni-predictive inter CUs. In some examples, intra-neighboring samples (neighboring samples predicted using intra prediction) of the current block can be used in LIC parameter derivation. In some examples, LIC is disabled for blocks with fewer than 32 luma samples. In some examples, for non-sub-block modes (e.g., non-affine modes), LIC parameter derivation is performed based on the template block samples of the current CU instead of the partial template block samples for the first top-left 16x16 unit. In some examples, LIC parameter derivation is performed based on partial template block samples, such as the partial template block samples for the first top-left 16x16 unit. In some examples, the template samples of the reference block are determined by using motion compensation (MC) with the MVs of the block without rounding to integer pel precision.

[0076] Some aspects of the present disclosure provide techniques for providing adjustments to LIC, so that the LIC can be flexibly adjusted for various scenarios and thus improve the accuracy of illumination compensation when LIC is enabled. In some examples, the techniques for adjusting LIC can include features that are similarly used in other model-derived based prediction modes, such as features in cross-component intra prediction modes.

[0077] Cross-component intra prediction modes can include a first technique called a cross-component linear model (CCLM), a second technique called a multi-model linear model (MMLM), a third technique called a convolutional cross-component model (CCCM), and a fourth technique called a gradient linear model (GLM).

[0078] In some cases (e.g., VVC), the first technique, CCLM, is used to reduce inter-component redundancy. In CCLM, chroma samples are predicted based on the reconstructed luma samples of the same CU by using a linear model such as that shown in Equation (1). pred C (i,j)=a·rec L '(i,j)+b Equation (1) Here pred C (i,j) represents the predicted chroma sample in the CU. L '(i,j) represents the downsampled reconstructed luma sample of the same CU. In one example, the CCLM linear model includes parameters (a and b) that can be derived using at most four neighboring chroma samples and their corresponding downsampled luma samples.

[0079] In some examples, based on the positions of neighboring chroma samples, CCLM can include different modes called LM_T (LM-up mode or above mode LM_A), LM_L (LM-left mode), and LM_LT (LM-upper-left mode or above-left mode LM_LA or simply LM mode). For example, if the current chroma block dimensions are W x H, W' and H' can be set for various modes in CCLM. When LM mode (also called LM_LT or LM_LA) is applied, W' = W, H' = H; when LM-A mode is applied, W' = W + H; and when LM-L mode is applied, H' = H + W.

[0080] The top adjacent positions are denoted as S[0,-1]...S[W'-1,-1], and the left adjacent positions are denoted as S[-1,0]...S[-1,H'-1]. Four positions are then selected. For example, when the LM mode is applied and both the upper and left adjacent samples are available, the four positions may include S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4] and S[-1,3*H' / 4]; when the LM_A mode is applied or only the upper adjacent sample is available, the four positions may include S[W' / 8,-1], S[3*W' / 8,-1], S[5*W' / 8,-1] and S[7*W' / 8,-1]; when the LM_L mode is applied or only the left adjacent sample is available, the four positions may include S[-1,H' / 8], S[-1,3*H' / 8], S[-1,5*H' / 8] and S[-1,7*H' / 8].

[0081] The four adjacent luma samples at the selected position are downsampled to x 0 A and x 1 A The two larger values ​​denoted by x 0 B and x 1 B The corresponding chroma sample values ​​are compared to find the smaller of the two values ​​denoted by y 0 A, y 1 A , y 0 B and y 1 B Then, the intermediate parameter X a , X b , Y a and Y b is derived as follows: X a =(x 0 A +x 1 A +1)>>1 X b =(x 0 B +x 1 B +1)>>1 Y a =(y 0 A +y 1 A +1)>>1 Y b =(y 0 B +y 1 B +1)>>1

[0082] Finally, the linear model parameters a and b are obtained according to equations (2) and (3). a=(Y a -Y b ) / (X a -X b ) Formula (2) b=Y b -a·X b Formula (3)

[0083] FIG. 8 shows a diagram of neighboring sample positions for parameter calculation for CCLM in some examples. Referring to FIG. 8, neighboring sample pairs (luma samples and chroma samples) for deriving parameters for CCLM prediction are indicated by shaded circles. In the example of FIG. 8, the current luma block size is 2N×2N, and the current chroma block size is N×N. The neighboring sample pairs may include 2N reference sample pairs, such as 2N reference samples neighboring the current chroma block and 2N reference samples neighboring the current luma block. According to one aspect of the present disclosure, the neighboring sample positions in FIG. 8 can be used when the LM mode (also referred to as LM_LT or LM_LA) is applied.

[0084] In some examples, in LM_T (also called LM_A) mode, a top template is used to calculate the parameters of the linear model. To obtain more samples, the top template is extended to (W'=W+H) samples. In some examples, in LM_L mode, a left template is used to calculate the parameters of the linear model. To obtain more samples, the left template is extended to (H'=H+W) samples.

[0085] In some examples, in LM_LT mode, left and top templates are used to calculate the parameters of the linear model. In one example, to match the chroma sample positions for a 4:2:0 video sequence, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filter can be specified by an SPS level flag.

[0086] In some examples (e.g., VVC), CCLM is extended by using a second technique, multi-model LM (MMLM). In MMLM mode, a threshold is calculated as the average of luma reconstructed neighboring samples. The reconstructed neighboring samples are then classified into two classes using the threshold: a first class of reconstructed neighboring samples that are greater than the threshold, and a second class of reconstructed neighboring samples that are less than the threshold. In one example, a linear model for each class is derived using a least-mean-square (LMS) method. In some examples, a slope adjustment can be applied to CCLM and MMLM. The slope adjustment can tilt the linear function that maps luma values ​​to chroma values ​​relative to a center point determined by the average luma value of reference samples within the neighboring samples.

[0087] In some examples, a third technique called a convolutional cross-component model (CCCM) can be used to predict chroma samples from reconstructed luma samples, such as in a manner similar to the CCLM mode in ECM-6.0. In CCCM, similar to CCLM, when chroma subsampling is used, the reconstructed luma samples are downsampled to match a lower-resolution chroma grid. Also, similar to CCLM, there is the option to use a single-model or multi-model variant of CCCM. In one example, the multi-model variant uses two models: one model is derived for samples above the average luma reference value, and another model is derived for the rest of the samples (in accordance with the CCLM design intent). In some examples, the multi-model CCCM mode can be selected for PUs with at least 128 reference samples available.

[0088] In CCCM, a convolution filter is used. In some examples, the convolution filter is a 7-tap convolution filter. The convolution 7-tap filter can include a first term of a 5-tap plus-sign shape spatial component (also called a spatial 5-tap component, and the spatial 5-tap component can have a cross shape, also called a plus-sign shape), a second term of a nonlinear term P, and a third term of a bias term.

[0089] 9 shows the location of luma sample positions that are input to the spatial 5-tap component in one example. The input to the spatial 5-tap component of the convolutional 7-tap filter includes a center (C) luma sample co-located with the predicted chroma sample and its top / north (N), bottom / south (S), left / west (W), and right / east (E) neighbors.

[0090] In some examples, the nonlinear term P is expressed as a power of two of the central luma sample C and scaled to the sample value range of the content, such as according to equation (4). P=(C×C+midVal)>>bitDepth Formula (4) In one example, for 10-bit content, the nonlinear term P is calculated according to equation (5). P=(C×C+512)>>10 Formula (5)

[0091] In some examples, the bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM), and in one example is set to the midpoint chroma value (512 for 10-bit content).

[0092] In some examples, the output of a convolution 7-tap filter is the filter coefficients c i and the input value, and is clipped to the range of valid chroma samples, e.g., according to equation (6). predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B Formula (6)

[0093] In some examples, the filter coefficients c ican be determined (computed) by minimizing the mean squared error (MSE) between the predicted and reconstructed chroma samples in the reference domain.

[0094] FIG. 10 shows a diagram illustrating a reference region containing six lines of chroma samples above and to the left of a PU. The reference region extends one PU width to the right of the PU boundary and one PU height below. In some examples, the reference region is adjusted to include only available samples. In some examples, the extension to the reference region is used to support "side samples" of a plus shape spatial filter, which are padded when they fall within unavailable areas.

[0095] In some examples, MSE minimization is performed by calculating an autocorrelation matrix of the luma input and a cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is ​​LDL decomposed, and the final filter coefficients are calculated using inverse substitution. In one example, the MSE minimization process for the filter coefficients is similar to the calculation of the ALF filter coefficients in ECM, but to avoid using square root operations, LDL decomposition is used instead of Cholesky decomposition in the MSE minimization process for the filter coefficients.

[0096] In some examples, a fourth technique, gradient linear model (GLM), is used. Compared to CCLM, instead of downsampled luma values, GLM utilizes luma sample gradients to derive a linear model. Specifically, when GLM is applied, the input to the CCLM process, i.e., the downsampled luma samples L, is replaced with the luma sample gradient G. Other parts of CCLM (e.g., parameter derivation, predicted sample linear transformation) remain unchanged. In one example, the prediction C is calculated according to Equation (7): C=a·G+b Equation (7)

[0097] In some examples, for signaling purposes, when CCLM mode is currently enabled for a CU, two flags are signaled separately for the Cb and Cr components to indicate whether GLM is enabled for each component. In some examples, when GLM is enabled for a component, one syntax element is further signaled to select one of four gradient filters for gradient calculation.

[0098] FIG. 11 shows examples of four Sobel-based gradient filter patterns that can be used in GLM.

[0099] Some aspects of the present disclosure provide an adjustment technique for LIC, so that the LIC can be flexibly adjusted for various scenarios, and thus the accuracy of illumination compensation when the LIC is enabled can be improved.

[0100] Note that various aspects of LIC can be modified through adjustments. In some examples, samples can be classified into different classes, and different classes can have different derived models; therefore, multiple models can be used for LIC (referred to as multi-model LIC). In some scenarios, a block may contain objects that experience different illumination changes; multi-model LIC can more accurately compensate for the illumination changes of the objects. In some examples, the configuration of the template can be adjusted; the template's position, size, etc. can be adjusted. In some scenarios, some of the neighboring samples of a block may experience similar illumination changes, while others may have different illumination changes; therefore, adjustments to the template can be used to more accurately compensate for the illumination changes. In some examples, different formulas for modeling LIC can be used to improve the accuracy of illumination compensation.

[0101] According to one aspect of the present disclosure, multiple models can be used for LIC (referred to as multi-model LIC). In some examples, samples can be classified into different classes, and the different classes can use different models derived from template samples of the different classes.

[0102] In some examples, for LIC, one of the current block (in the current picture) and the reference block (in the reference picture) is referred to as the first block, and the other of the current block and the reference block is referred to as the second block. For example, when the current block is referred to as the first block, the reference block is referred to as the second block, and similarly, when the reference block is referred to as the first block, the current block is referred to as the second block. In some examples, the classification can be determined based on samples related to the reference block, such as samples of the reference block and the reference block template, and then the samples related to the current block, such as samples of the current block and the current block template, can have the same class as the co-located samples related to the reference block. Similarly, in some examples, the classification can be determined based on samples related to the current block, such as samples of the current block and the current block template, and then the samples related to the reference block, such as samples of the reference block and the reference block template, can have the same class as the co-located samples related to the current block. Although some descriptions perform classification based on samples related to the reference block, the description can be modified to perform classification based on samples related to the current block.

[0103] In some examples, samples related to the reference block, such as samples of the reference block and samples of the reference block template, are classified into multiple classes based on specific classification criteria. Then, samples related to the current block, such as samples of the current block and samples of the current block template, are classified based on the class index of the co-located reference sample. For each class, a model is derived using samples of the current block template and samples of the reference block template that have the same class index. Then, predicted samples for each class are derived using the model and co-located reference samples (samples of the reference block) of the same class.

[0104] In some examples, an average sample value of the reference block (and / or reference block template) is calculated, and a classification criterion can classify samples into two classes based on the average sample value. For example, a sample is compared with the average sample value, and if the sample is smaller than (or equal to or smaller than) the average sample value, the sample is classified as the first class; otherwise, the sample is classified as the second class. Then, respective models can be derived for the multiple classes. Note that multi-model LIC based on the average sample value can be referred to as amplitude-based multi-model LIC.

[0105] In some examples, the model derivation for each class can be implemented using least-mean-square (LMS), which is similarly used in CCLM / MM-CCLM. In some examples, the model derivation for each class can be implemented using an autocorrelation matrix with LDL decomposition in the filter coefficient calculation for the convolution filter of CCCM.

[0106] In some examples, multi-model LIC (e.g., amplitude-based multi-model LIC, etc.) can replace single-model LIC without any syntax changes. In one example, a flag is signaled in high-level syntax such as SPS, PPS, picture header, slice header to indicate whether multi-model is used or not.

[0107] In some examples, when LIC is applied (e.g., indicated by the first flag), a second flag is signaled to indicate whether multi-model LIC (e.g., amplitude-based multi-model LIC, etc.) or single-model LIC is selected.

[0108] Note that various classification criteria can be used. In some examples, samples can be classified according to the gradient of the sample. In some examples, the gradient of the sample can be calculated as the difference between the sample value and an adjacent sample value. For example, the gradient of the sample is calculated as the difference between the sample value and the right adjacent sample value. In some other examples, the gradient of the sample is calculated based on a gradient filter pattern, such as one of the four Sobel-based gradient filter patterns in FIG. 11.

[0109] In some examples, a gradient value for each sample is calculated, and a cumulative gradient within the reference block (and / or reference block template) can be statically calculated. For example, a histogram of gradient values ​​within the reference block can be determined, and a classification threshold can be determined based on the histogram. For example, a median of the gradient values ​​can be determined based on the histogram. To classify a sample, the gradient value of the sample is compared to the median. If the gradient value of the sample is less than the median (or equal to or less than the median), the sample is classified into a first class; otherwise, the sample is classified into a second class. Note that multi-model LIC based on the gradients of samples can be referred to as gradient-based multi-model LIC.

[0110] In some examples, the mean (or median) of the gradient values ​​is used as the classification threshold, but if the mean (or median) of the gradient values ​​is lower than a predetermined threshold, the gradient-based multi-model LIC is presumed to be invalid.

[0111] We note that similar signaling techniques for amplitude-based multi-model LIC can be used for gradient-based multi-model LIC.

[0112] Furthermore, in some examples, gradient-based multi-model LIC can be combined with amplitude-based multi-model LIC, which can be referred to as combined amplitude and gradient-based multi-model LIC. For example, by using the combination of gradient-based multi-model and amplitude-based multi-model in the above example, four different models are used to code the block. The model selection for each sample is determined based on the amplitude of the sample value and the classification of the sample's gradient value.

[0113] We note that the combined amplitude and gradient-based multi-model LIC can use similar syntax signaling techniques as the amplitude-based multi-model LIC.

[0114] In some embodiments, syntax is signaled to indicate which multi-model LIC is selected for LIC, such as amplitude-based multi-model LIC, gradient-based multi-model LIC, combined amplitude-and-gradient-based multi-model LIC, etc. In one embodiment, a top template, a left template, or both a left template and a top template may be selected for LIC. In one embodiment, M rows may be used for the top template and / or N columns may be used for the left template, where both M and N are non-zero positive integer values. In one embodiment, syntax is signaled to indicate which template is selected for LIC. In one embodiment, template selection is implicitly derived based on coding information, including, but not limited to, block size, block shape, block aspect ratio, and neighboring reconstruction samples.

[0115] According to one aspect of the present disclosure, for example, different template sizes and / or template positions can be used in LIC to derive parameters of the LIC model. For example, a top template, a left template, and both a left template and a top template can be selected for LIC. For ease of explanation, when a template includes a subset of reconstructed neighboring samples above and to the left of a current block, in some examples, the template can be referred to as a subset template.

[0116] 12A-12F show some examples of block templates and / or subset templates. In FIG. 12A, the block template includes the top row, left column, and top left position with respect to the block. In FIG. 12B, the block template (subset template) includes the top row and left column with respect to the block. In FIG. 12C, the block template (subset template) includes the column immediately to the left of the block. In FIG. 12D, the block template (subset template) includes the row immediately above the block. In FIG. 12E, the block template (subset template) includes multiple columns to the left of the block. In FIG. 12F, the block template (subset template) includes multiple rows above the block.

[0117] In some examples, the block template for the LIC can be selected from a variety of candidate templates, such as the examples in Figures 12A-12F.

[0118] In some examples, M rows can be used for the top template and / or N columns can be used for the left template, where both M and N are non-zero positive integer values.

[0119] In some examples, syntax is signaled to indicate which of the template candidates is selected as the template for the LIC.

[0120] In some examples, template selection is implicitly derived based on coding information, including but not limited to block size, block shape, block aspect ratio, neighboring reconstructed samples, and so on.

[0121] According to one aspect of the present disclosure, different equations can be used to model LIC. The equations for modeling LIC can be different from those used in CCLM.

[0122] In some examples, (p[x,y]) 2 A nonlinear term such as can be used in the equation to model the LIC according to equation (8) or the like. predC=α0×p[x,y]+α1×p[x,y] 2 +β Equation (8) where p[x,y] is the reference sample pointed to by the MV at position [x,y] on the reference block, and parameters α0, α1, and β are model parameters. In one example, an autocorrelation matrix using LDL decomposition is used to obtain the parameters α0, α1, and β.

[0123] In some examples, an n-tap spatial domain filter (also referred to as an n-tap filter, spatial domain filter, etc.) is applied to the reference samples to form an equation, such as according to equation (9). predC=Σ k=0 n-1 α k ×p[x k ,y k ]+β Equation (9) where p[x k ,y k ] is the position [x k ,y k ] are the reference samples pointed to by MV in the n-1 and β are model parameters.

[0124] 13A-13C show examples of filter shapes for LICs in some embodiments. Note that the n-tap filter shape can be, but is not limited to, a cross shape as shown in FIG. 13A, a diamond shape as shown in FIG. 13B, or a square shape as shown in FIG. 13C.

[0125] In some examples, syntax is signaled to indicate which n-tap filter shape is selected when an n-tap filter is used.

[0126] In some examples, the spatial domain filter is a symmetric filter with n / 2 filter coefficients when n is even, and (n+1) / 2 filter coefficients when n is odd.

[0127] In some examples, a nonlinear term may also be added along with the spatial domain filter, such as according to equation (10). predC=Σ k=0 n-1 α k ×p[x k ,y k ]+α n ×p[x,y] 2 +β Equation (10)

[0128] In some examples, more than one nonlinear term can be used. Note that in some examples, the number of nonlinear terms is less than or equal to the number of spatial filter taps, n.

[0129] In some cases, the autocorrelation matrix using the LDL decomposition is calculated using the parameter α k (k∈[0,n-1]) and is used to obtain β.

[0130] In some examples, an n-tap spatial domain filter is applied to the reference samples to perform LIC, such as according to equation (11). predC=Σ k=0 n-1 α k ×p[x k ,y k ] Formula (11) where p[x k ,y k ] is the position [x k ,y k ] is the reference sample pointed to by MV. Furthermore, the autocorrelation matrix using LDL decomposition is k and k∈[0,n-1].

[0131] In some embodiments, the gradient values ​​of the reference samples may be used as input to perform LIC according to equation (12), or the like. predC=α×G[x,y]+β Equation (12) where G[x,y] is the gradient of the reference sample pointed to by the MV at position [x,y] on the reference block. In one example, an autocorrelation matrix using LDL decomposition is used to obtain the parameters α and β.

[0132] In some examples, the gradient term of the reference sample mentioned in equation (12) can be added to other equations such as equations (8), (9), (10), and (11).

[0133] In some examples, pixel downsampling in the current block template and the reference block template can be used for model derivation.

[0134] In some examples, syntax is signaled to indicate which model is used for a coding block, such as equation (8), equation (9), equation (10), or equation (11).

[0135] According to one aspect of the present disclosure, different LIC adjustment techniques can be combined, for example, multi-model derivation techniques and different equation for modeling techniques can be combined to perform LIC.

[0136] In some examples, each class may have its own model derivation method or formula, such as one of Equations (8)-(12) and an associated model derivation method. In some examples, a model derivation method is always used for that class.

[0137] In some examples, syntax is signaled for each class to indicate which model derivation method or formula is used, including but not limited to Equations (8)-(12) and related model derivation methods.

[0138] In some examples, all possible combinations or a subset of all possible combinations can be predefined in a list, and an index is signaled to indicate which combination in the list is to be used to perform LIC for coding the block.

[0139] 14 shows a flowchart outlining a process (1400) according to one embodiment of the present disclosure. The process (1400) can be used in a video encoder. In various embodiments, the process (1400) is performed by a processing circuit, such as a processing circuit performing the functions of a video encoder (103), a processing circuit performing the functions of a video encoder (303), etc. In some embodiments, the process (1400) is implemented with software instructions, and thus, the processing circuit performs the process (1400) when it executes the software instructions. The process begins at (S1401) and proceeds to (S1410).

[0140] At (S1410), it is determined to apply local illumination compensation (LIC) to the current block in the current picture.

[0141] In (S1420), parameters of the LIC model are derived according to a first template (or a first subset template) of the current block and a second template (or a second subset template) of a reference block in the reference picture. The position of the reference block is determined based on the motion vector of the current block. The first template includes a subset of reconstructed neighboring samples above and to the left of the current block, and the second template includes samples at the same positions relative to the subset of reconstructed neighboring samples.

[0142] It should be noted that in some embodiments of the present disclosure, LIC is allowed to use different template sizes and / or template positions. In one example, templates such as the first template (or the first subset template) can have different template sizes and / or template positions. In one embodiment, the top template, the left template, and both the left template and the top template may be selected for LIC. The second template (or the second subset template) corresponds to the first template and includes, for example, samples at the same positions as the first template.

[0143] At (S1430), the LIC model is applied to the current block according to the reference block to generate compensation samples for the current block.

[0144] In one embodiment, M rows may be used for the top template and / or N columns may be used for the left template, where both M and N are non-zero positive integer values. In one embodiment, syntax is signaled to indicate which template is selected for LIC. In one embodiment, the template selection is implicitly derived based on coding information, including but not limited to block size, block shape, block aspect ratio, and neighboring reconstructed samples.

[0145] In some examples, the first template includes a row of reconstructed neighboring samples immediately above the current block, and in some examples, the first template includes a column of reconstructed neighboring samples immediately to the left of the current block.

[0146] In some examples, the first template includes one or more rows of reconstructed neighboring samples above the current block, hi some examples, the first template includes one or more columns of reconstructed neighboring samples to the left of the current block.

[0147] In some examples, syntax is encoded into the coding bitstream for at least the current picture, the syntax indicating a first template for selection from a plurality of candidate templates.

[0148] In some examples, the first template for deriving parameters of the LIC model is determined according to at least one of the size of the current block, the shape of the current block, the aspect ratio of the current block, or the reconstructed neighboring samples.

[0149] The process then proceeds to (S1499) and ends.

[0150] The process 1400 may be adapted as appropriate. Steps of the process 1400 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0151] 15 shows a flowchart outlining a process (1500) according to one embodiment of the present disclosure. The process (1500) can be used in a video encoder. In various embodiments, the process (1500) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1500) is implemented with software instructions, and thus, the processing circuit performs the process (1500) when it executes the software instructions. The process begins at (S1501) and proceeds to (S1510).

[0152] At (S1510), coding information for a current block in a current picture is received from a coding video bitstream, where the coding information indicates applying local illumination compensation (LIC).

[0153] At (S1520), parameters of the LIC model are derived according to a first template (or a first subset template) of the current block and a second template (or a second subset template) of a reference block in the reference picture. The position of the reference block is determined based on a motion vector. The first template includes subsets of reconstructed neighboring samples above and to the left of the current block, and the second template includes samples at the same positions relative to the subset of reconstructed neighboring samples.

[0154] It should be noted that in some embodiments of the present disclosure, LIC is allowed to use different template sizes and / or template positions. In one example, templates such as the first template (or the first subset template) can have different template sizes and / or template positions. In one embodiment, the top template, the left template, and both the left template and the top template may be selected for LIC. The second template (or the second subset template) corresponds to the first template and includes, for example, samples at the same positions as the first template.

[0155] At (S1530), the LIC model is applied to the current block according to the reference block to generate compensation samples for the current block.

[0156] In one embodiment, M rows may be used for the top template and / or N columns may be used for the left template, where both M and N are non-zero positive integer values. In one embodiment, syntax is signaled to indicate which template is selected for LIC. In one embodiment, the template selection is implicitly derived based on coding information, including but not limited to block size, block shape, block aspect ratio, and neighboring reconstructed samples.

[0157] In some examples, the first subset template includes a row of reconstructed neighboring samples immediately above the current block, and in some examples, the first subset template includes a column of reconstructed neighboring samples immediately to the left of the current block.

[0158] In some examples, the first subset template includes one or more rows of reconstructed neighboring samples above the current block, hi some examples, the first subset template includes one or more columns of reconstructed neighboring samples to the left of the current block.

[0159] In some examples, a syntax is decoded, the syntax indicating a first template for selection from a plurality of candidate templates.

[0160] In some examples, the first subset template for deriving the parameters of the LIC model is determined according to at least one of the size of the current block, the shape of the current block, the aspect ratio of the current block, or the reconstructed neighboring samples.

[0161] The process then proceeds to (S1599) and ends.

[0162] The process 1500 may be adapted as appropriate. Steps of the process 1500 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0163] 16 shows a flowchart outlining a process (1600) according to one embodiment of the present disclosure. The process (1600) can be used in a video encoder. In various embodiments, the process (1600) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1600) is implemented with software instructions, and thus, the processing circuit performs the process (1600) when it executes the software instructions. The process begins at (S1601) and proceeds to (S1610).

[0164] At (S1610), it is determined to apply local illumination compensation (LIC) to the current block in the current picture.

[0165] At (S1620), a reference block in a reference picture is determined based on the motion vector of the current block.

[0166] At (S1630), samples in a first block are classified into at least a first class and a second class according to a classification criterion. The first block is one of a current block and a reference block, and samples in the second block are classified according to samples at the same positions in the first block so as to have the same class as the samples at the same positions in the first block, and the second block is another one of the current block and the reference block. In one example, the first block is the current block and the second block is the reference block. In another example, the first block is the reference block and the second block is the current block.

[0167] In (S1640), the template samples of the first block are classified into at least a first class and a second class according to a classification criterion, and the template samples of the second block are classified according to the corresponding co-located template samples of the first block so as to have the same class as the co-located samples of the first block.

[0168] In (S1650), a first parameter of a first LIC model is derived according to the first class template sample of the first block and the first class template sample of the second block.

[0169] In (S1660), second parameters of a second LIC model are derived according to the second class template sample of the first block and the second class template sample of the second block.

[0170] At (S1670), a first LIC model is applied to samples of a first class in the current block, and a second LIC model is applied to samples of a second class in the current block to generate compensated samples for the current block.

[0171] It should be noted that in some embodiments of the present disclosure, multiple models for LIC can be used. In one embodiment, samples of the reference block and reference block template are classified into multiple classes based on specific classification criteria. Then, samples of the current block and current block template are classified based on the class index of the co-located reference sample. For each class, a model is derived using samples of the current block template and reference block template associated with the same class index. Finally, predicted samples for each class are derived using the model and co-located reference samples associated with the same class. In one embodiment, the mean value of the reference block (and / or reference block template) is calculated, and the classification criterion for a sample is to compare the sample with this mean value. If the sample value is smaller than (or equal to or smaller than) this mean value, the sample is classified as the first class; otherwise, the sample is classified as the second class. In one embodiment, model derivation for each class may be implemented by using the autocorrelation matrix using the least mean square (LMS) method or LDL decomposition in CCLM / MM-CCLM. In one embodiment, multi-model LIC can replace single-model LIC without any syntax changes. In one example, a flag is signaled in a high-level syntax such as SPS, PPS, picture header, or slice header to indicate whether multi-model LIC is used. In another embodiment, a separate flag is signaled to indicate whether multi-model LIC or single-model LIC is selected when LIC is applied. In one embodiment, a gradient value for each sample is calculated, and a cumulative gradient value is calculated in a reference block (and / or reference block template). An average value of the cumulative gradient is calculated, and the classification criterion for a sample is to compare the sample with this average value. If the sample value is smaller than (or equal to or smaller than) this average value, the sample is classified as the first class; otherwise, the sample is classified as the second class.In one embodiment, when the average gradient value is lower than a predetermined threshold, the gradient-based multi-model is inferred to be invalid. In one embodiment, syntax signaling can be used for gradient-based multi-model LIC. In one embodiment, the gradient-based multi-model can be combined with the amplitude-based multi-model. By using a combination of these two multi-models, four different models are used for the coding block. The model selection for each sample can be based on the classification of its amplitude and gradient values.

[0172] In some examples, the first LIC model and the second LIC model have at least one different parameter value.

[0173] In some examples, an amplitude threshold is determined based on an average of the sample values ​​in the first block to classify the samples in the first block into at least a first class and a second class according to the classification criterion, and the samples are classified into the first class or the second class based on a comparison of the samples with the amplitude threshold.

[0174] In some examples, a gradient threshold is determined based on gradient values ​​of the samples in the first block to classify the samples in the first block into at least a first class and a second class according to the classification criterion, and the samples are classified into the first class or the second class based on a comparison of the gradient values ​​of the samples with the gradient threshold.

[0175] In some examples, to derive the first parameters of the first LIC model, at least one of a least mean squares operation and an autocorrelation matrix using LDL decomposition operation can be performed to derive the first parameters.

[0176] In some examples, syntax is encoded into a coding bitstream carrying at least the current picture, where the syntax indicates applying more than one LIC model to generate compensation samples for the current block.

[0177] The process then proceeds to (S1699) and ends.

[0178] The process 1600 may be adapted as appropriate. Steps of the process 1600 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0179] 17 shows a flowchart outlining a process (1700) according to one embodiment of the present disclosure. The process (1700) can be used in a video encoder. In various embodiments, the process (1700) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1700) is implemented with software instructions, and thus, the processing circuit performs the process (1700) when it executes the software instructions. The process begins at (S1701) and proceeds to (S1710).

[0180] At (S1710), coding information for a current block in a current picture is received from a coding video bitstream, where the coding information indicates applying local illumination compensation (LIC).

[0181] At (S1720), a reference block in a reference picture is determined based on the motion vector of the current block.

[0182] At (S1730), samples in a first block are classified into at least a first class and a second class according to a classification criterion. The first block is one of a current block and a reference block, and samples in the second block are classified according to samples at the same positions in the first block so as to have the same class as the samples at the same positions in the first block, and the second block is another one of the current block and the reference block. In one example, the first block is the current block and the second block is the reference block. In another example, the first block is the reference block and the second block is the current block.

[0183] In (S1740), the template samples of the first block are classified into at least a first class and a second class according to a classification criterion, and the template samples of the second block are classified according to the corresponding co-located template samples of the first block so as to have the same class as the co-located samples of the first block.

[0184] In (S1750), a first parameter of a first LIC model is derived according to the template sample of the first class of the first block and the template sample of the first class of the second block.

[0185] In (S1760), second parameters of a second LIC model are derived according to the second class template samples of the first block and the second class template samples of the second block.

[0186] At (S1770), a first LIC model is applied to samples of a first class in the current block, and a second LIC model is applied to samples of a second class in the current block to generate compensated samples for the current block.

[0187] It should be noted that in some embodiments of the present disclosure, multiple models for LIC can be used. In one embodiment, samples of the reference block and reference block template are classified into multiple classes based on specific classification criteria. Then, samples of the current block and current block template are classified based on the class index of the co-located reference sample. For each class, a model is derived using samples of the current block template and reference block template associated with the same class index. Finally, predicted samples for each class are derived using the model and co-located reference samples associated with the same class. In one embodiment, the mean value of the reference block (and / or reference block template) is calculated, and the classification criterion for a sample is to compare the sample with this mean value. If the sample value is smaller than (or equal to or smaller than) this mean value, the sample is classified as the first class; otherwise, the sample is classified as the second class. In one embodiment, model derivation for each class may be implemented by using the autocorrelation matrix using the least mean square (LMS) method or LDL decomposition in CCLM / MM-CCLM. In one embodiment, multi-model LIC can replace single-model LIC without any syntax changes. In one example, a flag is signaled in a high-level syntax such as SPS, PPS, picture header, or slice header to indicate whether multi-model LIC is used. In another embodiment, a separate flag is signaled to indicate whether multi-model LIC or single-model LIC is selected when LIC is applied. In one embodiment, a gradient value for each sample is calculated, and a cumulative gradient value is calculated in a reference block (and / or reference block template). An average value of the cumulative gradient is calculated, and the classification criterion for a sample is to compare the sample with this average value. If the sample value is smaller than (or equal to or smaller than) this average value, the sample is classified as the first class; otherwise, the sample is classified as the second class.In one embodiment, when the average gradient value is lower than a predetermined threshold, the gradient-based multi-model is inferred to be invalid. In one embodiment, syntax signaling can be used for gradient-based multi-model LIC. In one embodiment, the gradient-based multi-model can be combined with the amplitude-based multi-model. By using a combination of these two multi-models, four different models are used for the coding block. The model selection for each sample can be based on the classification of its amplitude and gradient values.

[0188] In some examples, the first LIC model and the second LIC model have at least one different parameter value.

[0189] In some examples, an amplitude threshold is determined based on an average of the sample values ​​in the first block to classify the samples in the first block into at least a first class and a second class according to the classification criterion, and the samples are classified into the first class or the second class based on a comparison of the samples with the amplitude threshold.

[0190] In some examples, a gradient threshold is determined based on gradient values ​​of the samples in the first block to classify the samples in the first block into at least a first class and a second class according to the classification criterion, and the samples are classified into the first class or the second class based on a comparison of the gradient values ​​of the samples with the gradient threshold.

[0191] In some examples, to derive the first parameters of the first LIC model, at least one of a least mean squares operation and an autocorrelation matrix using LDL decomposition operation can be performed to derive the first parameters.

[0192] In some instances, syntax is decoded, and the syntax indicates applying more than one LIC model to generate the compensation samples for the current block.

[0193] The process then proceeds to (S1799) and ends.

[0194] The process 1700 may be adapted as appropriate. Steps of the process 1700 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0195] 18 shows a flowchart outlining a process (1800) according to one embodiment of the present disclosure. The process (1800) can be used in a video encoder. In various embodiments, the process (1800) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1800) is implemented with software instructions, and thus, the processing circuit performs the process (1800) when it executes the software instructions. The process begins at (S1801) and proceeds to (S1810).

[0196] At (S1810), it is determined to apply local illumination compensation (LIC) to the current block in the current picture.

[0197] In step S1820, parameters of an LIC model are derived according to a first template of the current block and a second template of a reference block in the reference picture. The reference block is determined based on the motion vector of the current block. The LIC model is different from a linear model based on the amplitude of a single reference sample.

[0198] At (S1830), the LIC model is applied to the current block according to the reference block to generate compensation samples for the current block.

[0199] Note that in some embodiments, different equations can be used to model LIC, and these equations are different from those used in CCLM. In one embodiment, a nonlinear term can be used to form an equation such as Equation (8). In one example, an autocorrelation matrix using LDL decomposition is used to obtain the parameters α, α, and β. In one embodiment, an n-tap spatial domain filter is applied to the reference samples to form an equation such as Equation (9). In one embodiment, the n-tap filter shape can be, but is not limited to, a diamond, cross, square, etc., as shown in FIGS. 13A-13C. In one embodiment, syntax is signaled to indicate which n-tap filter shape is selected when an n-tap filter is used. In one embodiment, the spatial domain filter can be a symmetric filter with n / 2 filter coefficients if n is even, or (n+1) / 2 filter coefficients if n is odd. In one embodiment, a nonlinear term can also be added, as in Equation (10). In another embodiment, more than one nonlinear term can be used, and the number of nonlinear terms should be less than or equal to the number of spatial filter taps, n. In another embodiment, the autocorrelation matrix using LDL decomposition is calculated using the parameter α k and β. In another embodiment, an n-tap spatial domain filter is applied to the reference samples to form an equation such as equation (11). The autocorrelation matrix using the LDL decomposition is obtained by subtracting the parameters α k and k∈[0,n−1]. In one embodiment, the gradient values ​​of the reference samples may be used as input to form an equation such as equation (12). An autocorrelation matrix using LDL decomposition is used to obtain the parameters α and β. In one embodiment, the gradient values ​​of the reference samples may be added to equations (8)-(11). In one embodiment, pixel downsampling in the current block template and the reference block template may be used for model derivation. In one embodiment, syntax is signaled to indicate which modeling method is used for a coding block.

[0200] In some instances, the LIC model includes nonlinear terms.

[0201] In some examples, the LIC model includes an n-tap spatial domain filter, where n is greater than 1. The n-tap spatial domain filter has at least one of the following filter shapes: a cross, a diamond, and a square.

[0202] To derive the parameters of the LIC model, in some examples, a mean square error minimization is performed according to a first template of the current block and a second template of the reference block to calculate the filter coefficients of an n-tap spatial domain filter.

[0203] In some examples, the LIC model includes a slope term that is linearly based on the slope of the reference sample.

[0204] In some examples, syntax is encoded into a coding bitstream carrying at least the current picture, the syntax indicating a LIC model from a plurality of LIC model candidates.

[0205] The process then proceeds to (S1899) and ends.

[0206] The process 1800 may be adapted as appropriate. Steps of the process 1800 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0207] 19 shows a flowchart outlining a process (1900) according to one embodiment of the present disclosure. The process (1900) can be used in a video encoder. In various embodiments, the process (1900) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1900) is implemented with software instructions, and thus, the processing circuit performs the process (1900) when it executes the software instructions. The process begins at (S1901) and proceeds to (S1910).

[0208] At (S1910), coding information for a current block in a current picture is received from a coding video bitstream, where the coding information indicates applying local illumination compensation (LIC).

[0209] In (S1920), parameters of an LIC model are derived according to a first template of the current block and a second template of a reference block in the reference picture. The reference block is determined based on the motion vector of the current block. The LIC model is different from a linear model based on the amplitude of a single reference sample.

[0210] In (S1930), the LIC model is applied to the current block according to the reference block to generate compensation samples for the current block.

[0211] Note that in some embodiments, different equations can be used to model LIC, and these equations are different from those used in CCLM. In one embodiment, a nonlinear term can be used to form an equation such as Equation (8). In one example, an autocorrelation matrix using LDL decomposition is used to obtain the parameters α, α, and β. In one embodiment, an n-tap spatial domain filter is applied to the reference samples to form an equation such as Equation (9). In one embodiment, the n-tap filter shape can be, but is not limited to, a diamond, cross, square, etc., as shown in FIGS. 13A-13C. In one embodiment, syntax is signaled to indicate which n-tap filter shape is selected when an n-tap filter is used. In one embodiment, the spatial domain filter can be a symmetric filter with n / 2 filter coefficients if n is even, or (n+1) / 2 filter coefficients if n is odd. In one embodiment, a nonlinear term can also be added, as in Equation (10). In another embodiment, more than one nonlinear term can be used, and the number of nonlinear terms should be less than or equal to the number of spatial filter taps, n. In another embodiment, the autocorrelation matrix using LDL decomposition is calculated using the parameter α k and β. In another embodiment, an n-tap spatial domain filter is applied to the reference samples to form an equation such as equation (11). The autocorrelation matrix using the LDL decomposition is obtained by subtracting the parameters α k and k∈[0,n−1]. In one embodiment, the gradient values ​​of the reference samples may be used as input to form an equation such as equation (12). An autocorrelation matrix using LDL decomposition is used to obtain the parameters α and β. In one embodiment, the gradient values ​​of the reference samples may be added to equations (8)-(11). In one embodiment, pixel downsampling in the current block template and the reference block template may be used for model derivation. In one embodiment, syntax is signaled to indicate which modeling method is used for a coding block.

[0212] In some instances, the LIC model includes nonlinear terms.

[0213] In some examples, the LIC model includes an n-tap spatial domain filter, where n is greater than 1. The n-tap spatial domain filter has at least one of the following filter shapes: a cross, a diamond, and a square.

[0214] To derive the parameters of the LIC model, in some examples, a mean square error minimization is performed according to a first template of the current block and a second template of the reference block to calculate the filter coefficients of an n-tap spatial domain filter.

[0215] In some examples, the LIC model includes a slope term that is linearly based on the slope of the reference sample.

[0216] In some examples, the syntax is decoded, and the syntax indicates a LIC model from a plurality of LIC model candidates.

[0217] The process then proceeds to (S1999) and ends.

[0218] The process 1900 may be adapted as appropriate. Steps of the process 1900 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.

[0219] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 20 illustrates a computer system (2000) suitable for implementing certain embodiments of the disclosed subject matter.

[0220] Computer software may be coded using any suitable machine code or computer language that may be subject to assembly, compilation, linking, or similar mechanisms to create code that includes instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.

[0221] The instructions may be executed in various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0222] The components shown in Figure 20 for the computer system (2000) are exemplary in nature and are not intended to suggest any limitation regarding the scope of use or functionality of the computer software implementing the embodiments of the present disclosure, nor should the arrangement of components be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of the computer system (2000).

[0223] The computer system 2000 may include certain human interface input devices. Such human interface input devices may respond to input by one or more human users, for example, through tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), or olfactory input (not shown). Human interface input devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic vision).

[0224] The human interface input devices may include one or more of a keyboard (2001), a mouse (2002), a trackpad (2003), a touch screen (2010), a data glove (not shown), a joystick (2005), a microphone (2006), a scanner (2007), and a camera (2008) (only one of each is shown).

[0225] The computer system (2000) may also include certain human interface output devices. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (2010), data gloves (not shown), or joystick (2005), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (2009), headphones (not shown), etc.), visual output devices (e.g., screens (2010), including CRT, LCD, plasma, and OLED screens, each with or without touchscreen input capability and each with or without haptic feedback capability, some of which may provide two-dimensional visual output or greater than three-dimensional output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0226] The computer system (2000) may also include human-accessible storage devices and their associated media, such as optical media or similar media (2021), including CD / DVD ROM / RW (2020) with CD / DVD, thumb drives (2022), removable hard drives or solid state drives (203), legacy magnetic media such as tape and floppy disks (not shown), and special ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0227] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0228] The computer system (2000) may also include an interface (2054) to one or more communication networks (2055). The networks may be, for example, wireless, wired, or optical. The networks may further be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet and WLAN; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial TV; and vehicular and industrial networks including CANBus and the like. Certain networks typically require an external network interface adapter (e.g., a USB port on the computer system (2000)) attached to a particular general-purpose data port or peripheral bus (2049); others are typically integrated into the core of the computer system (2000) by attachment to a system bus (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system), as described below. Using any of these networks, the computer system (2000) can communicate with other entities. Such communications can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., from a particular CANbus to a particular CANbus device), or bidirectional to other computer systems, using, for example, local or wide-area digital networks. As noted above, specific protocols and protocol stacks can be used in each of these networks and network interfaces.

[0229] The above-mentioned human interface devices, human-accessible storage devices and network interfaces can be attached to the core (2040) of the computer system (2000).

[0230] The core (2040) may include one or more central processing units (CPUs) (2041), graphics processing units (GPUs) (2042), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (2043), task-specific hardware accelerators (2044), graphics adapters (2050), etc. These devices may be connected through a system bus (2048), along with read-only memory (ROM) (2045), random access memory (RAM) (2046), and internal mass storage (2047), such as an internal non-user-accessible hard drive or SSD. In some computer systems, the system bus (2048) is accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (2048) or through a peripheral bus (2049). In one example, a screen (2010) may be connected to the graphics adapter (2050). Peripheral bus architectures include PCI, USB, and the like.

[0231] The CPU (2041), GPU (2042), FPGA (2043), and accelerator (2044) can execute specific instructions that, in combination, can constitute the above-mentioned computer code. The computer code can be stored in a ROM (2045) or a RAM (2046). Also, temporary data can be stored in the RAM (2046), while permanent data can be stored, for example, in internal mass storage (2047). High-speed storage and retrieval from any of the memory devices can be enabled through the use of cache memory, which can be closely associated with one or more of the CPU (2041), GPU (2042), mass storage (2047), ROM (2045), RAM (2046), etc.

[0232] The computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those skilled in the computer software arts.

[0233] By way of example and not limitation, a computer system having the architecture (2000) and particularly the core (2040) can provide functionality as a result of the processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as introduced above, as well as media associated with the core's (2040) specific storage of a non-transitory nature, such as the core's internal mass storage (2047) or ROM (2045). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (2040). The computer-readable media can include one or more memory devices or chips according to particular needs. The software can cause the core (2040) and particularly the processor (including a CPU, GPU, FPGA, etc.) therein to perform particular processes or particular portions of particular processes described herein, including defining data structures stored in RAM (2046) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (2044)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software include logic, and vice versa, where appropriate. References to computer-readable media may encompass circuitry (such as an integrated circuit (IC)) that stores software for execution, circuitry embodying logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.

[0234] The use of "at least one" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof.

[0235] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. It will thus be appreciated that those skilled in the art will be able to devise various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.

Claims

1. 1. A method of video decoding, comprising: receiving coding information for a current block in a current picture from a coded video bitstream, the coding information indicating applying local illumination compensation (LIC) to the current block in the current picture; deriving parameters of an LIC model according to a first template of a current block and a second template of a reference block in a reference picture, the reference block being pointed to based on a motion vector, the first template including a subset of reconstructed neighboring samples above and to the left of the current block, and the second template including co-located samples with respect to the subset of reconstructed neighboring samples; applying the LIC model to the current block according to the reference block to generate a compensation sample for the current block; A method comprising:

2. The method of claim 1 , wherein the first template includes rows of the reconstructed neighboring samples adjacent above the current block.

3. The method of claim 1 , wherein the first template comprises a column of the reconstructed neighboring samples adjacent to the left of the current block.

4. The method of claim 1 , wherein the first template includes one or more rows of the reconstructed neighboring samples above the current block.

5. The method of claim 1 , wherein the first template comprises one or more columns of the reconstructed neighboring samples to the left of the current block.

6. The method of claim 1 , further comprising decoding syntax indicating the first template for selection from a plurality of candidate templates.

7. 2. The method of claim 1, further comprising: determining to use the first template to derive the parameters of the LIC model according to at least one of a size of the current block, a shape of the current block, an aspect ratio of the current block, or the reconstructed neighboring samples.

8. classifying samples in a first block into at least a first class and a second class according to a classification criterion, wherein the first block is one of the current block and the reference block, and samples in a second block are classified according to samples at the same position in the first block, and the second block is another one of the current block and the reference block; classifying template samples of the first block into at least the first class and the second class according to the classification criterion, wherein the template samples of the second block are classified according to corresponding co-located template samples of the first block; Deriving first parameters of a first LIC model according to the first class of template samples of the first block and the first class of template samples of the second block; deriving second parameters of a second LIC model according to the second class of template samples of the first block and the second class of template samples of the second block; applying the first LIC model to the first class of samples in the current block and applying the second LIC model to the second class of samples in the current block to generate compensated samples for the current block; The method of claim 1 further comprising:

9. The method of claim 8 , wherein the first LIC model and the second LIC model have at least one different parameter value.

10. classifying the samples in the first block into at least the first class and the second class according to the classification criterion, determining an amplitude threshold based on an average of the sample values ​​in the first block; classifying the sample into the first class or the second class based on a comparison of the sample with the amplitude threshold; The method of claim 8 further comprising:

11. classifying the samples in the first block into at least the first class and the second class according to the classification criterion, determining a gradient threshold based on gradient values ​​of samples in the first block; classifying the sample into the first class or the second class based on a comparison of the gradient value of the sample with the gradient threshold; The method of claim 8 further comprising:

12. The step of deriving the first parameters of the first LIC model includes:

9. The method of claim 8, further comprising deriving the first parameters of the first LIC model according to at least one of an autocorrelation matrix using a least mean squares operation and an LDL decomposition operation.

13. The method of claim 8 , further comprising: decoding syntax indicating applying more than one LIC model to generate the compensation samples for the current block.

14. The method of claim 1 , wherein the LIC model is different from a linear model based on a single reference sample.

15. The method of claim 14 , wherein the LIC model includes nonlinear terms.

16. The method of claim 14 , wherein the LIC model includes an n-tap spatial domain filter, where n is greater than 1.

17. 17. The method of claim 16, wherein the n-tap spatial domain filter has at least one of the following filter shapes: a cross, a diamond, and a square.

18. The step of deriving the parameters of the LIC model includes:

17. The method of claim 16, further comprising: calculating filter coefficients of the n-tap spatial domain filter by performing mean square error minimization according to the first template of the current block and the second template of the reference block.

19. 15. The method of claim 14, wherein the LIC model includes a gradient term that is linearly based on the gradient of a reference sample.

20. The method of claim 1 , further comprising decoding syntax indicating the LIC model from a plurality of candidates.

21. 1. An apparatus for video decoding, comprising a processing circuit, 21. Apparatus, wherein the processing circuitry is configured to perform a method according to any one of claims 1 to 20.

22. A computer program causing a computer to carry out the method of any one of claims 1 to 20.

23. 1. A method of video encoding, comprising: determining to apply local illumination compensation (LIC) to a current block in a current picture; determining parameters of an LIC model according to a first template of the current block and a second template of a reference block in a reference picture, the reference block being pointed to based on a motion vector, the first template including a subset of reconstructed neighboring samples above and to the left of the current block, and the second template including co-located samples with respect to the subset of reconstructed neighboring samples; applying the LIC model to the current block according to the reference block to generate a compensation sample for the current block; transmitting coding information of the current block in the current picture, the coding information indicating that the LIC is applied to the current block in the current picture; A method comprising:

Citation Information

Patent Citations

  • Multiple-model local illumination compensation

    US20190215522A1

  • Method and apparatus for video encoding or decoding

    US20200195976A1

  • Local illumination compensation in video coding

    US20200228796A1

  • Illumination compensation in video coding

    US20210266582A1

  • Method and apparatus for video coding

    US20220312004A1