Method, apparatus and program for local illumination compensation for biprediction
Local illumination compensation techniques address illumination variations in video coding by using weighted sums and offsets, enhancing compression efficiency and quality in both uni-predictive and bi-predictive scenarios.
Patent Information
- Application Number
- JP2025522103
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2026-01-15
AI Technical Summary
Existing video coding technologies face challenges in effectively addressing local illumination variations between current and reference blocks, particularly in bi-prediction scenarios, which can lead to suboptimal compression efficiency and quality.
Implement local illumination compensation (LIC) techniques that utilize weighted sums and offsets based on multiple parameter values, including non-linear terms, to enhance prediction accuracy in both uni-predictive and bi-predictive scenarios, using neighboring samples for parameter derivation.
Improves video coding efficiency by better modeling local illumination changes, leading to enhanced compression performance and quality in video encoding and decoding processes.
Smart Images

Figure 2026501434000001_ABST
Abstract
Description
[Technical Field]
[0001] Incorporation by Reference This application claims the benefit of priority to U.S. patent application Ser. No. 18 / 381,538, entitled "Local Illumination Compensation for Bi-Prediction," filed Oct. 18, 2023, which in turn claims the benefit of priority to U.S. Provisional Application Ser. No. 63 / 417,930, entitled "Local Illumination Compensation for Bi-Prediction," filed Oct. 20, 2022. The disclosures of the prior applications are incorporated herein by reference in their entireties.
[0002] Technical Field This disclosure describes embodiments generally related to video coding. [Background technology]
[0003] The background discussion provided herein is intended to generally set forth the context of the present disclosure. The work of the inventors identified in this application, to the extent that their work is described in this background section, as well as aspects of this specification that may not qualify as prior art as of the filing date, are not admitted expressly or impliedly as prior art to the present disclosure.
[0004] Image / video compression can help transmit image / video data across different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention [Problem to be solved by the invention]
[0005] Aspects of the present disclosure include methods and apparatus for video encoding / decoding. In some examples, the apparatus for video decoding includes a processing circuit. The processing circuit receives coded information for a current block in a current picture from a coded video bitstream, the coded information indicating applying local illumination compensation (LIC) to the current block according to at least a first reference block in a first reference picture. The processing circuit determines, for a sample in the current block, at least a first reference sample in the first reference block, the first reference sample being co-located with the sample in the current block. The processing circuit calculates a weighted sum and offset of multiple terms for LIC according to multiple parameter values for multiple parameters used in LIC. The multiple parameter values include at least a first weighting value for a first weighting applied to a nonlinear term of the first reference sample raised to the kth power, where k is a power value not equal to 1. The processing circuit reconstructs the sample in the current block according to the weighted sum. In some examples, the power value is 2.
[0006] In some examples, the processing circuit determines the parameter values for the parameters by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the parameters.
[0007] In some examples, the processing circuit decodes the first weighting value from the coded video bitstream. In some examples, the processing circuit decodes an index from the coded video bitstream, the index indicating the first weighting value. In some examples, the first weighting value is inherited from a neighboring block of the current block.
[0008] In some examples, the plurality of parameters includes the first weighting, the offset, and a weighting for at least a linear term. A processing circuit determines first values for a first subset of the plurality of parameters and second values for a second subset of the plurality of parameters by minimizing a prediction error between neighboring reconstructed samples of a current block and a prediction of the neighboring reconstructed samples by LIC using the first values for the first subset of the plurality of parameters and the second subset of the plurality of parameters. In one example, the processing circuit decodes the first values from the coded video bitstream. In another example, the processing circuit decodes at least an index from the coded video bitstream, the at least index indicating the first value. In another example, the first values for the first subset of the plurality of parameters are inherited from neighboring blocks of the current block.
[0009] In some examples, the first subset of the plurality of parameters includes at least linear weightings for linear terms, and the second subset of the plurality of parameters includes the offset and the non-linear terms.
[0010] In some examples, a first subset of the plurality of parameters includes the offset and a second subset of the plurality of parameters includes at least a linear term and a linear weighting for the non-linear term.
[0011] In some examples, the encoded information indicates applying bi-predictive LIC to the current block according to the first reference block in the first reference picture and a second reference block in a second reference picture. In some examples, the parameter values for the parameters include a first linear weighting value for a first linear weighting of a first linear term of the first reference sample and a second linear weighting value for a second linear weighting of a second linear term of a second reference sample co-located in the second reference block with respect to the sample in the current block.
[0012] In some examples, the parameter values for the parameters include first linear weighting values for first linear weightings of first linear terms for first samples in the first reference block, the first samples including the first reference sample and one or more neighboring samples of the first reference sample. The first samples may form any suitable shape, such as a cross, a vertical bar, a horizontal bar, a diamond, a rectangle, a diagonal line, etc.
[0013] In some examples, the encoded information indicates applying bi-predictive LIC to the current block according to the first reference block in the first reference picture and a second reference block in a second reference picture, and the parameter values for the parameters include a first linear weighting value for a first linear weighting of a first linear term for a first sample in the first reference block and a second linear weighting value for a second linear weighting of a second linear term for a second sample in the second reference block, the first sample including the first reference sample and one or more neighboring samples of the first reference sample, and the second sample including a second reference sample in the second reference block and one or more neighboring samples of the second reference sample.
[0014] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform a method for video decoding / encoding. [Brief explanation of the drawings]
[0015] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings.
[0016] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication facility (100).
[0017] [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder.
[0018] [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder.
[0019] [Figure 4] 1 shows a diagram for illustrating bi-predictive local illumination compensation (LIC) in some examples.
[0020] [Figure 5] 1 shows diagrams for illustrating LIC bi-prediction in some examples.
[0021] [Figure 6] 1 shows diagrams for illustrating LIC bi-prediction in some examples.
[0022] [Figure 7A] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7B] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7C] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7D] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7E] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7F] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7G] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7H]1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7I] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7J] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7K] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7L] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7M] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7N] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7O] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples. [Figure 7P] 1 is a diagram illustrating some sample shapes that can be used in LIC in some examples.
[0023] [Figure 8] 1 shows a flowchart outlining another process according to some embodiments of the present disclosure.
[0024] [Figure 9] 1 shows a flowchart outlining a process according to some embodiments of the present disclosure.
[0025] [Figure 10] 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0026] 1 illustrates a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of an application of the disclosed subject matter, which is a video encoder and video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, and storing compressed video on digital media, including CDs, DVDs, memory sticks, etc.
[0027] The video processing system (100) includes a capture subsystem (113), which may include a video source (101), such as a digital camera, generating a stream of uncompressed video pictures (102). In one example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102), shown as a thick line to emphasize its large amount of data compared to the encoded video data (104) (or encoded video bitstream), may be processed by an electronic device (120) including a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or encoded video bitstream), shown as a thin line to emphasize its small amount of data compared to the stream of video pictures (102), may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems (106) and (108) of FIG. 1, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include a video decoder (110), for example, within an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an outgoing stream of video pictures (111) that can be rendered on a display (112) (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) can be encoded according to some video encoding / compression standard.Examples of these standards include ITU-T Recommendation H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.
[0028] It should be noted that the electronic devices (120) and (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may include a video encoder (not shown).
[0029] 2 shows an exemplary block diagram of a video decoder (210). The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., a receiving circuit). The video decoder (210) can be used in place of the video decoder (110) in the example of FIG. 1.
[0030] The receiver (231) can receive one or more coded video sequences, e.g., included in a bitstream, to be decoded by the video decoder (210). In some embodiments, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of the other coded video sequences. The coded video sequences can be received from a channel (201), which can be a hardware / software link to a storage device that stores the encoded video data. The receiver (231) can receive the encoded video data along with other data, e.g., coded audio data and / or auxiliary data streams, which can be forwarded to respective usage entities (not shown). The receiver (231) can separate the coded video sequences from other data. To address network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter, "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory may be external to the video decoder (210) (not shown). In still other applications, there may be a buffer memory (not shown) external to the video decoder (210), for example, to deal with network jitter, and another buffer memory (215) internal to the video decoder (210), for example, to handle playback timing. If the receiver (231) receives data from a storage / forwarding device with sufficient bandwidth and controllability or from an isosynchronous network, the buffer memory (215) may be unnecessary or may be small. For use with best-effort packet networks such as the Internet, the buffer memory (215) may be necessary, may be relatively large, may be advantageously adaptively sized, and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (210).
[0031] The video decoder (210) may include a parser (220) that reconstructs symbols (221) from the coded video sequence. These symbol categories, as shown in FIG. 2, include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a rendering device (212) (e.g., a display screen) that is not an integral part of the electronic device (230) but may be coupled to the electronic device (230). The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) can parse and entropy decode the received coded video sequence. The coding of the coded video sequence may follow a variety of video coding techniques or standards, including variable-length coding, Huffman coding, arithmetic coding with or without context sensitivity, and the like. The parser (220) can extract from the coded video sequence a set of subgroup parameters for at least one subgroup of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroup can include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) can also extract information from the coded video sequence, such as transform coefficients, quantizer parameter values, and motion vectors.
[0032] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).
[0033] The reconstruction of the symbols (221) can involve several different units, depending on the type of coded video picture or part thereof (e.g., inter / intra picture, inter / intra block, etc.) and other factors. Which units are involved and how can be controlled by subgroup control information parsed from the coded video sequence by the parser (220). The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0034] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into several functional units, as described below. In a practical implementation operating under commercial constraints, many of these units will interact closely with each other and may be, at least partially, integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0035] The first unit is a scaler / inverse transform unit (251), which receives quantized transform coefficients and control information from the parser (220) as symbols (221), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks containing sample values that can be input to an aggregator (255).
[0036] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (255) may add, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0037] In other cases, the output samples of the scaler / inverse transform unit (251) may relate to an inter-coded, potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (253) may access a reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) associated with the block, these samples may be added by an aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches the prediction samples may be controlled by motion vectors, which are available to the motion-compensated prediction unit (253) in the form of symbols (221), which may have, for example, X, Y, and reference picture components. Motion compensation may also include interpolation of sample values fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, motion vector prediction mechanisms, and the like.
[0038] The output samples of the aggregator (255) can be subjected to various loop filtering techniques in a loop filtering unit (256). Video compression techniques can include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called a coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression can also respond to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and can also respond to previously reconstructed, loop-filtered sample values.
[0039] The output of the loop filter unit (256) can be a sample stream that can be output to a rendering device (212) and can also be stored in a reference picture memory (257) for use in future inter-picture prediction.
[0040] Certain coded pictures, once fully reconstructed, can be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting reconstruction of the next coded picture.
[0041] The video decoder (210) can perform decoding operations according to a given video compression technology or standard, such as ITU-T Recommendation H.265. The coded video sequence can conform to the syntax specified by the video compression technology or standard being used. This means that the coded video sequence conforms to both the syntax of the video compression technology or standard and the profile described in the video compression technology or standard. Specifically, the profile can select certain tools from all tools available in the video compression technology or standard as the only tools available under that profile. Also required for compliance is that the complexity of the coded video sequence must be within a range defined by the level of the video compression technology or standard. In some cases, the level constrains the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further constrained through a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0042] In some embodiments, the receiver (231) can receive additional (redundant) data along with the encoded video. The additional data may be included as part of the coded video sequence(s). The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) improvement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0043] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used in place of the video encoder (103) in the example of FIG. 1.
[0044] The video encoder (303) can receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of Figure 3) that can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320).
[0045] The video source (301) can provide a source video sequence to be encoded by the video encoder (303) in the form of a digital video sample stream, which can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) can be a storage device that stores prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. The video data can be provided as multiple individual pictures that, when viewed in sequence, give the impression of motion. The pictures themselves can be organized as a spatial array of pixels, each of which can contain one or more samples, depending on the sampling structure, color space, etc., in use. The following discussion focuses on samples.
[0046] According to some embodiments, the video encoder (303) may encode and compress pictures of a source video sequence into a coded video sequence (343) in real time, or under any other time constraints as needed. Enforcing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, as described below. Coupling is not shown for clarity. Parameters set by the controller (350) may include rate control-related parameters (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions for optimizing the video encoder (303) for a particular system design.
[0047] In some embodiments, the video encoder (303) is configured to operate in an encoding loop. As a very simplified explanation, in one example, the encoding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be encoded and reference picture(s)) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to that generated by a (remote) decoder. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-accurate results independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local and remote encoders. In other words, the predictive portion of the encoder "sees" the exact same sample values as the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronism (and the resulting drift if synchronism cannot be maintained, eg, due to channel errors) is also used in several related techniques.
[0048] The operation of the "local" decoder (333) may be the same as a "remote" decoder, such as the video decoder (210) already described in detail in connection with Figure 2. However, briefly referring also to Figure 2, because symbols are available and the encoding / decoding of symbols into an encoded video sequence by the entropy coder (345) and parser (220) may be lossless, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (333).
[0049] In some embodiments, decoder technology, with the exception of parsing / entropy decoding, present in a decoder is present in the same or substantially identical functional form in the corresponding encoder. Thus, the disclosed subject matter focuses on the operation of the decoder. A description of the encoder technology can be omitted, as it is the inverse of the decoder technology, which is comprehensively described. In certain areas, more detailed descriptions are provided below.
[0050] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as “reference pictures.” In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of reference pictures that may be selected as predictive references for the input picture.
[0051] The local video decoder (333) can decode the coded video data of a picture that may be designated as a reference picture based on the symbols generated by the source coder (330). The operation of the coding engine (332) can advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence may be a replica of the source video sequence, typically with some errors. The local video decoder (333) can replicate the decoding process that may be performed on the reference picture by the video decoder and store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has common content with the reconstructed reference picture that would be obtained by the far-end video decoder (in the absence of transmission errors).
[0052] The predictor (335) can perform the prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or types of metadata, such as reference picture motion vectors, block shapes, etc., that may serve as suitable prediction references for the new picture. The predictor (335) can operate on a sample block or pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture can have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0053] The controller (350) can manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0054] The output of all the functional units described above may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0055] The transmitter (340) can buffer the coded video sequence produced by the entropy coder (345) and prepare it for transmission over a communication channel (360), which can be a hardware or software link to a storage device that stores the encoded video data. The transmitter (340) can merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0056] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain coding picture type to each coded picture, which can affect the coding technique that can be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0057] Intra pictures (I pictures) can be coded and decoded without using other pictures in the sequence as a source of prediction. Some video coders allow various types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0058] Predictive pictures (P pictures) may be encoded and decoded using intra prediction or inter prediction, using motion vectors and reference indices to predict the sample values of each block.
[0059] Bidirectionally predicted pictures (B pictures) may be coded and decoded using intra- or inter-prediction, using two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-predictive pictures may use three or more reference pictures and associated metadata for the reconstruction of a single block.
[0060] A source picture is typically spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples each) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks, as determined by the coding assignment applied to the block's respective picture. For example, blocks of an I picture may be non-predictively coded or predictively coded with reference to previously coded blocks of the same picture (spatial prediction or intra-prediction). Pixel blocks of a P picture may be predictively coded via spatial prediction or via temporal prediction with reference to one previously coded reference picture. Blocks of a B picture may be predictively coded via spatial prediction or via temporal prediction with reference to one or two previously coded reference pictures.
[0061] The video encoder (303) may perform encoding operations in accordance with a predetermined video encoding technique or standard, such as ITU-T Recommendation H.265. In its operations, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0062] In some embodiments, the transmitter (340) can transmit additional data along with the encoded video. The source coder (330) can include such data as part of the coded video sequence. The additional data may include other forms of redundant data, such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0063] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. In one example, a particular picture being encoded / decoded, called the current picture, is divided into blocks. If a block in the current picture is similar to a reference block in a previously coded and still buffered reference picture in the video, the block in the current picture can be coded by a vector called a motion vector. The motion vector points to a reference block in the reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0064] In some embodiments, inter-picture prediction may use a bi-prediction technique. According to the bi-prediction technique, two reference pictures, such as a first reference picture and a second reference picture, are used, both of which precede a current picture in a video in decoding order (but may be past and future in display order, respectively). A block in the current picture may be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. A block may be predicted by a combination of the first and second reference blocks.
[0065] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0066] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed block-by-block. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU can be recursively quadtree-decomposed into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine a prediction type for the CU (e.g., inter prediction type or intra prediction type). The CU is divided into one or more prediction units (PUs) depending on temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels of 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0067] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In one embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0068] Aspects of this disclosure provide techniques that can be used with an inter-prediction technique called local illumination compensation (LIC), which can enable the use of LIC in bi-prediction.
[0069] Various inter-prediction modes can be used in video coding. For example, in VVC, for an inter-predicted CU, motion parameters can include MV(s), one or more reference picture indices, a reference picture list usage index, and additional information about certain coding features to be used for inter-predicted sample generation. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU can be associated with a PU, have no significant residual coefficients, and may not have coded motion vector deltas or MV differences (e.g., MVDs), or reference picture indices. A merge mode can be specified, in which motion parameters for the current CU are obtained from neighboring CU(s), including spatial and / or temporal candidates, and optionally additional information as introduced in VVC. The merge mode can be applied to inter-predicted CUs as well as skip mode. In one example, an alternative to merge mode is explicit transmission of motion parameters, where the MV(s), corresponding reference picture index and reference picture list usage flag for each reference picture list, and other information are explicitly signaled per CU.
[0070] In embodiments such as VVC, the VVC Test Model (VTM) reference software supports enhanced merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode, and more. It includes one or more sophisticated inter-prediction coding tools, including GPM (Gigabit Partitioning Mode).
[0071] In some examples, local illumination compensation (LIC) is used as an inter-prediction technique to model local illumination variations between a current block and its predicted block (also called a reference block) by using a linear function. The predicted block may be in a reference picture and pointed to by the motion vector (MV) of the current block. Parameters of the linear function may include a scale α and an offset β, and the linear function may be represented by α×p[x,y]+s to compensate for illumination changes, where p[x,y] indicates a reference sample at position [x,y] in the reference block (also called the predicted block), which is pointed to by the MV. In some examples, the scale α and offset s may be derived based on the template of the current block and the corresponding reference template of the reference block by using a least-squares method, and thus no signaling overhead is required except that a LIC flag may be signaled to indicate the use of LIC. The scale α and offset s derived based on the template of the current block may be referred to as a template-based parameter set.
[0072] In some examples, LIC is used for uni-predictive inter CUs. In some examples, intra-neighboring samples (neighboring samples predicted using intra prediction) of the current block may be used in LIC parameter derivation. In some examples, LIC is disabled for blocks with fewer than 32 luma samples. In some examples, for non-subblock modes (including non-affine modes), LIC parameter derivation is performed based on the template block samples of the current CU instead of the partial template block samples for the first top-left 16x16 unit. In some examples, LIC parameter derivation is performed based on partial template block samples, such as the partial template block samples for the first top-left 16x16 unit. In some examples, the template samples of the reference block are determined by using motion compensation (MC) using the MVs of the block without rounding to integer pixel precision.
[0073] It should be noted that the current design of LIC in ECM is only applied to uni-predicted blocks, which limits the coding performance of LIC. Aspects of the present disclosure provide various techniques for improving LIC performance, such as adding a non-linear term, enabling LIC in bi-prediction, and enabling the use of neighboring samples of co-located samples (also referred to as applying a filter with multiple taps to the co-located sample). In some examples, to add the non-linear term, the encoder / decoder respectively determine multiple parameter values for multiple parameters used in LIC, the multiple parameter values including at least a first weighting value for a first weighting of a k-th power non-linear term, where k is a power value and not equal to 1. In some examples, k may be any suitable integer or floating-point number. In some examples, k is greater than 1. In one example, k is a positive integer greater than or equal to 2. The encoder / decoder can determine, for a sample in a current block, at least a first co-located sample (also referred to as a first reference sample) that is co-located in a first reference block with respect to the sample in the current block, and calculate a weighted sum of multiple terms based on the first co-located sample, where the weighted sum includes a first weighting value applied to the k-th power of the first co-located sample. The sample in the current block can be reconstructed according to the weighted sum and an offset value for the offset.
[0074] Some aspects of this disclosure provide techniques for enabling LIC for bi-prediction, such that LIC-based prediction of a current block in a current picture can incorporate two reference blocks in respective reference pictures.
[0075] According to one aspect of the present disclosure, a sample in a current block in a current picture is predicted as a combination of prediction samples from two reference blocks along with offsets.
[0076] In some embodiments, the prediction of a sample in the current block is generated as a linearly weighted sum and offset of co-located reference samples from two reference blocks, as represented by equation (1).
number
[0077] In one embodiment, a linear weighting value (also referred to as a linear weighting value) and an offset value (also referred to as an offset value), i.e., α i (e.g., including α0 and α1) and s are derived using neighboring reconstructed samples of the current block and the reference block. In one example, a neighboring reconstructed sample (template) area may be used for the derivation. The derivation is to find an α that minimizes the prediction error in the neighboring reconstructed sample (template) area. iTo find the values of α and s, techniques such as least squares or least mean squares can be used. The template area can include reconstructed samples from the upper neighbor, the left neighbor, and / or the upper-left neighbor. In one example, the prediction of samples within the template area can be expressed using weighting and offset (α0, α1, and s) parameters (variables). Changing the weighting and / or offset can change the prediction error between the reconstructed samples within the template area and the prediction of samples within the template area using LIC. In one example, the derivation can find values of the weighting and offset (α0, α1, and s) that minimize the prediction error.
[0078] In another embodiment, a linear weighting value, i.e., α i (e.g., including the true linear weighting value used in LIC or an index indicating the true linear weighting value used in LIC) is signaled or inherited from a neighboring block, and the offset value s is calculated by multiplying the reconstructed samples of the neighbors of the current block and the reference block by α i The signaled values of and are used to derive the offset(s). In one example, the prediction of samples within the template area can be expressed using a linear weighting value and a parameter (variable) of offset(s). Changing the offset can change the prediction error between the reconstructed samples within the template area and the prediction of samples within the template area using LIC. In one example, the derivation can find the value of offset(s) that minimizes the prediction error.
[0079] In another embodiment, the offset value s (e.g., including the true offset value or an index indicating the offset value) is signaled or inherited from a neighboring block and is weighted by a linear weighting value, i.e., α iis derived using the reconstructed samples of the current block and the reference block neighborhood and the signaled value of s. In one example, the prediction of the samples in the template area can be expressed using the offset value s and weighting parameters (variables) α0 and α1. Changing the weighting can change the prediction error between the reconstructed samples in the template area and the prediction of the samples in the template area using LIC. In one example, the derivation can find the weighting values (α0 and α1) that minimize the prediction error.
[0080] In another embodiment, the linear weighting and offset values, i.e., α i and s (which may contain the true linear weighting and offset values, or one or more indices indicating the linear weighting and offset values) are all signaled or inherited from neighboring blocks.
[0081] In some embodiments, the prediction is generated as a linearly weighted sum and offset of multiple reference samples from two reference blocks, as represented by equation (2).
number
[0082] For example, coordinates are defined relative to the upper left corner of each block. If the upper left corner of a block is defined as coordinate (0,0), then p(x,y) represents the sample at coordinate (x,y) in the current block, p0(x',y') represents the sample at (x',y') in the first reference block in the first reference picture from reference list 0, p1(x',y') represents the sample at (x',y') in the second reference block in the second reference picture from reference list 1, α0(x',y') represents the weighting applied to p0(x',y'), and α1(x',y') represents the weighting applied to p1(x',y').
[0083] In one embodiment, a non-zero weighting α i The reference samples with (x,y) form a particular shape (denoted as S(x,y)) around the sample (x,y) co-located with the current sample to be predicted. Note that any suitable shape formed from multiple reference samples may be used.
[0084] In one example, the particular shape is a cross shape that includes five non-zero weighting factors in each of the list 0 and list 1 reference pictures.
[0085] FIG. 4 shows a diagram illustrating LIC bi-prediction in some examples. In FIG. 4, a current picture includes a current block (410). A first motion vector MV0 of the current block points to a first reference block (420) in a first reference picture (e.g., reference picture 0), and a second motion vector MV1 of the current block points to a second reference block (430) in a second reference picture (e.g., reference picture 1). For a sample (411) in the current block (410), the reference block (420) includes a co-located sample (also called a collocated sample or reference sample) (421), and the reference block (430) also includes a co-located sample (also called a collocated sample or reference sample) (431). In one example, the co-located sample corresponding to the sample (411) may be identified by the first motion vector MV0 and the second motion vector MV1. In one example, to predict sample (411), five samples (425) in a first reference block (420) that form a cross and five samples (435) in a second reference block (430) that form a cross are used, for example, according to equation (2), to generate a prediction for sample (411).
[0086] In another example, the particular shape is a vertical bar shape with three vertical taps (eg, three non-zero weighting coefficients in the vertical direction) in each of the list 0 and list 1 reference pictures.
[0087] FIG. 5 shows a diagram illustrating LIC bi-prediction in some examples. In FIG. 5, a current picture includes a current block (510). A first motion vector MV0 of the current block points to a first reference block (520) in a first reference picture (e.g., reference picture 0), and a second motion vector MV1 of the current block points to a second reference block (530) in a second reference picture (e.g., reference picture 1). For a sample (511) in the current block (510), the reference block (520) includes a co-located sample (also called a collocated sample or reference sample) (521), and the reference block (530) also includes a co-located sample (also called a collocated sample or reference sample) (531). In one example, the co-located sample corresponding to the sample (511) may be identified by the first motion vector MV0 and the second motion vector MV1. In one example, to predict sample (511), three samples (525) in a first reference block (520) forming a vertical bar and three samples (535) in a second reference block (530) forming a vertical bar are used, for example, according to equation (2), to generate a prediction for sample (511).
[0088] In another embodiment, the particular shape is a horizontal bar shape with three horizontal taps (eg, three non-zero weighting coefficients in the horizontal direction) in each of the List 0 and List 1 reference pictures.
[0089] FIG. 6 shows a diagram illustrating LIC bi-prediction in some examples. In FIG. 6, a current picture includes a current block (610). A first motion vector MV0 of the current block points to a first reference block (620) in a first reference picture (e.g., reference picture 0), and a second motion vector MV1 of the current block points to a second reference block (630) in a second reference picture (e.g., reference picture 1). For a sample (611) in the current block (610), the reference block (620) includes a co-located sample (also called a collocated sample or reference sample) (621), and the reference block (630) also includes a co-located sample (also called a collocated sample or reference sample) (631). In one example, the co-located sample corresponding to the sample (611) may be identified by the first motion vector MV0 and the second motion vector MV1. In one example, to predict sample (611), three samples (625) in a first reference block (620) forming a bar and three samples (635) in a second reference block (630) forming a bar are used, for example, according to equation (2), to generate a prediction of sample (611).
[0090] It should be noted that any suitable shape of samples in the reference pictures can be used in LIC bi-prediction.
[0091] Figures 7A-7P show several shapes of samples that can be used in LIC. The gray samples in each shape indicate the co-located (reference) samples in the reference picture that correspond to the samples to be predicted in the current picture. Figures 7A and 7N show, as examples, a shape called a cross shape. Figure 7B shows a shape called a vertical bar shape. Figure 7C shows a shape called a horizontal bar shape. Figures 7D-7E show a shape called a diagonal shape. Figures 7F-7I show a shape called a Z shape. Figure 7J shows a shape called an I shape. Figure 7K shows a shape called an H shape. Figures 7L and 7O show a shape called an X shape. Figure 7M shows a shape called a rectangular shape. Figure 7P shows a shape called a diamond shape.
[0092] In one embodiment, the linear weighting and offset values, i.e., α i (x', y') and s are derived using neighboring reconstructed samples of the current block and the reference block. In one example, a neighboring reconstructed sample (template) area can be used for the derivation. The derivation is performed using α i To find the values of s and s, techniques such as least squares or least mean squares can be used. The template area can include the reconstructed samples from the top neighbor, and / or the reconstructed samples from the left neighbor, and / or the reconstructed samples from the top-left neighbor.
[0093] In another embodiment, the linear weighting value, i.e., α i The (x',y') (or index to the value) is signaled or inherited from the neighboring blocks, and the offset value s is the sum of the reconstructed samples of the neighboring current and reference blocks and α i The signaled values of (x', y') are used to derive the
[0094] In another embodiment, the offset value s (or an index to the offset value) is signaled or inherited from a neighboring block and is used as the linear weighting value, i.e., α i (x',y') is derived using neighboring reconstructed samples of the current and reference blocks and the signaled value of s.
[0095] In another embodiment, the linear weighting value (or index to the value) and the offset value, i.e., α i (x',y') and s are all signaled or inherited from neighboring blocks.
[0096] In some embodiments, the prediction is generated as a nonlinear weighted sum and offset of co-located reference samples from two reference blocks, as represented by equation (3).
number
[0097] In one embodiment, the linear and non-linear weighting values and the offset value, i.e., α i , β i , and s are derived using neighboring reconstructed samples of the current block and the reference block. An example of the derivation is α , which minimizes the prediction error in the neighboring reconstructed sample (template) area. i , β iThe method of derivation is to use techniques such as least squares, least mean squares, etc. to find the values of , and s. In one example, a nearby reconstructed sample (template) area can be used for derivation. The derivation is to find the α that minimizes the prediction error in the nearby reconstructed sample (template) area. i , β i A least-squares method can be used to find values for α, β, and s. The template area can include reconstructed samples from an upper neighbor, a left neighbor, and / or an upper-left neighbor. In one example, predictions of samples within the template area can be expressed using linear weighting, nonlinear weighting, and offset parameters (variables) (α0, α1, β0, β1, and s). Changing the linear weighting, nonlinear weighting, and / or offsets can change the prediction error between the reconstructed samples within the template area and the prediction of samples within the template area using LIC. In one example, the derivation can find the weighting and offset values (α0, α1, β0, β1, and s) that minimize the prediction error.
[0098] In another embodiment, the linear weighting value, i.e., α i (or index to the value) is signaled or inherited from neighboring blocks and is weighted nonlinearly, i.e., β i and the offset value s is the reconstructed sample of the current block and the neighboring reference block, and α i The signaled values of β and β are used to derive the prediction error. In one example, the prediction of samples within the template area can be expressed using linear weighting values, nonlinear weighting parameters (variables), and offsets (β0, β1, and s). Changing the nonlinear weighting and / or offsets can change the prediction error between the reconstructed samples within the template area and the prediction of samples within the template area using LIC. In one example, the derivation can find values of the nonlinear weighting and offsets (β0, β1, and s) that minimize the prediction error.
[0099] In another embodiment, the offset value s (or index to the offset value) is signaled or inherited from neighboring blocks and is weighted by linear and non-linear weighting, i.e., α i and β i is derived using reconstructed samples from neighboring current and reference blocks and the signaled offset value of s. In one example, the prediction of samples within the template area can be expressed using offset values, linear weighting, nonlinear weighting, and offset parameters (variables) (α0, α1, β0, β1). Changing the linear and / or nonlinear weighting can change the prediction error between the reconstructed samples within the template area and the prediction of samples within the template area using LIC. In one example, the derivation can find the weighting values (α0, α1, β0, β1) that minimize the prediction error.
[0100] In another embodiment, the linear weighting value (or index to the value) and the offset value, i.e., α i , β i and s are all signaled or inherited from neighboring blocks.
[0101] In some embodiments, the prediction is generated as a combination of a linearly weighted sum of multiple reference samples from two reference blocks, a non-linearly weighted sum of co-located reference samples from two reference blocks, and an offset, for example, as represented by equation (4).
number
[0102] In one embodiment, a non-zero weighting α i The reference samples with (x,y) form a particular shape, i.e., S(x,y), around the sample (x,y) that is co-located with the current sample being predicted. Exemplary shapes include, but are not limited to, a cross, a diamond, a rectangle, a vertical shape, and a horizontal shape. Figures 7A-7P illustrate the use of non-zero weighting α i Some of the shapes that can be used for the reference sample with (x,y) are shown.
[0103] In one embodiment, the linear weighting value, i.e., α i (x',y'), the value of the nonlinear weighting, i.e., β i , and the offset value s are derived using neighboring reconstructed samples of the current block and the reference block. An example of the derivation is α i (x',y'),β i , the least mean square method is used to find the value of s.
[0104] In another embodiment, the linear weighting value, i.e., α i (x',y') (or index to the value) is signaled or inherited from neighboring blocks and is the value of the nonlinear weighting, i.e., β i and the offset value s is the reconstructed sample of the current block and the neighboring reference block, and α i The signaled values of (x', y') are used to derive the
[0105] In another embodiment, the offset value s (or an index to the offset value) is signaled or inherited from a neighboring block and is weighted by a linear weighting value, i.e., α i (x',y'), and the value of the nonlinear weighting, i.e., β i is derived using the reconstructed samples of the current and reference block neighborhoods and the signaled value of s.
[0106] In another embodiment, the linear weighting value, i.e., α i (x',y') (or index to the value), the value of the nonlinear weighting, i.e., β i , and the offset value s are all signaled or inherited from neighboring blocks.
[0107] 8 shows a flowchart outlining a process (800) according to one embodiment of the present disclosure. The process (800) may be used in a video decoder. In various embodiments, the process (800) is performed by a processing circuit, such as a processing circuit that performs the functions of the video decoder (110), a processing circuit that performs the functions of the video decoder (210), or the like. In some embodiments, the process (800) is implemented with software instructions, such that the processing circuit performs the process (800) when it executes the software instructions. The process begins at (S801) and proceeds to (S810).
[0108] At (S810), coded information of a current block in a current picture is received from a coded video bitstream, the coded information indicating applying local illumination compensation (LIC) to the current block according to at least a first reference block in a first reference picture, the first reference block being pointed to by a first motion vector of the current block.
[0109] In (S820), a plurality of parameter values for each of a plurality of parameters used in the LIC are determined. The plurality of parameter values include at least a first weighting value for a first weighting of a nonlinear term raised to the kth power, where k is a power value and is not equal to 1. In some examples, k may be any suitable integer or floating-point number. In some examples, k is greater than 1. In one example, k is a positive integer greater than or equal to 2.
[0110] In (S830), for a sample in the current block, at least a first reference sample is determined that is co-located in a first reference block with respect to the sample in the current block. The first reference sample is also referred to as a first co-located sample in one example.
[0111] At (S840), a weighted sum of a plurality of terms and an offset is calculated based on the first reference sample, the weighted sum including a first weighting value applied to the kth power of the first reference sample, as in Equation (3).
[0112] At (S850), the samples in the current block are reconstructed according to a weighted sum such as equation (3).
[0113] In some examples, the exponent value k is one of 2 or 3.
[0114] To determine the plurality of parameter values, in some examples, the plurality of parameter values for the plurality of parameters are determined by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the plurality of parameters.
[0115] In one example, the first weighting value is decoded directly from the coded video bitstream. In another example, an index is decoded from the coded video bitstream, and the index indicates the first weighting value from a plurality of weighting value candidates. In another example, the first weighting value is inherited from a neighboring block of the current block.
[0116] In some examples, the plurality of parameters includes a first weighting, an offset, and a weighting for at least a linear term. In one example, a first value for a first subset of the plurality of parameters is determined. A second value for a second subset of the plurality of parameters is determined by minimizing a prediction error between a neighboring reconstructed sample of the current block and a prediction of the neighboring reconstructed sample by LIC using the first value for the first subset of the plurality of parameters and the second subset of the plurality of parameters. In one example, the first value is decoded directly from the coded video bitstream. In another example, at least an index is decoded from the coded video bitstream, the at least index indicating the first value. In another example, the first value for the first subset of the plurality of parameters is inherited from a neighboring block of the current block.
[0117] In some examples, the first subset of the plurality of parameters includes linear weightings for the linear terms, and the second subset of the plurality of parameters includes offsets and non-linear terms.
[0118] In some examples, the first subset of the plurality of parameters includes offsets and the second subset of the plurality of parameters includes linear weightings for the linear and non-linear terms.
[0119] In some examples, the coded information indicates applying bi-predictive LIC to the current block according to a first reference block in a first reference picture and a second reference block in a second reference picture.
[0120] In some examples, the multiple parameter values for the multiple parameters include a first linear weighting value for a first linear weighting of a first linear term of a first reference sample and a second linear weighting value for a second linear weighting of a second linear term of a second reference sample co-located in a second reference block with respect to the sample in the current block.
[0121] In some examples, the parameter values for the parameters include first linear weighting values for first linear weightings of first linear terms for each first sample in the first reference block, the first sample including the first reference sample and one or more neighboring samples of the reference sample, and the first samples form any suitable shape, such as a cross, a vertical bar, a horizontal bar, a diamond, a rectangle, and a diagonal line.
[0122] In some examples, the encoded information indicates applying bi-predictive LIC to the current block according to a first reference block in a first reference picture and a second reference block in a second reference picture, and the multiple parameter values for the multiple parameters include a first linear weighting value for a first linear weighting of a first linear term for each first sample in the first reference block and a second linear weighting value for a second linear weighting of a second linear term for each second sample in the second reference block, the first sample including the first reference sample and one or more neighboring samples of the first reference sample, and the second sample including a second reference sample in the second reference block and one or more neighboring samples of the second reference sample.
[0123] Then, the process (800) proceeds to (S899) and ends.
[0124] Process 800 may be adapted as appropriate. Steps in process 800 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0125] 9 shows a flowchart outlining a process (900) according to one embodiment of the present disclosure. The process (900) may be used in a video encoder. In various embodiments, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of the video encoder (103), a processing circuit that performs the functions of the video encoder (303), or the like. In some embodiments, the process (900) is implemented with software instructions, such that the processing circuit performs the process (900) when it executes the software instructions. The process begins at (S901) and proceeds to (S910).
[0126] At (S910), it is determined to apply local illumination compensation (LIC) to a current block in a current picture according to at least a first reference block in a first reference picture for prediction of the current block.
[0127] In (S920), a plurality of parameter values for each of a plurality of parameters used in the LIC are determined. The plurality of parameter values include at least a first weighting value for a first weighting of a nonlinear term raised to the kth power, where k is a power value and is not equal to 1. In some examples, k may be any suitable integer or floating number. In some examples, k is greater than 1. In one example, k is a positive integer greater than or equal to 2.
[0128] At (S930), for a sample in the current block, at least a first reference sample co-located in a first reference block with respect to the sample in the current block is determined.
[0129] At (S940), a weighted sum of multiple terms is calculated based on the first reference sample, the weighted sum including a first weighting value applied to the kth power of the first reference sample, as in Equation (3).
[0130] In (S950), the samples in the current block are reconstructed according to the weighted sum and the offset value for the offset, as in equation (3).
[0131] In some examples, the exponent value k is one of 2 or 3.
[0132] To determine the parameter values, in some examples, the parameter values for the parameters are determined by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the parameters.
[0133] In one example, the first weighting value is encoded into the coded video bitstream. In another example, an index is encoded into the coded video bitstream, the index indicating the first weighting value from a plurality of weighting value candidates. In another example, the first weighting value is inherited from a neighboring block of the current block.
[0134] In some examples, the plurality of parameters includes a first weighting, an offset, and a weighting for at least a linear term. In one example, a first value for a first subset of the plurality of parameters is determined. A second value for a second subset of the plurality of parameters is determined by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the first value for the first subset of the plurality of parameters and the second subset of the plurality of parameters. In one example, the first value is encoded into the coded video bitstream. In another example, at least an index is encoded into the coded video bitstream, the at least index indicating the first value. In another example, the first value for the first subset of the plurality of parameters is inherited from a neighboring block of the current block.
[0135] In some examples, the first subset of the plurality of parameters includes linear weightings for the linear terms, and the second subset of the plurality of parameters includes offsets and non-linear terms.
[0136] In some examples, the first subset of the plurality of parameters includes offsets and the second subset of the plurality of parameters includes linear weightings for the linear and non-linear terms.
[0137] In some examples, it is determined to apply bi-predictive LIC to the current block according to a first reference block in a first reference picture and a second reference block in a second reference picture.
[0138] In some examples, the multiple parameter values for the multiple parameters include a first linear weighting value for a first linear weighting of a first linear term of a first reference sample and a second linear weighting value for a second linear weighting of a second linear term of a second reference sample co-located in a second reference block with respect to the sample in the current block.
[0139] In some examples, the parameter values for the parameters include first linear weighting values for first linear weightings of first linear terms for respective first samples in the first reference block, the first samples including the first reference sample and one or more neighboring samples of the first reference sample, and the first samples form any suitable shape, such as a cross, a vertical bar, a horizontal bar, a diamond, a rectangle, and a diagonal line.
[0140] In some examples, it is determined to apply bi-predictive LIC to a current block according to a first reference block in a first reference picture and a second reference block in a second reference picture, and the plurality of parameter values for the plurality of parameters include a first linear weighting value for a first linear weighting of a first linear term for each first sample in the first reference block and a second linear weighting value for a second linear weighting of a second linear term for each second sample in the second reference block, the first sample including the first reference sample and one or more neighboring samples of the first reference sample, and the second sample including a second reference sample in the second reference block and one or more neighboring samples of the second reference sample.
[0141] Then, the process proceeds to (S999) and ends.
[0142] Process 900 may be adapted as appropriate. Steps of process 900 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0143] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 10 illustrates a computer system (1000) suitable for implementing certain embodiments of the disclosed subject matter.
[0144] Computer software may be coded using any suitable machine code or computer language and may apply assembly, compilation, linking, or similar mechanisms to create code containing instructions that can be executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. directly, or through interpretation, microcode execution, etc.
[0145] The instructions may be executed on various types of computers or components thereof including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, and the like.
[0146] 10 for computer system 1000 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither the arrangement of components should be interpreted as having any dependency or requirement regarding any one or combination of components shown in the exemplary embodiment of computer system 1000.
[0147] The computer system (1000) may include certain human interface input devices that can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that do not necessarily involve direct human conscious input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still cameras), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0148] The input human interface devices may include one or more (only one of each is shown) of a keyboard (1001), a mouse (1002), a trackpad (1003), a touchscreen (1010), a data glove (not shown), a joystick (1005), a microphone (1006), a scanner (1007), and a camera (1008).
[0149] The computer system (1000) may also include some type of human interface output device. Such human interface output devices may stimulate one or more of the human user's senses, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1010), data gloves (not shown), or joystick (1005) (although haptic feedback devices may also function as input devices), audio output devices (e.g., speakers (1009), headphones (not shown)), visual output devices (e.g., screens (1010), including CRT screens, LCD screens, plasma screens, and OLED screens; each may or may not have touchscreen input capabilities, each may or may not have haptic feedback capabilities, some of which may output two-dimensional visual output or output in greater than three dimensions through means such as stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0150] The computer system (1000) may also include human-accessible storage devices and associated media, such as optical media including CD / DVD ROM / RW (1020) along with CD / DVD or similar media (1021), thumb drives (1022), removable hard drives or solid state drives (1023), legacy magnetic media such as tape and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices (not shown) such as security dongles, etc.
[0151] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.
[0152] The computer system (1000) may also include an interface (1054) to one or more communication networks (1055). Networks may be, for example, wireless, wired, or optical. Networks may further be local, wide-area, metropolitan, in-vehicle, and industrial, real-time, delay-tolerant, and the like. Examples of networks include Ethernet, WLAN, cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wired or wireless wide-area digital networks including cable, satellite, and terrestrial broadcast television; and in-vehicle and industrial networks including CANbus. Some networks typically require an external network interface adapter attached to some kind of general-purpose data port or peripheral bus (1049) (e.g., a USB port on the computer system (1000)). Others are typically integrated into the core of the computer system (1000) by attachment to a system bus, as described below (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1000) can communicate with other entities. Such communication may be unidirectional, receive-only (e.g., broadcast television), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional, for example, to other computer systems using local or wide-area digital networks. Each of these networks and network interfaces, as described above, may use certain protocols and protocol stacks.
[0153] The aforementioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1040) of the computer system (1000).
[0154] The core (1040) may include one or more central processing units (CPUs) (1041), graphics processing units (GPUs) (1042), specialized programmable processing units in the form of field programmable gate arrays (FPGAs) (1043), hardware accelerators for certain tasks (1044), graphics adapters (1050), etc. These devices may be connected through a system bus (1048), along with read-only memory (ROM) (1045), random access memory (1046), and internal mass storage devices (1047) such as internal non-user-accessible hard drives or solid-state drives (SSDs). In some computer systems, the system bus (1048) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the core's system bus (1048) or through a peripheral bus (1049). In one example, the screen 1010 can be connected to a graphics adapter 1050. Architectures for peripheral buses include PCI, USB, and the like.
[0155] The CPU (1041), GPU (1042), FPGA (1043), and accelerator (1044) may execute certain instructions that, in combination, may constitute the above-mentioned computer code. The computer code may be stored in ROM (1045) or RAM (1046). Temporary data may also be stored in RAM (1046), while persistent data may be stored, for example, in internal mass storage device (1047). Rapid storage and retrieval to any of the memory devices may be enabled through the use of cache memory, which may be closely associated with one or more of the CPU (1041), GPU (1042), mass storage device (1047), ROM (1045), RAM (1046), etc.
[0156] The computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.
[0157] By way of example and not limitation, the architecture (1000), and in particular a computer system having a core (1040), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be user-accessible mass storage, as discussed above, as well as media associated with some type of storage of the core (1040) that is non-transitory, such as a core-internal mass storage device (1047) or ROM (1045). Software implementing various embodiments of the present disclosure can be stored on such devices and executed by the core (1040). The computer-readable media can include one or more memory devices or chips, depending on particular needs. The software can cause the core (1040) and in particular the processor (including a CPU, GPU, FPGA, etc.) therein to perform certain processes or certain portions thereof described herein, including defining data structures stored in RAM (1046) and modifying such data structures according to software-defined processes. Additionally or alternatively, a computer system may provide functionality as a result of logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1044)), which may operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. Reference to software includes logic, and vice versa, as appropriate. Reference to a computer-readable medium may encompass circuitry (e.g., an integrated circuit (IC)) that stores software for execution, circuitry that embodies logic for execution, or both, as appropriate. The present disclosure encompasses any suitable combination of hardware and software.
[0158] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.
[0159] While this disclosure has described several exemplary embodiments, there are alterations, permutations, and various substitute equivalents that fall within the scope of this disclosure. Thus, those skilled in the art will be able to devise numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within its spirit and scope.
Claims
1. 1. A method of video decoding comprising: receiving coded information of a current block in a current picture from a coded video bitstream, the coded information indicating applying local illumination compensation (LIC) to the current block according to at least a first reference block in a first reference picture; determining, for a sample in a current block, at least a first reference sample in the first reference block, the first reference sample being co-located with the sample in the current block; calculating a weighted sum of a plurality of terms plus an offset for the LIC according to a plurality of parameter values for a plurality of parameters used in the LIC, the plurality of parameter values including at least a first weighting value for a first weighting applied to a nonlinear term of the first reference sample raised to the kth power, where k is a power value not equal to 1; and reconstructing the samples in the current block according to the calculated values. method.
2. The method of claim 1 , wherein the exponent value is 2.
3. determining the parameter values for the plurality of parameters by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the plurality of parameters; The method of claim 1.
4. decoding the first weighting values from the encoded video bitstream; decoding an index from the encoded video bitstream, the index indicating the first weighting value; and inheriting the first weighting value from a neighboring block of the current block; The method of claim 1 , further comprising at least one of:
5. The plurality of parameters includes the first weighting, the offset, and a weighting for at least a linear term, and the method includes: determining first values for a first subset of the plurality of parameters; determining second values for a second subset of the plurality of parameters, wherein the second values are determined by minimizing a prediction error between neighboring reconstructed samples of the current block and a prediction of the neighboring reconstructed samples by LIC using the first values for the first subset of the plurality of parameters and the second subset of the plurality of parameters; The method of claim 1.
6. Determining first values for a first subset of the plurality of parameters includes: decoding the first value from the encoded video bitstream; decoding at least an index from the encoded video bitstream, wherein at least the index indicates the first value; and Inheriting the first values for the first subset of the plurality of parameters from a neighboring block of the current block. The method of claim 5 , comprising at least one of:
7. 6. The method of claim 5, wherein the first subset of the plurality of parameters includes linear weighting values for linear terms, and the second subset of the plurality of parameters includes linear weighting values for the offset and the non-linear terms.
8. The method of claim 5 , wherein the first subset of the plurality of parameters includes the offsets and the second subset of the plurality of parameters includes linear weighting values for linear terms and non-linear terms.
9. 2. The method of claim 1, wherein the encoded information indicates applying bi-predictive LIC to a current block according to the first reference block in the first reference picture and a second reference block in a second reference picture.
10. 10. The method of claim 9, wherein the plurality of parameter values for the plurality of parameters includes a first linear weighting value for a first linear weighting of a first linear term of the first reference sample and a second linear weighting value for a second linear weighting of a second linear term of a second reference sample co-located in the second reference block with respect to the sample in the current block.
11. 2. The method of claim 1, wherein the plurality of parameter values for the plurality of parameters include first linear weighting values for first linear weightings of first linear terms for respective first samples in the first reference block, the first samples including the first reference sample and one or more neighboring samples of the first reference sample.
12. The first sample: cross; Vertical bar shape; Horizontal bar shape; Diamond shape; Rectangle; and Diagonal Linear 12. The method of claim 11, comprising at least one of:
13. 2. The method of claim 1 , wherein the encoded information indicates applying bi-predictive LIC to the current block according to the first reference block in the first reference picture and a second reference block in a second reference picture, and the parameter values for the parameters include a first linear weighting value for a first linear weighting of a first linear term for each first sample in the first reference block and a second linear weighting value for a second linear weighting of a second linear term for each second sample in the second reference block, the first sample including the first reference sample and one or more neighboring samples of the first reference sample, and the second sample including a second reference sample in the second reference block and one or more neighboring samples of the second reference sample.
14. Apparatus for video decoding, comprising processing circuitry configured to perform a method according to any one of claims 1 to 13.
15. A computer program which, when executed by a processor, causes the processor to carry out a method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and apparatus for lighting compensation of intra-predicted video
JP2011509639A
System and method for adaptively determining template size for illumination compensation
JP2019531029A