Improvements in coding information derivation for inter prediction

By employing a format rule to process video bitstreams with candidate lists and LIC information, the method addresses inefficiencies in existing video coding technologies, improving compression efficiency and quality of reconstructed video blocks.

CN120323023APending Publication Date: 2025-07-15TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005372.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-05-10
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing video encoding technology has problems such as inefficiency and high encoding complexity in inter-frame prediction, especially when processing complex video content, it is difficult to effectively utilize time and spatial redundancy for efficient compression.

Method used

Function-based prediction methods such as local illumination compensation (LIC), cross component linear model (CCLM), multi-model linear model (MMLM), convolutional cross component model (CCCM) and gradient linear model (GLM) are used to construct candidate lists, select appropriate decoding blocks and inherit their decoding information for inter-frame prediction, and use local and global illumination changes to predict video blocks.

Benefits of technology

It improves the efficiency and quality of video encoding, reduces the encoding complexity, and enhances the accuracy and compression performance of inter-frame prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120323023A_ABST
    Figure CN120323023A_ABST
Patent Text Reader

Abstract

In one embodiment, an apparatus for video decoding includes processing circuitry configured to construct a candidate list, the candidate list including one or more coded blocks associated with a current block. A coding block associated with the current block is a candidate of coding information that provides a function-based prediction method for prediction of the current block. The processing circuitry may select a particular coding block from the candidate list, the particular coding block being coded with particular coding information of the function-based prediction method. Further, the processing circuitry may inherit specific coding information of the function-based prediction method from the specific coding block to the current block, derive parameters of a function used in the function-based prediction method from the specific coding information, and reconstruct samples of at least the current block from the parameters of the function.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application claims the benefit of the priority of U.S. Provisional Application No. 63 / 546,916, "Improvement Of Coded Information Derivation For Inter Prediction", filed on November 1, 2023, which is hereby incorporated by reference in its entirety. Technical Field

[0002] The present disclosure describes aspects generally related to video coding. Background Art

[0003] The background art description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work described in this background art section, the work of the presently named inventors, and aspects of the description that may not otherwise be considered prior art at the time of filing are neither expressly nor implicitly admitted to be prior art against the present disclosure.

[0004] Image / video compression can help to transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video decoder techniques can compress video based on spatial redundancy and temporal redundancy. In an example, a video decoder can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from a reconstructed current picture for sample prediction. In another example, a video decoder can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can utilize motion compensation to predict samples in a current picture based on a previously reconstructed picture. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0005] Aspects of the present disclosure include bitstreams, methods, and devices for video coding / decoding. In some examples, a device for video coding / decoding includes processing circuitry.

[0006] Some aspects of the present disclosure provide a method for processing visual media data. The method includes: processing a bitstream of visual media data according to formatting rules. The bitstream includes decoding information of one or more pictures, and the one or more pictures include a current block in a current picture. The formatting rules specify constructing a candidate list including one or more decoded blocks associated with the current block, and the decoded blocks associated with the current block are candidates for providing local illumination compensation (LIC) decoding information for prediction of the current block. The LIC decoding information includes at least one of a control flag of the LIC, a template type of the LIC, and a model type of the LIC. The formatting rules also specify selecting a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the LIC. For example, the first decoding information of the LIC includes at least one of a first control flag of the LIC, a first template type of the LIC, and a first model type of the LIC. In addition, the formatting rules specify: inheriting the first decoding information of the LIC from the first decoded block to the current block; deriving parameters of the LIC according to the first decoding information; and reconstructing samples of at least the current block according to the parameters of the LIC.

[0007] Some aspects of the present disclosure provide an apparatus for video decoding. The apparatus includes processing circuitry configured to construct a candidate list including one or more decoded blocks associated with a current block. The decoded blocks associated with the current block are candidates for providing decoding information of a function-based prediction method for prediction of the current block. The function-based prediction method uses a function having parameters derived based on a template of the current block. The processing circuitry may select a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the function-based prediction method. In addition, the processing circuitry may inherit the first decoding information of the function-based prediction method from the first decoded block to the current block, derive parameters of the function used in the function-based prediction method according to the first decoding information, and reconstruct samples of at least the current block according to the parameters of the function.

[0008] In some examples, the one or more decoded blocks include at least one of an adjacent decoded block, a non-adjacent decoded block, and a decoded block whose decoding information is in a buffer.

[0009] In some examples, the function-based prediction method is one of the following: local illumination compensation (LIC), cross-component linear model (CCLM), multi-model linear model (MMLM), convolutional cross-component model (CCCM), and gradient linear model (GLM).

[0010] In some examples, the function-based prediction method is an inter-frame prediction method, and the processing circuitry is configured to select a first decoded block from a candidate list based on the inter-frame prediction mode information of the first decoded block and the current block. In an example, the inter-frame prediction mode information includes at least one of a prediction mode, a prediction direction, a reference list, and a reference index.

[0011] In some examples, the processing circuitry is configured to select the first decoded block from the candidate list when the current block and the first decoded block share the same prediction mode, the same prediction direction, the same reference list, and the same reference index.

[0012] In some examples, the processing circuitry is configured to select the first decoded block from the candidate list when the current block and the first decoded block share the same prediction mode.

[0013] In some examples, the processing circuitry is configured to select the first decoded block from the candidate list when the current block and the first decoded block share the same prediction mode and the same reference picture list.

[0014] In some examples, the processing circuitry is configured to select the first decoded block from the candidate list when the current block is a uni-directional prediction block using a first reference list and the first reference list is available at the first decoded block. In an example, the processing circuitry is configured to inherit only the first decoded information of the inter-frame prediction method of the first reference list from the first decoded block to the current block.

[0015] In some examples, the processing circuitry is configured to: select the first decoded block from the candidate list when the current block is a uni-directional prediction block using a first reference list and a first reference index, and the first decoded block is a bi-directional prediction block using the used first reference list and the first reference index. In an example, the processing circuitry is configured to inherit only the first decoded information of the inter-frame prediction method at the first reference list and the first reference index from the first decoded block to the current block.

[0016] In some examples, the current block is a uni-directional prediction block and the current block uses a first reference list. The processing circuitry is configured to: when the first reference list is available at the first decoded block, inherit the first decoded information of the inter-frame prediction method at the first reference list from the first decoded block to the current block; and when the first reference list is not available at the first decoded block, inherit the first decoded information of the inter-frame prediction method at the second reference list from the first decoded block to the current block.

[0017] In some examples, the current block is a bi - directional prediction block, and the processing circuitry is configured to: inherit first decoding information of an inter - prediction method at a reference list to the current block when the reference list of the current block is available at a first decoded block; and inherit default decoding information of the inter - prediction method to the current block when the reference list of the current block is not available at the first decoded block. In an example, the default decoding information includes predefined values. In another example, the default decoding information includes a syntax element signaled at one of a sequence level, a picture level, a slice level, a coding tree unit (CTU) level, and a coding unit (CU) level.

[0018] In some examples, the function - based prediction method is local illumination compensation (LIC), and the first decoding information includes at least one of a control flag of LIC, a template type of LIC, and a model type of LIC.

[0019] Some aspects of the present disclosure provide a method for video coding. The method includes determining to use an inter - prediction method for prediction of a current block in a current picture, the inter - prediction method being a function - based prediction method having parameters derived from a template of the current block. The method further includes constructing a candidate list including one or more decoded blocks associated with the current block, the decoded blocks associated with the current block being candidates that provide decoding information of the inter - prediction method for prediction of the current block. The method further includes selecting a first decoded block from the candidate list, the first decoded block being decoded with first decoding information of the inter - prediction method. In addition, the method includes inheriting the first decoding information of the inter - prediction method from the first decoded block to the current block; deriving parameters of a function used in the inter - prediction method according to the first decoding information; and reconstructing samples of the current block according to the parameters of the function.

[0020] In some examples, the one or more decoded blocks include at least one of an adjacent decoded block, a non - adjacent decoded block, and a decoded block whose decoding information is in a buffer. In an example, the function - based prediction method is local illumination compensation (LIC).

[0021] Aspects of the present disclosure also provide an apparatus for video coding. The apparatus for video coding includes processing circuitry configured to implement any of the methods for video coding described.

[0022] Aspects of the present disclosure also provide a method for video decoding. The method includes any of the methods implemented by an apparatus for video decoding.

[0023] Aspects of the present disclosure also provide a non - transitory computer - readable medium storing instructions that, when executed by a computer, cause the computer to execute any of the methods for video decoding / encoding described. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Other features, qualities, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0025] Figure 1 is a schematic diagram showing an example of a block diagram of a communication system (100).

[0026] Figure 2 is a schematic diagram showing an example of a block diagram of a decoder.

[0027] Figure 3 is a schematic diagram showing an example of a block diagram of an encoder.

[0028] Figure 4 shows the positions of spatial merge candidates according to an embodiment of the present disclosure.

[0029] Figure 5 shows candidate pairs for redundancy checks considered for spatial merge candidates according to an embodiment of the present disclosure.

[0030] Figure 6 shows an example motion vector scaling for temporal merge candidates.

[0031] Figure 7 shows an example candidate position of a temporal merge candidate for a current block.

[0032] Figure 8 shows a diagram of a template in some examples.

[0033] Figure 9 shows a flowchart outlining a decoding process according to some aspects of the present disclosure.

[0034] Figure 10 shows a flowchart outlining an encoding process according to some aspects of the present disclosure.

[0035] Figure 11 is a schematic diagram of a computer system according to an aspect. DETAILED DESCRIPTION

[0036] Figure 1 shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of the application of a video encoder and a video decoder in a streaming environment. The disclosed subject matter can be equivalently applied to other video-supported applications, including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media such as CDs, DVDs, memory sticks, etc.

[0038] The video processing system (100) includes a capture subsystem (113), which can include a video source (101), such as a digital imaging device, that creates a video picture stream (102), such as an uncompressed video picture stream. In an example, the video picture stream (102) includes samples taken by the digital imaging device. The video picture stream (102) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (104) (or decoded video bitstream). The video picture stream (102) can be processed by an electronic device (120) coupled to the video source (101) and including a video encoder (103). The video encoder (103) can include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower data volume when compared to the video picture stream (102). The encoded video data (104) can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 the client subsystems (106) and (108) in

[0039] Note that the electronic devices (120) and (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can also include a video encoder (not shown).

[0040] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) can be included in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., receiving circuitry). The video decoder (210) can be used instead of Figure 1 the video decoder (110) in the example.

[0041] A receiver (231) may receive one or more decoded video sequences to be decoded by a video decoder (210), such as included in a bitstream. In an aspect, one decoded video sequence is received at a time, where the decoding of each decoded video sequence is independent of the decoding of other decoded video sequences. The decoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data together with other data, such as decoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not depicted). The receiver (231) may separate the decoded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory may be external to the video decoder (210) (not depicted). In yet some other applications, a buffer memory (not depicted) may exist external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may additionally exist inside the video decoder (210) to, for example, handle presentation timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be small. For use on a best-effort packet network such as the Internet, a buffer memory (215) may be required, which may be relatively large and may advantageously have an adaptive size and may be implemented at least partially in an operating system or a similar element (not depicted) external to the video decoder (210).

[0042] The video decoder (210) may include a parser (220) to reconstruct symbols (221) according to the decoded video sequences. The categories of these symbols include: information for managing the operation of the video decoder (210), and possibly information for controlling a rendering device such as a rendering device (212) (e.g., a display screen), which is not part of the electronic device (230) but may be coupled to the electronic device (230), as Figure 2As shown. The control information for the rendering device may be in the form of a supplementary enhancement information (SEI) message or a video usability information (VUI) parameter set segment (not depicted). The parser (220) may perform parsing / entropy decoding on the received decoded video sequence. The decoding of the decoded video sequence may be performed according to a video decoding technology or a video decoding standard and may follow various principles, including variable length decoding, Huffman coding, arithmetic decoding with or without context sensitivity, etc. The parser (220) may extract subgroup parameter sets for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the subgroup. The subgroup may include a group of pictures (GOP), a picture, a tile, a slice, a macroblock, a coding unit (CU), a block, a transform unit (TU), a prediction unit (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the decoded video sequence.

[0043] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0044] Depending on the type of the decoded video picture or a part thereof (e.g., an inter picture and an intra picture, an inter block and an intra block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the manner of involvement may be controlled by subgroup control information parsed by the parser (220) from the decoded video sequence. For clarity, this subgroup control information flow between the parser (220) and the following multiple units is not depicted.

[0045] In addition to the functional blocks already mentioned, the video decoder (210) may be conceptually subdivided into multiple functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated into each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.

[0046] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) may output a block including sample values, and the block may be input into the aggregator (255).

[0047] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed part of the current picture. Such predictive information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses the reconstructed information around it obtained from the current picture buffer (258) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information already generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0048] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to an inter-coded and possibly motion-compensated block. In such a case, the motion compensation prediction unit (253) may access the reference picture memory (257) to obtain samples for prediction. After motion compensating the obtained samples according to the symbols (221) belonging to the block, these samples may be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The address at which the motion compensation prediction unit (253) in the reference picture memory (257) obtains the prediction samples may be controlled by a motion vector, which is available to the motion compensation prediction unit (253) in the form of symbols (221) that may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (257) when using sub-sample accurate motion vectors, a motion vector prediction mechanism, and the like.

[0049] The output samples of the aggregator (255) may be subject to various loop filtering techniques in the loop filter unit (256). The video compression technique may include in-loop filter techniques that are controlled by parameters included in the decoded video sequence (also referred to as the decoded video bitstream) as symbols (221) from the parser (220) and that are available to the loop filter unit (256). Video compression may also respond to meta-information obtained during decoding of a previous part (in decoding order) of the decoded picture or decoded video sequence, as well as to previously reconstructed and loop-filtered sample values.

[0050] The output of the loop filter unit (256) may be a sample stream, which may be output to the rendering device (212) and stored in the reference picture memory (257) for use in future inter-picture prediction.

[0051] Once fully reconstructed, some decoded pictures can be used as reference pictures for future prediction. For example, once the decoded picture corresponding to the current picture is fully reconstructed and the decoded picture has been identified as a reference picture (e.g., by a parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting the reconstruction of subsequent decoded pictures.

[0052] The video decoder (210) may perform decoding operations according to standards such as ITU-T H.265 Recommendations or predetermined video compression techniques. In the sense that the decoded video sequence follows both the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard, the decoded video sequence can conform to the syntax specified by the used video compression technique or standard. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as tools that can only be used under that profile. For compliance, it is also required that the complexity of the decoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in millions of samples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further restricted by the hypothetical reference decoder (HRD) specification and the metadata of the HRD buffer management signaled in the decoded video sequence.

[0053] In an aspect, the receiver (231) may receive additional (redundant) data together with the encoded video. The additional data may be included as part of the decoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0054] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used instead of Figure 1 the video encoder (103) in the example.

[0055] A video encoder (303) may receive video samples from a video source (301) that is not part of the electronic device (320) in the example. The video source (301) may capture video images to be decoded by the video encoder (303). In another example, the video source (301) is part of the electronic device (320). Figure 3 The video source (301) may provide a source video sequence in the form of a stream of digital video samples to be decoded by the video encoder (303). The stream of digital video samples may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, …), any color space (e.g., BT.601 Y CrCb, RGB, …), and any suitable sampling structure (e.g., Y CrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared video. In a video conferencing system, the video source (301) may be a camera device that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that convey motion when viewed in sequence. The pictures themselves may be organized as a spatial array of pixels, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0056] According to an aspect, the video encoder (303) may decode and compress pictures of the source video sequence into a decoded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate decoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to other functional units. For simplicity, the couplings are not depicted. Parameters set by the controller (350) may include rate control related parameters (picture skip, quantizer, λ value of rate distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured to have other suitable functions that are specific to the video encoder (303) optimized for a particular system design.

[0057]

[0058] ​In some aspects, the video encoder (303) is configured to operate in a decoding loop. As a highly simplified description, in an example, the decoding loop may include a source decoder (330) (e.g., responsible for creating symbols such as a symbol stream based on an input picture to be decoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to what a (remote) decoder would also create. The reconstructed sample stream (sample data) is input to the reference picture memory (334). Since the decoding of the symbol stream results in a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact between the local encoder and the remote encoder. In other words, the reference picture samples that the prediction part of the encoder "sees" are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. The basic principle of this reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, for example, due to channel errors) is also used in some related technologies.

[0059] The operation of the "local" decoder (333) can be the same as that of the "remote" decoder, such as the video decoder (210) that has been described in detail above in conjunction with Figure 2 However, briefly referring also to Figure 2 , since the symbols are available and the encoding of the symbols into the decoded video sequence by the entropy decoder (345) and the decoding of the symbols by the parser (220) can be lossless, the entropy decoding part of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0060] In an aspect, decoder techniques other than the parsing / entropy decoding present in the decoder exist in the corresponding encoder in the same or substantially the same functional form. Thus, the disclosed subject matter focuses on decoder operations. Since the encoder techniques are the opposite of the fully described decoder techniques, the description of the encoder techniques can be simplified. More detailed descriptions are provided in some places below.

[0061] During operation, in some examples, the source decoder (330) may perform motion-compensated predictive decoding that predicts an input picture by referring to one or more previously decoded pictures designated as "reference pictures" from a video sequence. In this way, the decoding engine (332) decodes the difference between a pixel block of the input picture and a pixel block of the reference picture that can be selected as a prediction reference for the input picture.

[0062] The local video decoder (333) can decode the decoded video data of pictures that can be designated as reference pictures based on the symbols created by the source decoder (330). The operation of the decoding engine (332) can advantageously be lossy processing. When the decoded video data can be decoded at the video decoder ( Figure 3 not shown in the figure), the reconstructed video sequence can generally be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that the video decoder can perform on the reference pictures, and can store the reconstructed reference pictures in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference pictures, which has the same content (without transmission errors) as the reconstructed reference pictures that will be obtained by the remote video decoder.

[0063] The predictor (335) can perform a prediction search for the decoding engine (332). That is, for a new picture to be decoded, the predictor (335) can search in the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata that can be used as a suitable prediction reference for the new picture, such as reference picture motion vectors, block shapes, etc. The predictor (335) can operate block by block based on sample blocks to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture can have prediction references extracted from multiple reference pictures stored in the reference picture memory (334).

[0064] The controller (350) can manage the decoding operations of the source decoder (330), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0065] The outputs of all the above-mentioned functional units can undergo entropy coding in the entropy coder (345). The entropy coder (345) converts these symbols into a decoded video sequence by applying lossless compression to the symbols generated by various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0066] The transmitter (340) can buffer the decoded video sequence created by the entropy coder (345) to prepare for transmission via the communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge the decoded video data from the video encoder (303) with other data to be transmitted, such as decoded audio data and / or an auxiliary data stream (source not shown).

[0067] The controller (350) may manage the operation of the video encoder (303). During decoding, the controller (350) may assign a specific decoded picture type to each decoded picture, which may affect the decoding techniques that can be applied to the corresponding picture. For example, pictures may typically be assigned to one of the following picture types:

[0068] An intra picture (I picture) may be decoded and decoded without using any other picture in the sequence as a prediction source. Some video decoders allow different types of intra pictures, including for example Independent Decoder Refresh ("IDR") pictures.

[0069] A predictive picture (P picture) may be decoded and decoded using intra prediction or inter prediction that utilizes motion vectors and reference indices to predict the sample values of each block.

[0070] A bi-predictive picture (B picture) may be decoded and decoded using intra prediction or inter prediction that utilizes two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures may use more than two reference pictures and associated metadata for the reconstruction of a single block.

[0071] Source pictures may typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples respectively), and decoded on a block-by-block basis. These blocks may be predictively decoded with reference to other (already decoded) blocks, which are determined by the decoding assignment applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively decoded, or may be predictively decoded (spatial prediction or intra prediction) with reference to already decoded blocks of the same picture. Pixel blocks of a P picture may be predictively decoded with reference to one previously decoded reference picture via spatial prediction or via temporal prediction. Blocks of a B picture may be predictively decoded with reference to one or two previously decoded reference pictures via spatial prediction or via temporal prediction.

[0072] The video encoder (303) may perform decoding operations according to a predetermined video decoding technique or standard such as ITU-T Recommendation H.265. In its operation, the video encoder (303) may perform various compression operations, including predictive decoding operations that exploit temporal and spatial redundancies in the input video sequence. Accordingly, the decoded video data may conform to the syntax specified by the video decoding technique or standard being used.

[0073] In an aspect, the transmitter (340) may transmit additional data along with the encoded video. The source decoder (330) may include such data as part of the decoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set segments, etc.

[0074] Video may be captured as a plurality of source pictures (video pictures) in a time sequence. Intra-picture prediction (commonly abbreviated as intra prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In an example, a particular picture in encoding / decoding, which is referred to as the current picture, is partitioned into blocks. In a case where a block in the current picture is similar to a reference block in a reference picture that has been previously decoded and is still buffered in the video, the block in the current picture may be decoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in a case where multiple reference pictures are used, the motion vector may have a third dimension identifying the reference picture.

[0075] In some aspects, bidirectional prediction techniques may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future respectively in display order). A block in the current picture may be decoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0076] In addition, the merge mode technique may be used for inter-picture prediction to improve decoding efficiency.

[0077] According to some aspects of the present disclosure, predictions such as inter - picture prediction and intra - picture prediction are performed in units of blocks. For example, according to the HEVC (High Efficiency Video Coding) standard, pictures in a video picture sequence are segmented into Coding Tree Units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three Coding Tree Blocks (CTBs), namely one luminance CTB and two chrominance CTBs. Each CTU can be recursively partitioned into one or more Coding Units (CUs) in a quadtree manner. For example, a 64×64 - pixel CTU can be partitioned into one 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an example, each CU is analyzed to determine a prediction type for the CU, such as an inter - prediction type or an intra - prediction type. According to temporal and / or spatial predictability, the CU is partitioned into one or more Prediction Units (PUs). Generally, each PU includes a luminance Prediction Block (PB) and two chrominance PBs. In an aspect, prediction operations in encoding / decoding are performed in units of prediction blocks. Using a luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0078] Note that any suitable technology can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In an aspect, one or more integrated circuits can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210). In another aspect, one or more processors that execute software instructions can be used to implement the video encoders (103) and (303) and the video decoders (110) and (210).

[0079] Various inter-frame prediction modes can be used for video coding. For example, in VVC, for an inter-frame prediction CU, the motion parameters may include an MV, one or more reference picture indices, a reference picture list usage index, and additional information of certain coding features to be used for generating inter-frame prediction samples. The motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU may be associated with a PU and does not have significant residual coefficients, no coded motion vector delta or MV difference (e.g., MVD) or reference picture index. A merge mode can be specified, in which the motion parameters of the current CU are obtained from neighboring CUs, including spatial candidates and / or temporal candidates, and optionally including additional information such as introduced in VVC. The merge mode can be applied to the CUs for inter-frame prediction, not only for the skip mode. In an example, an alternative to the merge mode is the explicit transmission of motion parameters, in which the MV, the corresponding reference picture index for each reference picture list, the reference picture list usage flag, and other information are signaled explicitly for the CU.

[0080] In an embodiment, for example, in VVC, the VVC Test Model (VTM) reference software includes one or more refined inter-frame prediction coding tools, and the one or more refined inter-frame prediction coding tools include: extended merge prediction, merged motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luminance sample MV storage and 8×8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter-frame and intra-frame prediction (CIIP), geometric partitioning mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0081] In some examples, extended merge prediction can be used. In an example, such as in VTM4, a merge candidate list is constructed by sequentially including the following five types of candidates: a spatial motion vector predictor (MVP) from a spatially adjacent CU, a temporal MVP from a collocated CU, a history-based MVP (HMVP) from a first-in first-out (FIFO) table, a pairwise average MVP, and a zero MV.

[0082] The size of the merge candidate list can be signaled in the slice header. In an example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU decoded in the merge mode, a truncated unary binarization (TU) can be used to encode the index of the best merge candidate (e.g., the merge index). The first binary bit (bin) of the merge index can be decoded using context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass decoding can be used for the other binary bits.

[0083] Some examples of the generation process for each category of merge candidates are provided below. In an embodiment, spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as the derivation of spatial merge candidates in HEVC. In an example, up to four merge candidates are selected from the candidates at the positions depicted in Figure 4 .

[0084] Figure 4 shows the positions of spatial merge candidates according to an embodiment of the present disclosure. Referring to Figure 4 , the order of derivation is B1, A1, B0, A0, and B2. Position B2 is considered only if any of the CUs at positions A0, B0, B1, and A1 are unavailable (e.g., because the CU belongs to another slice or another tile) or are intra-coded. After adding the candidate at position A1, the addition of the remaining candidates is subject to a redundancy check that ensures the exclusion of candidates with the same motion information from the candidate list, thereby improving decoding efficiency.

[0085] To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by arrows in Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used for the redundancy check do not have the same motion information.

[0086] Figure 5 shows the candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to Figure 5, the pairs linked by corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0087] In an embodiment, time candidates are derived as follows. In an example, only one temporal merge candidate is added to the candidate list. Figure 6 An example motion vector scaling for temporal merge candidates is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) can be derived based on a co-located CU (612) belonging to a collocated reference picture (604) (e.g., as indicated by the dashed line in Figure 6 . The reference picture list for deriving the co-located CU (612) can be signaled explicitly in the slice header. As indicated by the dashed line in Figure 6 , the scaled MV (621) of the temporal merge candidate can be obtained. The scaled MV (621) can be scaled from the MV of the co-located CU (612) using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the collocated reference picture (604) of the collocated picture (603) and the collocated picture (603). The reference picture index of the temporal merge candidate can be set to zero. The collocated picture is a reference picture used as a source picture for temporal motion information derivation. The collocated picture can be identified in one of two lists called list 0 or list 1. In some examples, the encoder can determine the collocated picture and signal it using appropriate syntax techniques.

[0088] Figure 7 An example candidate positions (e.g., C0 and C1) for the temporal merge candidate of a current CU are shown. The position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is at the lower right corner of the co-located CU (710) of the current CU. Candidate position C1 is at the center of the co-located CU (710) of the current CU. If the CU at candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, then candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, is inter-coded, and is in the current row of the CTU, then candidate position C0 is used to derive the temporal merge candidate.

[0089] According to some aspects of the present disclosure, a function-based prediction method can be used for inter-frame prediction or intra-frame prediction. For inter-frame prediction, the function-based prediction method can use a function to generate samples of a current block in a current picture based on reference samples of a reference block in a reference picture. For intra-frame prediction, the function-based prediction method can use a function to generate a first color component of a current block based on a second color component of the current block. In some examples, the parameters in the function of the function-based prediction method can be derived based on a template of the current block.

[0090] In some examples, local illumination compensation (LIC) is used as an inter-frame prediction technique to model local illumination changes between a current block and a predicted block (also referred to as a reference block) of the current block by using a linear function. The predicted block is in a reference picture and can be pointed to by a motion vector (MV). The parameters of the linear function can include a scale α and an offset β, and the linear function can be represented by α×p[x,y]+β to compensate for illumination changes, where p[x,y] represents a reference sample at the location [x,y] in the reference block (also referred to as the predicted block), and the reference block is pointed to by the MV. In some examples, the scale α and the offset β can be derived by using the least squares method based on a template of the current block and a corresponding reference template of the reference block, so that no signaling overhead is required except that the LIC flag can be signaled to indicate the use of LIC. The scale α and the offset β derived based on the template of the current block can be referred to as a template-based parameter set.

[0091] In some examples, LIC is used for unidirectional prediction of an inter-frame CU. In some examples, intra-frame adjacent samples of the current block (adjacent samples predicted using intra-frame prediction) can be used for deriving LIC parameters. In some examples, LIC is disabled for blocks with fewer than 32 luma samples. In some examples, for non-subblock modes (e.g., non-affine modes), LIC parameter derivation is performed based on the modular block samples of the current CU rather than the partial modular block samples of the first top-left 16×16 unit. In some examples, LIC parameter derivation is performed based on partial modular block samples, such as the partial modular block samples of the first top-left 16×16 unit. In some examples, the template samples of the reference block are determined by motion compensation (MC) of the MV of the block without rounding it to integer pixel precision.

[0092] In some examples, cross-component prediction can be used as an intra-frame prediction technique. Cross-component prediction can include a first technique called cross-component linear model (CCLM), a second technique called multi-model linear model (MMLM), a third technique called convolutional cross-component model (CCCM), and a fourth technique called gradient linear model (GLM).

[0093] For example, the first technique CCLM is used to reduce cross-component redundancy. In CCLM, by using a linear model, such as using Equation (1), chrominance samples are predicted based on the reconstructed luma samples of the same CU: pred C (i, j) = a · rec L ′(i, j) + b Equation (1) where pred C (i, j) represents the predicted chrominance sample in the CU, and rec L ′(i, j) represents the downsampled reconstructed luma sample of the same CU. The CCLM linear model includes parameters (a and b) that can be derived using at most four adjacent chrominance samples and their corresponding downsampled luma samples in the example. In the example, at most four adjacent chrominance samples and their corresponding downsampled luma samples are referred to as the template of the CU.

[0094] In some examples, based on the positioning of adjacent chrominance samples, CCLM can include different modes called LM_T (LM top mode or upper mode LM_A), LM_L (LM left mode), and LM_LT (LM left-top mode or upper-left mode LM_LA or just LM mode). For example, if the size of the current chrominance block is W × H, then W’ and H’ can be set for various modes in CCLM. When the LM mode (also known as LM_LT or LM_LA) is applied, W’ = W and H’ = H; when the LM-A mode is applied, W’ = W + H; when the LM-L mode is applied, H’ = H + W.

[0095] Note that MMLM, CCCM, and GLM also use functions for prediction. The parameters of the functions can be derived based on the template.

[0096] Note that the following description uses inter prediction to illustrate the technique of obtaining decoding information for function-based prediction methods, and the technique can be appropriately used for obtaining decoding information for intra prediction.

[0097] Some aspects of the present disclosure provide techniques for improving the derivation of decoding information for function-based prediction methods. For example, an encoder / decoder may construct a candidate list including one or more decoded blocks associated with a current block. The decoded blocks associated with the current block in the candidate list are candidates for providing decoding information for a function-based prediction method for predicting the current block, where the function-based prediction method uses a function with parameters derived based on a template of the current block. The encoder / decoder may select a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the function-based prediction method. In addition, the encoder / decoder may inherit the first decoding information of the function-based prediction method from the first encoded block to the current block and derive parameters of the function used in the function-based prediction method according to the first decoding information.

[0098] According to some aspects of the present disclosure, some inter prediction techniques are designed to minimize the distortion between a current block and its predicted block in a corresponding reference picture. For example, an inter prediction technique (also referred to as a first method of an inter prediction method) may apply a function such as a non-linear function, a linear function, etc. using an original predicted block in a reference picture as an input to generate a current block in a current picture. For example, an inter prediction technique may generate a prediction of samples in a current block based on a function of one or more predicted samples in a reference picture. The function may include a linear term or a non-linear term and may include one or more parameters that can be derived. Note that LIC is one of such inter prediction techniques.

[0099] In some examples, the function is a linear function and can be represented by where n is a non-negative integer, and p(x i , y i ) is a predicted sample at a location (x i , y i ) in a reference image, and the predicted sample is pointed to based on an MV associated with the current block. In addition, a set of predicted samples represented by p(x i , y i ) - where i = 0,..., n - may be a set of predicted samples around corresponding samples in the reference samples pointed to by the MV based on the current sample to be predicted. In some examples, parameters α i and β may be derived (e.g., by using the least squares method) by minimizing the difference between the current block template and its predicted block template based on the template of the current block (also referred to as the current block template) and the template of the predicted block of the current block (also referred to as the predicted block template). The template of the current block or the predicted block is respectively composed of spatially adjacent reconstructed samples of the current block or the predicted block.

[0100] Figure 8A diagram showing templates in some examples. For example, template (810) is referred to as an L-shaped template T L , and includes neighboring samples in the row above, the column to the left, and at the upper left corner of the current block (also referred to as the current decoding block); template (820) is referred to as an above and left template T a+1 , and includes neighboring samples in the row above and the column to the left of the current block; template (830) is referred to as an above template T a , and includes neighboring samples in the row above the current block; and template (840) is referred to as a left template T l , and includes neighboring samples in the column to the left of the current block. Note that the template may include Figure 8 neighboring samples of other suitable shapes not shown in

[0101] In some examples, multiple candidate template types may also be supported, and one candidate template type is selected to derive the parameters of the linear function. The syntax may be signaled in the bitstream (e.g., at the block level) to indicate which candidate template type is selected.

[0102] In some examples, a control flag may be signaled in the bitstream associated with the inter-frame prediction technique (e.g., at the block level) to indicate whether the inter-frame prediction technique is applied to the current block. Alternatively, the value of this control flag may also be inherited from other decoding blocks. More specifically, the first control flag of the inter-frame prediction technique associated with the current block is inherited from the second control flag of the inter-frame prediction technique associated with another decoding block or multiple decoding blocks. In addition, the control flag may be derived at the decoding block level to adaptively determine whether to apply inter-frame prediction.

[0103] In some examples, the first control flag of the inter-frame prediction technique associated with the current block is inherited from the second control flag of the inter-frame prediction technique associated with another decoding block or multiple decoding blocks. In some examples, the decoding information of the inter-frame prediction technique may be derived from neighboring decoding blocks, non-neighboring decoding blocks, or decoding blocks that store the decoding information in the buffer.

[0104] Note that in this disclosure, without limiting generality, in the examples, the term parameter refers to the parameters α i and β used to determine the linear function for deriving the prediction block. The term template type refers to different template shapes, such as but not limited to Figure 8 one of the template types shown for parameter derivation of non-linear or linear prediction functions.

[0105] Some aspects of the present disclosure provide techniques for constructing a candidate list by determining whether decoding information from an adjacent decoding block, a non - adjacent decoding block, or a decoding block storing decoding information in a buffer can be used as a candidate in the candidate list. And then, the decoding information of the selected candidate is inherited to the current block. In some examples, the decoding information refers to the information used by an inter - frame prediction technique to generate a prediction of a block. In some examples, the decoding information is associated with a specific inter - frame prediction technique. For example, when the inter - frame prediction technique is LIC, the decoding information may include a control flag, a template type, a model type, etc. The control flag is associated with the inter - frame prediction technique. The template type indicates adjacent samples used in the inter - frame prediction technique and can be any suitable template type, such as Figure 8 the various template types shown. The model type may indicate the type of function used by the inter - frame prediction technique, such as a linear function, a non - linear function, etc.

[0106] In some examples, the inter - frame prediction mode information of the decoding block and the current block is used to determine whether the decoding information (for the inter - frame prediction technique) of the decoding block can be used to derive the decoding information (for the inter - frame prediction technique) of the current block.

[0107] In some examples, the inter - frame prediction mode information includes but is not limited to a prediction direction, a reference list (also referred to as a reference picture list), a reference index (also referred to as a reference picture index), etc. In an example, the decoding information of the inter - frame prediction technique from the decoding block can be inherited to the current block when the prediction mode, the reference list, and the reference index of the decoding block are the same as those of the current block.

[0108] In some examples, the decoding information of the inter - frame prediction technique of the decoding block can be inherited to the current block when the prediction mode of the decoding block is the same as that of the current block.

[0109] In some examples, the decoding information of the inter - frame prediction technique of the decoding block can be inherited to the current block when the prediction mode and the reference list in the decoding block are the same as those of the current block.

[0110] In some examples, the decoding information of the inter prediction technique of the decoding block can be inherited to the current block when the current block is a uni - directional prediction block, and the reference list of the current block is also available in the decoding block. In the example, the current block is a uni - directional prediction block, and only reference list X (e.g., X is 0 or 1) is available. Then, the decoding information of the inter prediction technique of the decoding block at only reference list X is inherited to the current block. In the example, the decoding block is a uni - directional prediction block using reference list X, so the decoding information of the inter prediction technique of the decoding block at reference list X is inherited to the current block. In another example, the decoding block is a bi - directional prediction block using reference list X and reference list (1 - X), so the decoding information of the inter prediction technique of the decoding block at reference list X is inherited to the current block, while the decoding information of the inter prediction technique of the decoding block at reference list (1 - X) is not inherited to the current block.

[0111] In some examples, the decoding information of the inter prediction technique of the decoding block can be inherited to the current block when the decoding block is a bi - directional prediction block and the current block is a uni - directional prediction block, and the reference list and reference index of the current block are also used in the decoding block. In the example, the current block is a uni - directional prediction block, and only reference list X at reference index Y is available. The decoding information of the inter prediction technique of the decoding block at only reference list X and reference index Y is inherited to the current block.

[0112] In some examples, the decoding information of the inter prediction technique (e.g., the first method) of the selected decoding block is inherited to the current block regardless of the prediction mode, reference list, and reference index of the selected decoding block.

[0113] In some examples, the current block is a uni - directional prediction block, and reference list X is used in the current block. The decoding information of the inter prediction technique (e.g., the first method) in reference list X of the decoding block is inherited to the current block when reference list X is available for the decoding block. Otherwise, the decoding information of the inter prediction technique (e.g., the first method) in another reference list (reference list 1 - X) is inherited.

[0114] In some examples, for the case of bi - directionally predicting the current block, the default decoding information of the inter prediction technique (e.g., the first method) in reference list X is inherited to the current block when reference list X is not available for the selected decoding block. In the example, the default decoding information can be a predefined value, or can be signaled in syntax elements such as sequence - level syntax elements, picture - level syntax elements, slice - level syntax elements, CTU - level syntax elements, CU - level syntax elements, etc.

[0115] Figure 9A flowchart outlining a process (900) according to aspects of the present disclosure is shown. The process (900) can be used in a video decoder. In various aspects, the process (900) is performed by processing circuitry, such as processing circuitry that performs the functions of video decoder (110), processing circuitry that performs the functions of video decoder (210), etc. In some aspects, the process (900) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry performs the process (900). The process begins at (S901) and proceeds to (S910).

[0116] At (S910), a candidate list including one or more decoded blocks associated with the current block is constructed. The decoded blocks associated with the current block are candidates that provide decoded information for a function-based prediction method for predicting the current block. The function-based prediction method uses a function with parameters derived from a template based on the current block.

[0117] At (S920), a first decoded block is selected from the candidate list. The first decoded block is decoded with first decoded information of the function-based prediction method.

[0118] At (S930), the first decoded information of the function-based prediction method is inherited from the first decoded block to the current block.

[0119] At (S940), the parameters of the function in the function-based prediction method are derived according to the first decoded information.

[0120] At (S950), the samples of the current block are reconstructed according to the parameters of the function in the function-based prediction method.

[0121] In some examples, one or more decoded blocks include at least one of an adjacent decoded block, a non-adjacent decoded block, and a decoded block whose decoded information is in a buffer.

[0122] In some examples, the function-based prediction method is one of local illumination compensation (LIC), cross-component linear model (CCLM), multi-model linear model (MMLM), convolutional cross-component model (CCCM), gradient linear model (GLM).

[0123] In some examples, the function-based prediction method is an inter-frame prediction method, such as LIC. The first decoded block can be selected from the candidate list based on the inter-frame prediction mode information of the first decoded block and the current block. The inter-frame prediction mode information includes at least one of a prediction mode, a prediction direction, a reference picture list, and a reference index.

[0124] In some examples, when the current block and the first decoded block share the same prediction mode, the same prediction direction, the same reference list, and the same reference index, the first decoded block is selected from the candidate list.

[0125] In some examples, when the current block and the first decoded block share the same prediction mode, the first decoded block is selected from the candidate list.

[0126] In some examples, when the current block and the first decoded block share the same prediction mode and the same reference picture list, the first decoded block is selected from the candidate list.

[0127] In some examples, when the current block is a uni - directional prediction block using the first reference list and the first reference list is available at the first decoded block, the first decoded block is selected from the candidate list. In the example, only the first decoding information of the inter - prediction method of the first reference list is inherited from the first decoded block to the current block.

[0128] In some examples, when the current block is a uni - directional prediction block using the first reference list and the first reference index, and the first decoded block is a bi - directional prediction block using the first reference list and the first reference index used in decoding, the first decoded block is selected from the candidate list. In the example, only the first decoding information of the inter - prediction method at the first reference list and the first reference index is inherited from the first decoded block to the current block.

[0129] In some examples, the first decoding information of the inter - prediction method is inherited to the current block regardless of the prediction mode, reference list, and reference index of the decoded block.

[0130] In some examples, the current block is a uni - directional prediction block and the current block uses the first reference list. When the first reference list is available at the first decoded block, the first decoding information of the inter - prediction method at the first reference list is inherited from the first decoded block to the current block; and when the first reference list is not available at the first decoded block, the first decoding information of the inter - prediction method at the second reference list is inherited from the first decoded block to the current block.

[0131] In some examples, the current block is a bi - directional prediction block. When the reference list is available at the first decoded block, the first decoding information of the inter - prediction method at the reference list is inherited from the first decoded block to the current block; and when the reference list is not available at the first decoded block, the default decoding information of the inter - prediction method is inherited to the current block. In the example, the default decoding information includes predefined values. In another example, the default decoding information is indicated by a syntax element signaled at one of the sequence level, picture level, slice level, coding tree unit (CTU) level, and coding unit (CU) level.

[0132] In some examples, the function-based prediction method is local illumination compensation (LIC), and the first decoding information includes at least one of a control flag of the LIC, a template type of the LIC, and a model type of the LIC.

[0133] Then, the process proceeds to (S999) and terminates.

[0134] The process (900) can be appropriately adjusted. Steps in the process (900) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0135] Figure 10 A flowchart outlining a process (1000) according to aspects of the present disclosure is shown. The process (1000) can be used in a video encoder. In various aspects, the process (1000) is performed by processing circuitry, such as processing circuitry that performs the functions of video encoder (103), processing circuitry that performs the functions of video encoder (303), etc. In some aspects, the process (1000) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry performs the process (1000). The process starts at (S1001) and proceeds to (S1010) in parallel.

[0136] At (S1010), it is determined to use a function-based prediction method for prediction of a current block in a current picture. The function-based prediction method uses a function having parameters derived from a template of the current block.

[0137] At (S1020), a candidate list including one or more decoded blocks associated with the current block is constructed, and the decoded blocks associated with the current block are candidates for providing decoding information for the function-based prediction method for prediction of the current block.

[0138] At (S1030), a first decoded block is selected from the candidate list, and the first decoded block is decoded using first decoding information of the function-based method.

[0139] At (S1040), the first decoding information of the function-based prediction method is inherited from the first decoded block to the current block.

[0140] At (S1050), parameters of the function used in the function-based prediction method are derived according to the first decoding information.

[0141] At (S1060), samples of the current block are reconstructed according to the derived parameters. Then the current block is encoded accordingly.

[0142] In some examples, one or more decoding blocks include at least one of adjacent decoding blocks, non - adjacent decoding blocks, and decoding blocks in which decoding information is in a buffer.

[0143] In some examples, the function - based prediction method is one of local illumination compensation (LIC), cross - component linear model (CCLM), multi - model linear model (MMLM), convolutional cross - component model (CCCM), and gradient linear model (GLM).

[0144] In some examples, the function - based prediction method is an inter - frame prediction method, such as LIC. The first decoding block can be selected from a candidate list based on the inter - frame prediction mode information of the first decoding block and the current block. The inter - frame prediction mode information includes at least one of a prediction mode, a prediction direction, a reference picture list, and a reference index.

[0145] In some examples, when the current block and the first decoding block share the same prediction mode, the same prediction direction, the same reference list, and the same reference index, the first decoding block is selected from the candidate list.

[0146] In some examples, when the current block and the first decoding block share the same prediction mode, the first decoding block is selected from the candidate list.

[0147] In some examples, when the current block and the first decoding block share the same prediction mode and the same reference picture list, the first decoding block is selected from the candidate list.

[0148] In some examples, when the current block is a uni - directional prediction block using the first reference list and the first reference list is available at the first decoding block, the first decoding block is selected from the candidate list. In the example, only the first decoding information of the inter - frame prediction method of the first reference list is inherited from the first decoding block to the current block.

[0149] In some examples, when the current block is a uni - directional prediction block using the first reference list and the first reference index and the first decoding block is a bi - directional prediction block using the first reference list and the first reference index used in decoding, the first decoding block is selected from the candidate list. In the example, only the first decoding information of the inter - frame prediction method at the first reference list and the first reference index is inherited from the first decoding block to the current block.

[0150] In some examples, the first decoding information of the inter - frame prediction method is inherited to the current block regardless of the prediction mode, reference list, and reference index of the decoding block.

[0151] In some examples, the current block is a uni - directional prediction block and the current block uses a first reference list. When the first reference list is available at the first decoded block, the first decoded information of the inter - prediction method at the first reference list is inherited from the first decoded block to the current block; and when the first reference list is not available at the first decoded block, the first decoded information of the inter - prediction method at the second reference list is inherited from the first decoded block to the current block.

[0152] In some examples, the current block is a bi - directional prediction block. When the reference list is available at the first decoded block, the first decoded information of the inter - prediction method at the reference list is inherited from the first decoded block to the current block; and when the reference list is not available at the first decoded block, the default decoded information of the inter - prediction method is inherited to the current block. In an example, the default decoded information includes predefined values. In another example, the encoder may include a syntax element in the bitstream to indicate the default decoded information. The syntax element may be signaled at one of sequence level, picture level, slice level, coding tree unit (CTU) level, and coding unit (CU) level.

[0153] In some examples, the function - based prediction method is local illumination compensation (LIC), and the first decoded information includes at least one of a control flag of LIC, a template type of LIC, and a model type of LIC.

[0154] Then, the process proceeds to (S1099) and terminates.

[0155] The process (1000) can be modified appropriately. Steps in the process (1000) can be modified and / or omitted. Additional steps can be added. Any suitable order of implementation can be used.

[0156] According to an aspect of the present disclosure, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to formatting rules. For example, the bitstream can be a bitstream decoded / encoded by any of the decoding and / or encoding methods described herein. The formatting rules can specify one or more constraints of the bitstream and / or one or more processes to be performed by the decoder and / or encoder.

[0157] In an example, a bitstream includes decoding information of one or more pictures, the one or more pictures including a current block in a current picture. Format rules specify constructing a candidate list including one or more decoded blocks associated with the current block. The decoded blocks associated with the current block are candidates for decoding information that provides local illumination compensation (LIC) for prediction of the current block, and the decoding information of the LIC includes at least one of a control flag of the LIC, a template type of the LIC, and a model type of the LIC. The format rules also specify selecting a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the LIC. For example, the first decoding information of the LIC includes at least one of a first control flag of the LIC, a first template type of the LIC, and a first model type of the LIC. The format rules also specify that the first decoding information of the LIC is inherited from the first decoded block to the current block, LIC parameters are derived based on the first decoding information, and samples of the current block are reconstructed based on the derived LIC parameters.

[0158] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 11 FIG. shows a computer system (1100) suitable for implementing certain aspects of the disclosed subject matter.

[0159] The computer software may be decoded using any suitable machine code or computer language, and the machine code or computer language may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that may be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc. or executed through interpretation, microcode execution, etc.

[0160] The instructions may be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0161] Figure 11 The components shown in FIG. for the computer system (1100) are examples and are not intended to impose any limitation on the scope of use or functionality of the computer software for implementing aspects of the present disclosure. The configuration of the components should also not be construed as having any dependence or requirement related to any one component or combination of components shown in the example aspects of the computer system (1100).

[0162] A computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs made by one or more human users through, for example, tactile inputs (e.g., keystrokes, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious inputs made by humans, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from still image cameras), and video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0163] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touchscreen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera device (1108).

[0164] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile outputs, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touchscreen (1110), data glove (not shown), or joystick (1105), but there may also be tactile feedback devices that do not serve as input devices); audio output devices (e.g., speakers (1109), headphones (not depicted)); visual output devices (e.g., screens (1110), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual outputs or more than three-dimensional outputs through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and fog machines (not depicted)); and printers (not depicted).

[0165] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with media such as CD / DVD (1121), thumb drives (1122), removable hard disk drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.

[0166] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass a transmission medium, a carrier wave, or other transient signals.

[0167] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; vehicular and industrial networks including CAN bus (CANBus), etc. Some networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (1149) (such as, for example, the USB port of the computer system (1100)); other networks are typically integrated into the core of the computer system (1100) by attaching to the system bus (such as, for example, an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system) as described below. Using any of these networks, the computer system (1100) can communicate with other entities. Such communication can be one-way, receive-only (e.g., broadcast TV), one-way transmit-only (e.g., CAN bus to certain CAN bus devices), or two-way, e.g., to other computer systems using local or wide area digital networks. As described above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.

[0168] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1140) of the computer system (1100).

[0169] The core (1140) may include one or more central processing units (CPUs) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) (1143), a hardware accelerator for specific tasks (1144), a graphics adapter (1150), etc. These devices, together with a read-only memory (ROM) (1145), a random access memory (1146), and an internal mass storage device (1147) such as an internal hard disk drive or SSD that is not user-accessible, can be connected via a system bus (1148). In some computer systems, the system bus (1148) can be accessed in the form of one or more physical plugs to enable expansion through additional CPUs, GPUs, etc. Peripheral devices can be directly attached to the system bus (1148) of the core or attached to the system bus (1148) of the core through a peripheral bus (1149). In an example, a screen (1110) can be connected to the graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, etc.

[0170] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) can execute certain instructions, which can combinatorially form the computer code mentioned above. The computer code can be stored in the ROM (1145) or RAM (1146). Transient data can also be stored in the RAM (1146), while permanent data can be stored in, for example, the internal mass storage device (1147). Fast storage and retrieval to any of the memory devices in the memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (1141), GPU (1142), mass storage device (1147), ROM (1145), RAM (1146), etc.

[0171] A computer-readable medium can have computer code for performing various computer-implemented operations. The medium and the computer code can be media and computer code that are specially designed and constructed for the purposes of this disclosure, or they can be of the types that are known and available to those skilled in the art of computer software.

[0172] By way of example and not limitation, a computer system (1100) having an architecture and in particular a core (1140) can provide functionality due to a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage device as introduced above, and certain storage devices of the core (1140) having a non-transitory nature, such as a mass storage device (1147) inside the core or a ROM (1145). The software implementing various aspects of the present disclosure can be stored in such a device and executed by the core (1140). Depending on specific needs, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1140) - and in particular the processors therein (including the CPU, GPU, FPGA, etc.) - to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in the RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality due to logic hardwired or otherwise embodied in a circuit (e.g., an accelerator (1144)), which can operate in place of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, a reference to software can include logic and vice versa. In appropriate cases, a reference to computer-readable media can include a circuit (e.g., an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0173] The use of "at least one of... " or "one of... " in the present disclosure is intended to include any one or combination of the recited elements. For example, a reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. A reference to one of A or B and one of A and B is intended to include A or B or (A and B). The use of "one of... " does not exclude any combination of the recited elements in cases where, for example, the elements are not mutually exclusive.

[0174] Although the present disclosure has described several examples of aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Accordingly, it will be recognized that those skilled in the art will be able to design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for processing visual media data, the method comprising: Processing a bitstream of visual media data according to format rules, wherein: The bitstream includes decoding information of one or more pictures, and the one or more pictures include a current block in a current picture; and The format rules specify: Constructing a candidate list including one or more decoded blocks associated with the current block, and the decoded blocks associated with the current block are candidates for decoding information that provides local illumination compensation (LIC) for prediction of the current block, and the decoding information of the LIC includes at least one of the following: a control flag of the LIC, a template type of the LIC, and a model type of the LIC; Selecting a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the LIC, and the first decoding information of the LIC includes at least one of the following: a first control flag of the LIC, a first template type of the LIC, and a first model type of the LIC; The first decoding information of the LIC is inherited from the first decoded block to the current block; Deriving parameters of the LIC according to the first decoding information; and Reconstructing samples of at least the current block according to the parameters of the LIC.

2. A device for video decoding, comprising processing circuitry configured to: Construct a candidate list, the candidate list including one or more decoded blocks associated with a current block, and the decoded blocks associated with the current block are candidates for decoding information that provides a function-based prediction method for prediction of the current block, and the function-based prediction method uses a function with parameters derived from a template based on the current block; Selecting a first decoded block from the candidate list, and the first decoded block is decoded with first decoding information of the function-based prediction method; Inheriting the first decoding information of the function-based prediction method from the first decoded block to the current block; Deriving parameters of the function used in the function-based prediction method according to the first decoding information; And Reconstructing samples of at least the current block according to the parameters of the function.

3. The device according to claim 2, wherein The one or more decoded blocks include at least one of the following: an adjacent decoded block, a non-adjacent decoded block, and a decoded block whose decoding information is in a buffer.

4. The device according to claim 2, wherein The function-based prediction method is one of the following: local illumination compensation (LIC), cross-component linear model (CCLM), multi-model linear model (MMLM), convolutional cross-component model (CCCM), and gradient linear model (GLM).

5. The device according to claim 2, wherein, The function-based prediction method is an inter-frame prediction method, and the processing circuitry is configured to: Select the first decoded block from the candidate list based on inter-frame prediction mode information of the first decoded block and the current block.

6. The apparatus according to claim 5, wherein, The inter-frame prediction mode information includes at least one of the following: a prediction mode, a prediction direction, a reference list, and a reference index.

7. The device according to claim 6, wherein The processing circuitry is configured to: When the current block and the first decoded block share the same prediction mode, the same prediction direction, the same reference list, and the same reference index, select the first decoded block from the candidate list.

8. The device according to claim 6, wherein, The processing circuitry is configured to: When the current block and the first decoded block share the same prediction mode, select the first decoded block from the candidate list.

9. The device according to claim 6, wherein The processing circuitry is configured to: When the current block and the first decoded block share the same prediction mode and one or more identical reference lists, select the first decoded block from the candidate list.

10. The device according to claim 6, wherein The processing circuitry is configured to: When the current block is a uni-directional prediction block using a first reference list and the first reference list is available at the first decoded block, select the first decoded block from the candidate list.

11. The apparatus according to claim 10, wherein, The processing circuitry is configured to: Inherit only the first decoded information of the inter-frame prediction method of the first reference list from the first decoded block to the current block.

12. The device according to claim 6, wherein, The processing circuitry is configured to: When the current block is a uni-directional prediction block using a first reference list and a first reference index, and the first decoded block is a bi-directional prediction block using the used first reference list and the first reference index, select the first decoded block from the candidate list.

13. The apparatus according to claim 12, wherein, The processing circuitry is configured to: Inherit only the first decoded information of the inter-frame prediction method at the first reference list and the first reference index from the first decoded block to the current block.

14. The device according to claim 6, wherein, The current block is a uni-directional prediction block and the current block uses a first reference list. The processing circuitry is configured to: When the first reference list is available at the first decoded block, inherit the first decoded information of the inter-frame prediction method at the first reference list from the first decoded block to the current block; And When the first reference list is not available at the first decoded block, inherit the first decoded information of the inter-frame prediction method at a second reference list from the first decoded block to the current block.

15. A method for video coding, comprising: Determine to use an inter-frame prediction method for prediction of a current block in a current picture, the inter-frame prediction method being a function-based prediction method having parameters derived based on a template of the current block; Construct a candidate list including one or more decoded blocks associated with the current block, the decoded blocks associated with the current block being candidates that provide decoded information of the inter-frame prediction method for prediction of the current block; Select a specific decoded block from the candidate list, the specific decoded block being decoded with specific decoded information of the inter-frame prediction method; Inherit the specific decoded information of the inter-frame prediction method from the specific decoded block to the current block; Derive parameters of a function used in the inter-frame prediction method according to the specific decoded information; And Reconstruct samples of the current block according to the parameters of the function.