List construction of coding information for inter prediction

By adopting formula-based inter prediction technology in video encoding technology, using the information of reference blocks to generate more accurate prediction samples, the problem of low encoding efficiency in the prior art is solved, and more efficient video encoding is achieved.

CN120019647APending Publication Date: 2025-05-16TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004320.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-11
Filing Date
2024-07-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video encoding technology is difficult to effectively utilize the information of the reference block in inter-frame prediction, resulting in low encoding efficiency.

Method used

Using formula-based inter prediction technology, a predicted sample of the current block is generated by inputting the reconstruction sample of the reference block in the reference picture into a specific formula. The formula includes parameters derived based on the template of the current block and the reference block. And build a candidate list and select the appropriate encoded block to determine the prediction information of the current block.

Benefits of technology

The encoding efficiency of inter-frame prediction is improved, and through more accurate prediction sample generation, redundant information in the code stream is reduced, and the quality of video decoding is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019647A_ABST
    Figure CN120019647A_ABST
Patent Text Reader

Abstract

A method includes receiving a stream of encoded information for a sequence of pictures, the encoded information indicating inter prediction of a current block in a current picture using a formula-based inter prediction technique; the formula-based inter prediction technique generates a prediction sample of a current block by inputting a reconstructed sample of a reference block in a reference picture to a formula, based on the formula including parameters derived based on a current template of the current block and a reference template of the reference block. The method further includes constructing a candidate list, the candidate list including one or more coded blocks associated with the current block, a coded block of the one or more coded blocks being a candidate providing formula-based inter prediction information.
Need to check novelty before this filing date? Find Prior Art

Description

Incorporation by reference

[0001] This application claims the benefit of priority to U.S. Patent Application No. 18 / 770,485, filed on July 11, 2024, entitled “ON LIST CONSTRUCTION OF CODEDINFORMATION OF INTER PREDICTION,” which claims the benefit of priority to U.S. Provisional Application No. 63 / 526,437, filed on July 12, 2023, entitled “ON LIST CONSTRUCTION OF CODEDINFORMATION OF INTER PREDICTION.” The disclosure of the prior application is incorporated herein by reference in its entirety. Technical Field

[0002] This disclosure describes aspects generally related to video encoding. Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. To the extent that the work of the presently named inventors described in this background section and in various aspects of this specification was performed, it does not indicate that it qualifies as prior art at the time of filing, and it is never explicitly or implicitly admitted that it is prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video data between different devices, storage, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial redundancy and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from a current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can predict samples in a current picture based on a previously reconstructed picture using motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the invention

[0005] Aspects of the present disclosure include code streams, methods and apparatus for video encoding / decoding. In some examples, the apparatus for video encoding / decoding includes a processing circuit.

[0006] Some aspects of the present disclosure provide a method for video decoding. The method includes receiving a code stream of encoding information for a picture sequence, the encoding information indicating that a current block in a current picture is inter-predicted using a formula-based inter-frame prediction technique; the formula-based inter-frame prediction technique is based on a formula, and generates a prediction sample of the current block by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block. The method also includes constructing a candidate list, the candidate list including one or more encoding blocks associated with the current block, the encoding blocks in the one or more encoding blocks being candidates for providing formula-based inter-frame prediction information of the formula-based inter-frame prediction technique. In addition, the method includes selecting a first encoding block from the candidate list, the first encoding block being encoded using a first formula-based inter-frame prediction information of the formula-based inter-frame prediction technique. The method also includes: determining current formula-based inter-frame prediction information of the current block based on the first formula-based inter-frame prediction information; deriving one or more parameter values ​​for applying the formula-based inter-frame prediction technology to one or more parameters of the current block according to the current formula-based inter-frame prediction information; and generating at least one prediction sample of the current block based on a formula including one or more parameters set according to the one or more parameter values.

[0007] In some examples, the one or more coding blocks include at least one of an adjacent spatially adjacent coding block, a non-adjacent spatially adjacent coding block, and / or a temporally co-located coding block of the current block in a reference picture.

[0008] In some embodiments, the method includes constructing the candidate list according to a first-in-first-out (FIFO) queue, the FIFO queue storing historical formula-based inter-frame prediction information of one or more historically encoded blocks encoded using a formula-based inter-frame prediction technique before the current block.

[0009] In some examples, the method includes storing current formula-based inter prediction information applied to the current block in a FIFO queue in a FIFO manner.

[0010] In some examples, the method includes at least one of: resetting the FIFO queue at the beginning of a frame; resetting the FIFO queue at the beginning of a row of CTUs; resetting the FIFO queue at the beginning of a row of super blocks (SBs); resetting the FIFO queue at the beginning of a slice; resetting the FIFO queue at the beginning of a tile; and / or resetting the FIFO queue at the beginning of a fragment.

[0011] In some examples, the method includes decoding a first flag from a bitstream. When the first flag is true, the method includes decoding an index from the bitstream, the index indicating a first coding block in a candidate list. When the first flag is false, the method includes determining current formula-based inter-frame prediction information without using the candidate list. In one example, when the first flag is false, the method includes decoding current formula-based inter-frame prediction information of a formula-based inter-frame prediction technique from the bitstream. In another example, when the first flag is false, the method includes deriving current formula-based inter-frame prediction information of a formula-based inter-frame prediction technique in the absence of additional signaling in the bitstream.

[0012] In some examples, the method includes constructing a candidate list based on a scanning order of at least one adjacent spatially adjacent coding block, a first-in-first-out (FIFO) queue storing previously coded blocks via a formula-based inter-frame prediction technique, at least one non-adjacent spatially adjacent block, and / or at least one of a temporally co-located coding block of a current block in a reference picture.

[0013] In one example, the method includes: checking whether an adjacent spatially adjacent coding block is encoded using a formula-based inter-frame prediction technique; when an adjacent spatially adjacent coding block is encoded using a formula-based inter-frame prediction technique, inserting the adjacent spatially adjacent coding block into a candidate list; checking whether a FIFO queue is empty; when the FIFO queue is not empty, inserting the FIFO queue into a candidate list; checking whether a non-adjacent spatially adjacent block is encoded using a formula-based inter-frame prediction technique; when a non-adjacent spatially adjacent block is encoded using a formula-based inter-frame prediction technique, inserting the non-adjacent spatially adjacent block into a candidate list; checking whether a temporally co-located coding block of a current block in a reference picture is encoded using a formula-based inter-frame prediction technique; and when a temporally co-located coding block is encoded using a formula-based inter-frame prediction technique, inserting the temporally co-located coding block into the candidate list.

[0014] In some examples, the method includes: calculating template matching costs for one or more coding blocks respectively; and reordering one or more coding blocks in a candidate list according to the template matching costs.

[0015] In some examples, the candidate list includes a first candidate at the beginning of the candidate list and one or more coding blocks after the first candidate, the first candidate indicating current formula-based inter-frame prediction information using derivation and / or direct signaling. The method includes: decoding an index from a bitstream; and when the index indicates the first candidate, determining the current formula-based inter-frame prediction information without using the formula-based inter-frame prediction information of the one or more coding blocks.

[0016] In one example, when the index indicates the first candidate, the method includes decoding current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique from the code stream. In another example, the method includes deriving current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique in the absence of additional signaling from the code stream.

[0017] In some examples, the method includes: calculating template matching costs for one or more coding blocks respectively; and reordering one or more coding blocks located after the first candidate in the candidate list according to the template matching costs.

[0018] In some examples, the formula-based inter prediction information includes at least one of a template type and a formula type.

[0019] Some aspects of the present disclosure provide a method for video encoding. The method includes determining to encode a current block in a current picture by inter-frame prediction, and the inter-frame prediction has potential use of a formula-based inter-frame prediction technology; the formula-based inter-frame prediction technology is based on a formula, and generates a prediction sample of the current block by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, and the formula includes one or more parameters derived based on a current template of the current block and a reference template of the reference block. The method also includes constructing a candidate list, which includes one or more coding blocks associated with the current block. The coding blocks in the one or more coding blocks are candidates for providing formula-based inter-frame prediction information of the formula-based inter-frame prediction technology for inter-frame prediction of the current block. The method also includes: determining the current formula-based inter-frame prediction information of the current block based on the candidate list; and encoding the current block based on the current formula-based inter-frame prediction information. In one example, the method includes generating coding information of the current block to be included in the code stream.

[0020] Some aspects of the present disclosure provide a method for processing visual media data. The method includes processing a code stream of visual media data according to a format rule. The code stream includes encoding information of a picture sequence, the encoding information indicating that a current block in a current picture is inter-predicted using a formula-based inter-prediction technique; the formula-based inter-prediction technique is based on a formula, and generates a prediction sample of the current block by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block. The format rule specifies: constructing a candidate list, the candidate list includes one or more coding blocks associated with the current block, and the coding blocks in the one or more coding blocks are candidates for providing formula-based inter-prediction information for the formula-based inter-prediction technique. The format rule also specifies: selecting a first coding block from the candidate list, the first coding block being encoded using first formula-based inter-frame prediction information of a formula-based inter-frame prediction technology; determining current formula-based inter-frame prediction information of a current block based on the first formula-based inter-frame prediction information; determining one or more parameter values ​​for applying the formula-based inter-frame prediction technology to one or more parameters of the current block based on the current formula-based inter-frame prediction information; and generating at least one prediction sample of the current block based on a formula including one or more parameters set according to the one or more parameter values.

[0021] Aspects of the present disclosure also provide an apparatus for video encoding or video decoding. The apparatus for video encoding / decoding comprises a processing circuit configured to implement any of the described methods for video encoding or video decoding.

[0022] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer, cause the computer to perform any of the described methods for video decoding / encoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0024] Figure 1 is a schematic illustration of an example of a block diagram of a communication system (100).

[0025] Figure 2 is a schematic illustration of an example of a block diagram of a decoder.

[0026] Figure 3 is a schematic illustration of an example of a block diagram of an encoder.

[0027] Figure 4The positions of spatial merging candidates according to an embodiment of the present disclosure are shown.

[0028] Figure 5 Candidate pairs considered for redundancy check of spatial merging candidates according to an embodiment of the present disclosure are shown.

[0029] Figure 6 Exemplary motion vector scaling for temporal merging candidates is shown.

[0030] Figure 7 Exemplary candidate positions of temporal merging candidates for the current block are shown.

[0031] Figure 8 A diagram of some example templates is shown.

[0032] Fig. 9 A flow chart outlining a decoding process according to some aspects of the present disclosure is shown.

[0033] Fig.10 A flow chart outlining an encoding process according to some aspects of the present disclosure is shown.

[0034] Fig.11 is a schematic illustration of a computer system according to an aspect. DETAILED DESCRIPTION

[0035] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter is equally applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.

[0036] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, which creates, for example, an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples captured by a digital camera. The video picture stream (102), which is depicted as a thick line to emphasize the high amount of data compared to the encoded video data (104) (or the encoded video bitstream), can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. The encoded video data (104) (or the encoded video bitstream), which is depicted as a thin line to emphasize the lower amount of data compared to the video picture stream (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and creates an output video picture stream (111) that can be presented on a display (112) (e.g., a display screen) or other presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of such standards include ITU-T Recommendation H.265. In one example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0037] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).

[0038] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) in an example of FIG.

[0039] The receiver (231) may receive one or more encoded video sequences, for example, included in a bitstream to be decoded by the video decoder (210). In one aspect, the encoded video sequences are received one at a time, wherein the decoding of each encoded video sequence is independent of the decoding of the other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, for example, encoded audio data and / or auxiliary data streams that may be forwarded to their respective consuming entities (not depicted). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other applications, the buffer memory (215) may be located outside the video decoder (210) (not depicted). In still other applications, a buffer memory (not depicted) may be provided outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be provided inside the video decoder (210) to, for example, handle playback timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, the buffer memory (215) may not be required, or the buffer memory (215) may be made smaller. For use on a traffic packet network such as the Internet, the buffer memory (215) may be required, and the buffer memory (215) may be relatively large, advantageously may have an adaptive size, and may be implemented at least partially in an operating system or similar element (not depicted) external to the video decoder (210).

[0040] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210) and potential information for controlling a presentation device such as a presentation device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230), such as Figure 2As shown. The control information for the rendering device may be in the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not depicted). The parser (220) may parse / entropy decode the received coded video sequence. The encoding of the coded video sequence may be performed according to a video coding technique or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) may extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the coded video sequence based on at least one parameter corresponding to the group. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, etc.

[0041] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0042] Depending on the type of coded video picture or part of coded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by parser (220) through subgroup control information parsed from the coded video sequence. For clarity, such subgroup control information flow between parser (220) and the multiple units below is not depicted.

[0043] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following multiple functional units is appropriate.

[0044] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) may output a block including sample values, which may be input into an aggregator (255).

[0045] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture, but may use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by an intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) uses surrounding reconstructed information extracted from a current picture buffer (258) to generate a block of the same size and shape as the block being reconstructed. For example, the current picture buffer (258) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.

[0046] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and potentially motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to extract samples for prediction. After the extracted samples are motion compensated according to the symbols (221) belonging to the block, these samples may be added to the output of the scaler / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The extraction of prediction samples by the motion compensated prediction unit (253) from the address in the reference picture memory (257) may be controlled by a motion vector, which may be provided to the motion compensated prediction unit (253) in the form of symbols (221), which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values ​​extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0047] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video sequence (also referred to as the coded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a coded picture or a previous (in decoding order) portion of the coded video sequence, and to previously reconstructed and loop filtered sample values.

[0048] The output of the loop filter unit (256) may be a sample stream that may be output to a rendering device (212) and stored in a reference picture memory (257) for future inter-picture prediction.

[0049] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed, and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257), and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.

[0050] The video decoder (210) may perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265 Recommendation. The encoded video sequence may conform to the syntax specified by the video compression technology or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all the tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence may also be required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management signaled in the encoded video sequence.

[0051] In one aspect, the receiver (231) can receive additional (redundant) data when receiving the encoded video. The additional data can be included as part of the encoded video sequence. The additional data can be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data can take the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0052] Figure 3 An example of a block diagram of a video encoder (303) is shown. The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example of FIG.

[0053] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In an example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source (301) can capture video images to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).

[0054] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be constructed as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on the samples.

[0055] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required. Implementing an appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to the other functional units described. For clarity, the coupling is not depicted in the figure. The parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be configured to have other suitable functions that are related to the video encoder (303) optimized for a certain system design.

[0056] In some aspects, the video encoder (303) is configured to operate in an encoding loop. As an oversimplified description, in one example, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Since the decoding of the symbol stream produces a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurately corresponding between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values ​​that the decoder will "see" when using the prediction during decoding. This basic principle of reference picture synchronization (and the drift that occurs when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.

[0057] The operation of the "local" decoder (333) may be similar to that described above in conjunction with Figure 2 The "remote" decoder of the video decoder (210) described in detail is identical. However, additional brief reference is made to Figure 2 , since the symbols are available and the entropy encoder (345) and the parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210) including the buffer memory (215) and the parser (220) may not be fully implemented in the local decoder (333).

[0058] On the one hand, except for the parsing / entropy decoding present in the decoder, the decoder technology is present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the decoder operation. The description of the encoder technology can be simplified because the encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.

[0059] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture by referencing one or more previously encoded pictures from a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of a reference picture that may be selected as a prediction reference for the input picture.

[0060] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may advantageously be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may store a copy of the reconstructed reference picture locally that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.

[0061] The predictor (335) may perform a prediction search for the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may be used as appropriate prediction references for the new picture. The predictor (335) may operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory (334).

[0062] The controller (350) may manage encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0063] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into a coded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.

[0064] The transmitter (340) may buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which may be a hardware / software link to a storage device that may store the encoded video data. The transmitter (340) may combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or ancillary data streams (source not shown).

[0065] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:

[0066] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh ("IDR") pictures.

[0067] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses a motion vector and a reference index to predict the sample values ​​of each block.

[0068] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra prediction or inter prediction, which uses two motion vectors and reference indices to predict the sample values ​​of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata to reconstruct a single block.

[0069] The source picture may typically be spatially subdivided into blocks of samples (e.g. blocks of 4×4, 8×8, 4×8 or 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively coded, or blocks of an I picture may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.

[0070] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.

[0071] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0072] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is partitioned into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.

[0073] In some aspects, a bidirectional prediction technique may be used for inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that precede the current picture in the video in decoding order (but may be in the past and future, respectively, in display order). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.

[0074] In addition, merge mode technology can be used for inter-picture prediction to improve coding efficiency.

[0075] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the High-Efficiency Video Coding (HEVC) standard, pictures in a video picture sequence are divided into coding tree units (CTUs) for compression, and the CTUs in the pictures have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Typically, a CTU includes three coding tree blocks (CTBs), which are a luminance CTB and two chrominance CTBs. Each CTU can be recursively divided into one or more coding units (CUs) using a quadtree. For example, a 64×64 pixel CTU can be divided into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. The CU is divided into one or more prediction units (PUs) based on temporal and / or spatial predictability. Typically, each PU includes a luma prediction block (PB) and two chroma PBs. In one aspect, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Using the luma prediction block as an example of a prediction block, the prediction block includes a matrix of values ​​for pixels (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0076] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more integrated circuits. In another aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) may be implemented using one or more processors executing software instructions.

[0077] The present disclosure provides a technique for constructing a list of coding information in inter-frame prediction coding.

[0078] Various inter prediction modes can be used in video coding. For example, in VVC, for an inter-prediction CU, motion parameters may include MV, one or more reference picture indices, a reference picture list usage index, and additional information of certain coding features used to generate inter-prediction samples. Motion parameters may be signaled explicitly or implicitly. When a CU is encoded using skip mode, the CU may be associated with a PU and may not have significant residual coefficients, coded motion vector increments or MV differences (e.g., MVDs), or reference picture indices. A merge mode may be specified, in which the motion parameters for the current CU are obtained from neighboring CUs, including spatial and / or temporal candidates, optionally including additional information (e.g., information introduced in VVC). The merge mode is applicable to inter-prediction CUs, not just to skip mode. In one example, an alternative to merge mode is to explicitly transmit motion parameters, where for each CU, MV, the corresponding reference picture index for each reference picture list, a reference picture list usage flag, and other information are explicitly signaled.

[0079] In one embodiment, for example, in VVC, the VVC Test Model (VTM) reference software includes one or more refined inter-frame prediction coding tools, which include: extended merge prediction, merged motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, sub-block-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bidirectional prediction (BCW) with CU-level weights, bidirectional optical flow (BDOF), prediction refinement (PROF) using optical flow, decoder-side motion vector refinement (DMVR), joint inter-frame and intra-frame prediction (CIIP), geometric partition mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0080] In some examples, extended merge prediction may be used. In one example, such as in VTM4, the merge candidate list is constructed by including the following five types of candidates in order: spatial motion vector prediction (MVP) from spatially adjacent CUs, temporal MVP from co-located CUs, history-based MVP (HMVP) from a first-in-first-out (FIFO) table, pairwise average MVP, and zero MV.

[0081] The size of the merge candidate list can be written into the slice header. In one example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in merge mode, the index of the best merge candidate (e.g., merge index) can be encoded using truncated unary binarization (TU). The first bin of the merge index can be encoded using context (e.g., context adaptive binary arithmetic coding (CABAC)), and bypass encoding can be used for other bins.

[0082] Some examples of the generation process of merge candidates for each category are provided below. In one embodiment, the spatial candidate derivation is as follows. The derivation of spatial merge candidates in VVC can be the same as the derivation of spatial merge candidates in HEVC. In one example, at Figure 4 A maximum of four merge candidates are selected from the candidates for the depicted positions.

[0083] Figure 4 FIG. 4 shows the positions of spatial merging candidates according to an embodiment of the present disclosure. Figure 4 , the order of derivation is B1, A1, B0, A0, B2. Position B2 is considered only when any CU in positions A0, B0, B1, and A1 is unavailable (for example, because the CU belongs to another slice or another tile) or when intra coding is used. After adding the candidate at position A1, a redundancy check is performed when adding the remaining candidates, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0084] In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only Figure 5 The candidate pairs are connected by arrows in , and a candidate is added to the candidate list only if the corresponding candidates used for redundancy check do not have the same motion information.

[0085] Figure 5 FIG. 4 shows candidate pairs considered for redundancy check of spatial merging candidates according to an embodiment of the present disclosure. Figure 5, candidate pairs connected by corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, the candidates at positions B1, A0 and / or B2 can be compared with the candidate at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidate at position B1.

[0086] In one embodiment, the temporal candidates are derived as follows: In one example, only one temporal merge candidate is added to the candidate list. Figure 6 An exemplary motion vector scaling for a temporal merge candidate is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) may be derived based on a collocated CU (612) belonging to a collocated reference picture (604) (e.g., Figure 6 ). The reference picture list used to derive the co-located CU (612) may be explicitly written into the slice header. Figure 6 A scaled MV (621) for a temporal merge candidate is obtained as shown by the dotted line in . The scaled MV (621) can be scaled according to the MV of the co-located CU (612) by using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the co-located reference picture (604) of the co-located picture (603) and the co-located picture (603). The reference picture index of the temporal merge candidate can be set to zero. A co-located picture is a reference picture that is used as a source picture for derivation of temporal motion information. The co-located picture can be identified in one of the two lists (referred to as list 0 or list 1). In some examples, the encoder can determine the co-located picture and signal the co-located picture using appropriate syntax techniques.

[0087] Figure 7 Exemplary candidate positions (e.g., C0 and C1) for temporal merge candidates for the current CU are shown. The position of the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is located at the lower right corner of the co-located CU (710) of the current CU. Candidate position C1 is located at the center of the co-located CU (710) of the current CU. If the CU at candidate position C0 is not available, intra-coded, or located outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and located in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0088] According to some aspects of the present disclosure, a formula-based prediction method may be used for inter-frame prediction or intra-frame prediction. For inter-frame prediction, the formula-based prediction method may use a formula to generate samples of a current block in a current picture based on reference samples of a reference block in a reference picture. For intra-frame prediction, the formula-based prediction method may use a formula to generate a first color component of a current block based on a second color component of the current block. In some examples, parameters in a formula for a formula-based prediction method may be derived based on a template of the current block.

[0089] In some examples, local illumination compensation (LIC) is used as an inter-prediction technique to model local illumination changes between a current block and a prediction block (also referred to as a reference block) of the current block by using a linear function. The prediction block is located in a reference picture and may be pointed to by a motion vector (MV). Parameters of the linear formula may include a scale α and an offset β, which may be expressed as α×p[x, y]+β to compensate for illumination changes, where p[x, y] represents a reference sample at position [x, y] in a reference block (also referred to as a prediction block), which is pointed to by the MV from the current block. In some examples, the scale α and the offset β may be derived based on a template of the current block and a corresponding reference template of the reference block by using a least squares method, so no signaling overhead is required, except that a LIC flag may be signaled to indicate the use of LIC. The scale α and the offset β derived based on the template of the current block may be referred to as a template-based parameter set.

[0090] In some examples, LIC is used for unidirectionally predicted inter-frame CUs. In some examples, intra-frame neighboring samples of the current block (neighboring samples predicted using intra-frame prediction) can be used in LIC parameter derivation. In some examples, LIC is disabled for blocks whose luma samples are less than 32. In some examples, for non-sub-block modes (e.g., non-affine modes), LIC parameter derivation is performed based on template block samples of the current CU (rather than partial template block samples of the first upper left 16×16 unit). In some examples, LIC parameter derivation is performed based on partial template block samples (e.g., partial template block samples of the first upper left 16×16 unit). In some examples, the template samples of the reference block are determined by using motion compensation (MC) on the MV of the block, without rounding the MV of the block to integer pixel precision.

[0091] In some examples, cross-component prediction can be used as an intra-prediction technique. Cross-component prediction can include a first technique called a cross-component linear model (CCLM), a second technique called a multi-model linear model (MMLM), a third technique called a convolutional cross-component model (CCCM), and a fourth technique called a gradient linear model (GLM).

[0092] For example, the first technique CCLM is used to reduce cross-component redundancy. In CCLM, based on the reconstructed luma samples of the same CU, chroma samples are predicted by using a linear model (also called a linear formula), for example, using equation (1): pred c (i,j)=a·rec′ L (i,j)+b Equation (1) Among them, pred c (i,j) represents the predicted chroma sample in CU, rec L ′(i,j) represents the downsampled reconstructed luma sample of the same CU. The CCLM linear model includes parameters (a and b), which in one example can be derived from up to four adjacent chroma samples and their corresponding downsampled luma samples. In one example, up to four adjacent chroma samples and their corresponding downsampled luma samples are called a template of the CU.

[0093] In some examples, based on the positions of adjacent chroma samples, CCLM may include different modes, which are referred to as LM_T (LM top mode or upper mode LM_A), LM_L (LM left mode), and LM_LT (LM left top mode, or upper left mode LM_LA, or just LM mode). For example, if the size of the current chroma block is W×H, then in CCLM, W' and H' may be set for various modes. When the LM mode (also referred to as LM_LT or LM_LA) is applied, W'=W, H'=H. When the LM-A mode is applied, W'=W+H. When the LM-L mode is applied, H'=H+W.

[0094] It should be noted that MMLM, CCCM and GLM also use functions for prediction. The parameters of the function can be derived based on the template.

[0095] It should be noted that the following description uses inter-frame prediction to illustrate the technique for deriving encoding information for a formula-based prediction method, and that the technique can be appropriately used to derive encoding information for intra-frame prediction.

[0096] According to some aspects of the present disclosure, some inter-frame prediction techniques are designed to minimize the distortion between the current block and the prediction block of the current block in the corresponding reference picture. For example, the inter-frame prediction technique (also referred to as the first method of the inter-frame prediction method, the formula-based inter-frame prediction technique, the function-based inter-frame prediction technique, or the model-based inter-frame prediction technique) may apply a formula (e.g., a nonlinear formula, a linear formula, etc.) to generate the current block in the current picture using the original prediction block in the reference picture as an input of the formula. For example, the inter-frame prediction technique may generate a prediction of a sample in the current block based on a formula that takes one or more prediction samples in the reference picture as an input. The formula may include linear terms or nonlinear terms, and may include one or more parameters, which may be derived. It should be noted that LIC is one of such inter-frame prediction techniques.

[0097] In some examples, the formula is a linear formula and can be expressed as It means that n is a non-negative integer, p(x i ,y i ) is the position (x) in the reference image i ,y i ), the prediction sample is pointed to based on the MV associated with the current block. Further, a set of prediction samples (given by p(x i ,y i ), where i=0, ..., n) may be a set of prediction samples around the corresponding sample in the reference sample pointed to by the current sample to be predicted by MV. In some examples, the parameter α may be derived based on a template of the current block (also referred to as the current block template) and a template of a prediction block of the current block (also referred to as the prediction block template) (e.g., by using a least squares method) by minimizing the difference between the current block template and its prediction block template. i The template of the current block consists of the spatially adjacent reconstructed samples of the current block, and the template of the prediction block consists of the spatially adjacent reconstructed samples of the prediction block.

[0098] Figure 8 The diagrams of templates in some examples are shown. For example, the template (810) is called an L-shaped template T L , and includes adjacent samples located in a row above, a column to the left, and the upper left corner of the current block (also referred to as the current coding block). The template (820) is called the upper and left template T a+l , and includes adjacent samples located in one row above and one column to the left of the current block. The template (830) is called the upper template T a , and includes the adjacent samples located in the row above the current block. The template (840) is called the left template T l , and includes the neighboring samples located in the left column of the current block. It should be noted that the template may include Figure 8 Adjacent samples of other suitable shapes not shown.

[0099] In some examples, multiple candidate template types may be supported, and one candidate template type may be selected to derive the parameters of the linear formula. Syntax may be written in the code stream (eg, at the block level) to indicate which candidate template type is selected.

[0100] In some examples, a control flag associated with a formula-based inter-frame prediction technique may be written in a code stream (e.g., at a block level) to indicate whether the formula-based inter-frame prediction technique is applied to the current block. Alternatively, the value of the control flag may also be inherited from other coding blocks. More specifically, a first control flag associated with the current block of the formula-based inter-frame prediction technique is inherited from a second control flag associated with another one or more coding blocks of the formula-based inter-frame prediction technique. In addition, a control flag may be derived at the coding block level to adaptively determine whether to apply a formula-based inter-frame prediction technique.

[0101] In some examples, a first control flag associated with a current block for a formula-based inter-frame prediction technique is inherited from a second control flag associated with another one or more coding blocks for a formula-based inter-frame prediction technique. In some examples, encoding information for the formula-based inter-frame prediction technique may be derived from an adjacent coding block, a non-adjacent coding block, or a coding block storing encoding information in a buffer.

[0102] It should be noted that in the present disclosure, no general limitation is made. In one example, the term "parameter" refers to the parameter α used to determine the linear formula for deriving the prediction block. i and β. The term “template type” refers to different template shapes, such as but not limited to Figure 8 One of the template types shown for parameter derivation in nonlinear or linear formulas.

[0103] Some aspects of the present disclosure provide improved techniques for information derivation for formula-based inter-frame prediction techniques, and can improve the coding efficiency of inter-frame prediction. For example, an encoder / decoder may construct a candidate list, the candidate list including one or more coding blocks associated with a current block. In some examples, the one or more coding blocks include at least one non-adjacent spatially adjacent block of the current block. The coding blocks in the one or more coding blocks are candidates for providing formula-based inter-frame prediction information for formula-based inter-frame prediction techniques. Further, a first coding block is selected from the candidate list, and the first coding block is encoded using a first formula-based inter-frame prediction information for a formula-based inter-frame prediction technique. Then, the current formula-based inter-frame prediction information of the current block is determined based on the first formula-based inter-frame prediction information, and one or more parameter values ​​of one or more parameters for applying the formula-based inter-frame prediction technique to the current block can be derived based on the current formula-based inter-frame prediction information.

[0104] Some aspects of the present disclosure provide techniques for constructing a first-in, first-out (FIFO) queue for storing formula-based inter-frame prediction information for a coded block, for example, historical formula-based inter-frame prediction information for a block encoded using a formula-based inter-frame prediction technique. The formula-based inter-frame prediction information for a coded block refers to information that enables the use of a formula-based inter-frame prediction technique for the coded block, such as a template type, a formula type, etc. In some examples, when a current block is reconstructed using a formula-based inter-frame prediction technique during encoding / decoding (e.g., a control flag for the formula-based inter-frame prediction technique is true), information of the formula-based inter-frame prediction technique for the current block is stored in the FIFO queue and can be used to encode / decode subsequent blocks.

[0105] In some examples, the FIFO queue is defined to have a fixed size. In one example, when the FIFO queue has space, for example, when the occupied size of the FIFO queue is less than the fixed size of the FIFO queue, information of the formula-based inter-frame prediction technique of the current block that has been encoded using the formula-based inter-frame prediction technique can be pushed into the FIFO queue. In another example, when the current block is encoded (e.g., encoded or decoded) using the formula-based inter-frame prediction technique, and the total occupied size of the stored information of the formula-based inter-frame prediction technique is equal to the fixed size of the FIFO queue, the oldest information of the formula-based inter-frame prediction technique can be popped out before inserting new information of the formula-based inter-frame prediction technique of the current block.

[0106] In some embodiments, the FIFO queue of the formula-based inter-frame prediction technique is reset at the beginning of frame encoding / decoding. In one example, at the beginning of frame encoding / decoding, the FIFO queue is reset to empty.

[0107] In some embodiments, the FIFO queue of the formula-based inter-frame prediction technique is reset at the beginning of a row of a coding tree unit (CTU) or a super block (SB). For example, at the beginning of a row of a CTU or a SB, the FIFO queue of the formula-based inter-frame prediction technique is reset to empty.

[0108] In some embodiments, the FIFO queue of the formula-based inter-frame prediction technique is reset at the slice level, the tile level, or the segment level. In one example, at the beginning of the slice, the FIFO queue of the formula-based inter-frame prediction technique is reset to empty. In another example, at the beginning of the tile, the FIFO queue of the formula-based inter-frame prediction technique is reset to empty. In another example, at the beginning of the segment, the FIFO queue of the formula-based inter-frame prediction technique is reset to empty.

[0109] Some aspects of the present disclosure provide techniques for constructing a candidate list using formula-based inter-frame prediction information, where the formula-based inter-frame prediction information comes from a coding block associated with a current block, such as one or more spatially adjacent coding blocks of the current block, one or more non-adjacent spatially adjacent coding blocks of the current block, a temporally co-located coding block of a reference picture from the current block, a FIFO queue of a formula-based inter-frame prediction technique, etc. In some examples, when a formula-based inter-frame prediction technique is applied to the current block, an index may be written in a bitstream to indicate which coding block in the candidate list provides the formula-based inter-frame prediction information to be applied to the current block.

[0110] In some embodiments, a flag (e.g., a list usage flag) is written in the codestream to indicate whether to use the candidate list to provide formula-based inter-frame prediction information for the current block. When the flag is true, the candidate list is used, and then an index is written in the codestream to indicate which coding block in the candidate list provides the formula-based inter-frame prediction information to be applied to the current block. Otherwise, the flag is false, the candidate list is not used, and the formula-based inter-frame prediction information to be applied to the current block can be derived and / or written in the codestream. In one example, the formula-based inter-frame prediction information to be applied to the current block is written directly in the codestream. In another example, the formula-based inter-frame prediction information to be applied to the current block is derived in the absence of additional signaling in the codestream.

[0111] In one embodiment, a candidate list of formula-based inter-frame prediction information is constructed according to at least one of the following candidates in a predetermined scanning order: adjacent coding blocks of adjacent spaces, FIFO queues, adjacent coding blocks of non-adjacent spaces, and temporally co-located coding blocks from reference pictures. In some examples, the candidate list includes coding blocks in the following order: adjacent coding blocks of one or more adjacent spaces of the current block, a FIFO queue storing historical data of formula-based inter-frame prediction information, adjacent coding blocks of one or more non-adjacent spaces, and temporally co-located coding blocks from reference pictures.

[0112] In one embodiment, adjacent coding blocks in adjacent spaces are first scanned, and then inserted into the first FIFO queue when the first FIFO queue is not empty. Further, it is checked whether adjacent coding blocks in one or more non-adjacent spaces can be used to provide formula-based inter-frame prediction information, and then it is checked whether temporally co-located coding blocks from reference pictures can be used to provide formula-based inter-frame prediction information.

[0113] In some examples, a list usage flag is signaled to indicate whether the candidate list is used to provide formula-based inter-frame prediction information for the current block. When the flag is true, the candidate list is used, and then an index (also referred to as a first index) is written in the bitstream to indicate which coding block in the candidate list is used to provide formula-based inter-frame prediction information for the current block. Otherwise, when the flag is false, the derived and / or signaled formula-based inter-frame prediction information is used for the current block. In one example, when the flag is false, the formula-based inter-frame prediction information to be applied to the current block is decoded from the bitstream. In another example, when the flag is false, the formula-based inter-frame prediction information to be applied to the current block is derived in the absence of additional signaling in the bitstream.

[0114] In some embodiments, the candidate list may be reordered in ascending order based on the template matching cost value. In some examples, the encoding information of each candidate is applied to the template in each candidate prediction block to calculate the template matching cost. For example, for a candidate in the candidate list (e.g., a coding block), the formula-based inter-frame prediction information of the candidate is used to determine the candidate formula to be used in the formula-based inter-frame prediction technology of the current block. Further, the reconstructed samples of the reference template of the reference block for the current block can be input into the candidate formula to generate a candidate template. In one example, a difference measure between the candidate template and the current template of the current block is calculated as the template matching cost value of the candidate. In some examples, the candidates in the candidate list (e.g., coding blocks) are reordered in ascending order of the template matching cost values ​​associated with the candidates.

[0115] According to some aspects of the present disclosure, the functions of the list usage flag and the index may be combined into a single index (also referred to as a second index). In some embodiments, the formula-based inter-frame prediction information derived and / or signaled is placed at the beginning of the candidate list, for example as a first candidate corresponding to an index value of 0. After the first candidate, one or more coding blocks are placed in the candidate list. Then, the second index is signaled to indicate which candidate in the candidate list is used. In one example, when the value of the second index is 0, the formula-based inter-frame prediction information derived and / or signaled is used to apply the formula-based inter-frame prediction technique to the current block. When the second index is greater than 0, a coding block is selected from one or more coding blocks in the candidate list to provide formula-based inter-frame prediction information for applying the formula-based inter-frame prediction technique to the current block.

[0116] In one embodiment, the candidate list may be reordered in ascending order based on the cost of template matching, and the derived and / or signaled formula-based inter-frame prediction information is always placed at the beginning of the candidate list. In some examples, the candidate list includes a first candidate at the beginning of the candidate list, and one or more coding blocks as candidates after the first candidate. The first candidate indicates the derived and / or signaled formula-based inter-frame prediction information. One or more coding blocks may be reordered in ascending order of template matching cost values ​​associated with one or more coding blocks. For example, for a coding block in the candidate list, the formula-based inter-frame prediction information of the coding block is used to determine a candidate formula to be used in a formula-based inter-frame prediction technique for a current block. Further, a reconstructed sample of a reference template of a reference block for the current block may be input into the candidate formula to generate a candidate template. In one example, a difference metric between the candidate template and the current template of the current block is calculated as a template matching cost value for the coding block. In some examples, one or more coding blocks in the candidate list are reordered in ascending order of template matching cost values ​​associated with the one or more coding blocks.

[0117] Fig. 9 A flow chart outlining a process (900) according to an aspect of the present disclosure is shown. The process (900) may be used in a video decoder. In various aspects, the process (900) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (900) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (900). The process starts at (S901) and proceeds to (S910).

[0118] At (S910), a code stream of encoding information for a picture sequence is received, the encoding information indicating that a current block in a current picture is inter-predicted using a formula-based inter-prediction technique. The formula-based inter-prediction technique is based on a formula, and a prediction sample of the current block is generated by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block.

[0119] At (S920), a candidate list is constructed, the candidate list including one or more coding blocks associated with the current block. Coding blocks in the one or more coding blocks are candidates for providing formula-based inter prediction information of the formula-based inter prediction technique.

[0120] At (S930), a first coding block is selected from the candidate list, the first coding block being encoded using first formula-based inter prediction information of the formula-based inter prediction technique.

[0121] At (S940), current formula-based inter-frame prediction information of the current block is determined based on the first formula-based inter-frame prediction information. For example, the current block inherits the first formula-based inter-frame prediction information from the first coding block, and therefore, the current formula-based inter-frame prediction information is set to be the same as the first formula-based inter-frame prediction information.

[0122] At (S950), one or more parameter values ​​for one or more parameters for applying a formula-based inter prediction technique to the current block are derived based on the current formula-based inter prediction information.

[0123] At (S960), at least one prediction sample of the current block is generated based on a formula including one or more parameters set according to one or more parameter values.

[0124] In some examples, the formula-based inter prediction technique is local illumination compensation (LIC).In some examples, the one or more coding blocks include at least one of: a neighboring coding block in an adjacent space, a neighboring block in a non-adjacent space, and / or a temporally co-located coding block of the current block in a reference picture.

[0125] In some embodiments, the candidate list is constructed according to a first-in-first-out (FIFO) queue that stores historical formula-based inter prediction information of one or more historically encoded blocks that were encoded using a formula-based inter prediction technique before the current block.

[0126] In some examples, the current formula-based inter-frame prediction information applied to the current block is stored in a FIFO queue in a FIFO manner. When the number of stored historical coding blocks is equal to the size of the FIFO queue, the oldest historical coding block is popped out and the current block with the current formula-based inter-frame prediction information is inserted into the FIFO queue.

[0127] In one example, the FIFO queue is reset at the beginning of a frame (e.g., the FIFO queue is set to empty). In another example, the FIFO queue is reset at the beginning of a CTU row. In another example, the FIFO queue is reset at the beginning of a superblock (SB) row. In another example, the FIFO queue is reset at the beginning of a slice. In another example, the FIFO queue is reset at the beginning of a tile. In another example, the FIFO queue is reset at the beginning of a fragment.

[0128] In some examples, a first flag is decoded from a bitstream, the first flag indicating whether a candidate list is used to determine current formula-based inter-frame prediction information. When the first flag is true, an index is decoded from the bitstream, the index indicating a first coding block in the candidate list. The first formula-based inter-frame prediction information of the first coding block is used to determine the current formula-based inter-frame prediction information. When the first flag is false, the current formula-based inter-frame prediction information is determined without using the candidate list. In one example, the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique is decoded directly from the bitstream. In another example, the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique is derived in the absence of additional signaling in the bitstream.

[0129] In some examples, a candidate list is constructed based on a scanning order of at least one of at least one adjacent spatial adjacent coding block, a FIFO queue storing historical coding blocks by a formula-based inter-frame prediction technique, at least one non-adjacent spatial adjacent block, and / or a temporal co-location coding block of a current block in a reference picture. In one example, it is checked whether the adjacent spatial adjacent coding block is encoded using a formula-based inter-frame prediction technique. When the adjacent spatial adjacent coding block is encoded using a formula-based inter-frame prediction technique, the adjacent spatial adjacent coding block is inserted into the candidate list. Further, it is checked whether the FIFO queue is empty. When the FIFO queue is not empty, the FIFO queue is inserted into the candidate list. Further, it is checked whether the non-adjacent spatial adjacent blocks are encoded using a formula-based inter-frame prediction technique. When the non-adjacent spatial adjacent blocks are encoded using a formula-based inter-frame prediction technique, the non-adjacent spatial adjacent blocks are inserted into the candidate list. Then, it is checked whether the temporal co-location coding block of the current block in the reference picture is encoded using a formula-based inter-frame prediction technique. When the temporal co-location coding block is encoded using a formula-based inter-frame prediction technique, the temporal co-location coding block is inserted into the candidate list.

[0130] In some examples, template matching costs are calculated for one or more coding blocks respectively. For example, for a coding block that is a candidate in a candidate list, a candidate formula may be determined based on formula-based inter-frame prediction information associated with the coding block. Then, a template (reconstructed sample in the template area) of a reference block (prediction block) is input into the candidate formula to generate a test template for the current block. In some examples, a difference measure between the test template of the current block and the current template is calculated as a template matching cost associated with the coding block.

[0131] In some examples, one or more coding blocks in the candidate list may be reordered according to the template matching cost. For example, one or more coding blocks may be reordered in ascending order according to the template matching cost.

[0132] In some examples, the first flag is not used. A candidate list is constructed to have a first candidate located at the beginning of the candidate list and one or more coding blocks located after the first candidate, the first candidate indicating the current formula-based inter-frame prediction information using derivation and / or direct signaling. For example, an index is decoded from a bitstream. When the index indicates the first candidate, the current formula-based inter-frame prediction information is determined without using the formula-based inter-frame prediction information of one or more coding blocks. In one example, when the index indicates the first candidate, the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technology is decoded directly from the bitstream. In another example, when the index indicates the first candidate, the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technology is derived in the absence of additional signaling in the bitstream.

[0133] In some examples, template matching costs for one or more coding blocks are calculated respectively, and one or more coding blocks located after the first candidate in the candidate list are reordered according to the template matching costs.

[0134] In some examples, the formula-based inter prediction information includes at least one of a template type and a formula type.

[0135] Then, the process proceeds to (S999) and terminates.

[0136] The process (900) may be adapted as appropriate. Steps in the process (900) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0137] Fig.10 A flow chart outlining a process (1000) according to an aspect of the present disclosure is shown. The process (1000) may be used in a video encoder. In various aspects, the process (1000) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. According to some aspects, the process (1000) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (1000). The process starts at (S1001) and proceeds to (S1010).

[0138] At (S1010), it is determined to encode a current block in a current picture by inter-prediction, which has a potential use of a formula-based inter-prediction technique. The formula-based inter-prediction technique is based on a formula to generate a prediction sample of the current block by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block.

[0139] At (S1020), a candidate list is constructed, the candidate list including one or more coding blocks associated with the current block. Coding blocks in the one or more coding blocks are candidates for providing formula-based inter prediction information of the formula-based inter prediction technique for inter prediction of the current block.

[0140] At (S1030), current formula-based inter prediction information for the current block is determined based on the candidate list. In some examples, the current formula-based inter prediction information may be determined from the candidate list using an evaluation such as a rate-distortion evaluation.

[0141] At (S1040), the current block is encoded according to the current formula-based inter-frame prediction information. Coding information of the current block to be included in the code stream is generated. For example, the coding information may include a flag and / or an index indicating a candidate in the candidate list that provides the current formula-based inter-frame prediction information.

[0142] In some examples, the formula-based inter prediction technique is local illumination compensation (LIC).In some examples, the one or more coding blocks include at least one of a neighboring spatial neighboring coding block, a non-neighboring spatial neighboring block, and / or a temporally co-located coding block of the current block in a reference picture.

[0143] In some embodiments, the candidate list is constructed according to a first-in-first-out (FIFO) queue that stores historical formula-based inter prediction information of one or more historically encoded blocks that were encoded using a formula-based inter prediction technique before the current block.

[0144] In some examples, the current formula-based inter-frame prediction information applied to the current block is stored in a FIFO queue in a FIFO manner. When the number of stored historical coding blocks is equal to the size of the FIFO queue, the oldest historical coding block is popped out and the current block with the current formula-based inter-frame prediction information is inserted into the FIFO queue.

[0145] In one example, the FIFO queue is reset at the beginning of a frame (e.g., the FIFO queue is set to empty). In another example, the FIFO queue is reset at the beginning of a CTU row. In another example, the FIFO queue is reset at the beginning of a superblock (SB) row. In another example, the FIFO queue is reset at the beginning of a slice. In another example, the FIFO queue is reset at the beginning of a tile. In another example, the FIFO queue is reset at the beginning of a fragment.

[0146] In some examples, it is determined to use first formula-based inter-frame prediction information of a first coding block from a candidate list as current formula-based inter-frame prediction information of a current block. Then, a first flag with a true value and an index are encoded in a bitstream carrying coding information of a picture sequence, wherein the first flag with a true value indicates that the current formula-based inter-frame prediction information is determined from the candidate list, and the index indicates the first coding block.

[0147] In some examples, it is determined that one or more coding blocks from the candidate list are not used to determine the current formula-based inter-frame prediction information. Then, a first flag with a false value is encoded in the bitstream, and the first flag with a false value indicates that the current formula-based inter-frame prediction information is not determined based on the one or more coding blocks in the candidate list.

[0148] In one example, when the first flag has a false value, current formula-based inter prediction information of the formula-based inter prediction technique is directly encoded in the code stream.

[0149] In some examples, a candidate list is constructed based on a scanning order of at least one of at least one adjacent spatial adjacent coding block, a FIFO queue storing historical coding blocks by a formula-based inter-frame prediction technique, at least one non-adjacent spatial adjacent block, and / or a temporal co-location coding block of a current block in a reference picture. In one example, it is checked whether the adjacent spatial adjacent coding block is encoded using a formula-based inter-frame prediction technique. When the adjacent spatial adjacent coding block is encoded using a formula-based inter-frame prediction technique, the adjacent spatial adjacent coding block is inserted into the candidate list. Further, it is checked whether the FIFO queue is empty. When the FIFO queue is not empty, the FIFO queue is inserted into the candidate list. Further, it is checked whether the non-adjacent spatial adjacent blocks are encoded using a formula-based inter-frame prediction technique. When the non-adjacent spatial adjacent blocks are encoded using a formula-based inter-frame prediction technique, the non-adjacent spatial adjacent blocks are inserted into the candidate list. Then, it is checked whether the temporal co-location coding block of the current block in the reference picture is encoded using a formula-based inter-frame prediction technique. When the temporal co-location coding block is encoded using a formula-based inter-frame prediction technique, the temporal co-location coding block is inserted into the candidate list.

[0150] In some examples, template matching costs are calculated for one or more coding blocks respectively. For example, for a coding block that is a candidate in a candidate list, a candidate formula may be determined based on formula-based inter-frame prediction information associated with the coding block. Then, a template (reconstructed sample in the template area) of a reference block (prediction block) is input into the candidate formula to generate a test template for the current block. In some examples, a difference measure between the test template of the current block and the current template is calculated as a template matching cost associated with the coding block.

[0151] In some examples, one or more coding blocks in the candidate list may be reordered according to the template matching cost. For example, one or more coding blocks may be reordered in ascending order according to the template matching cost.

[0152] In some examples, the first flag is not used. A candidate list is constructed to have a first candidate at the beginning of the candidate list and one or more coding blocks after the first candidate, the first candidate indicating current formula-based inter-frame prediction information using derivation and / or direct signaling. For example, the index is encoded in the bitstream. When the index indicates the first candidate, the current formula-based inter-frame prediction information is determined without using the formula-based inter-frame prediction information of one or more coding blocks. In one example, when the index indicates the first candidate, the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique is directly encoded in the bitstream.

[0153] In some examples, template matching costs for one or more coding blocks are calculated respectively, and one or more coding blocks located after the first candidate in the candidate list are reordered according to the template matching costs.

[0154] In some examples, the formula-based inter prediction information includes at least one of a template type and a formula type.

[0155] Then, the process proceeds to (S1099) and terminates.

[0156] The process (1000) may be adapted as appropriate. Steps in the process (1000) may be modified and / or omitted. Additional steps may be added. Any suitable order of implementation may be used.

[0157] According to one aspect of the present disclosure, a method for processing visual media data is provided. In the method, a code stream of the visual media data is processed according to a format rule. For example, the code stream may be a code stream decoded / encoded in any of the decoding and / or encoding methods described herein. The format rule may specify one or more constraints of the code stream and / or one or more processes to be performed by a decoder and / or encoder.

[0158] In one example, a code stream includes encoding information of a picture sequence, the encoding information indicating that a current block in a current picture is inter-predicted using a formula-based inter-prediction technique, the formula-based inter-prediction technique is based on a formula, and a prediction sample of the current block is generated by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block. The format rule specifies: constructing a candidate list, the candidate list including one or more coding blocks associated with the current block, the coding blocks in the one or more coding blocks being candidates for providing formula-based inter-prediction information of the formula-based inter-prediction technique. The format rule also specifies: selecting a first coding block from the candidate list, the first coding block being encoded using first formula-based inter-prediction information of the formula-based inter-prediction technique. The format rule further specifies: determining current formula-based inter-frame prediction information for the current block based on the first formula-based inter-frame prediction information; determining one or more parameter values ​​for applying the formula-based inter-frame prediction technique to the current block based on the current formula-based inter-frame prediction information; and reconstructing at least one sample of the current block based on the formula-based inter-frame prediction technique using the one or more parameter values. For example, generating at least one predicted sample of the current block based on a formula including one or more parameters set according to the one or more parameter values.

[0159] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Fig.11 A computer system (1100) suitable for implementing certain aspects of the disclosed subject matter is shown.

[0160] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to assembly, compilation, linking or similar mechanisms to create code comprising instructions, which may be executed directly by one or more computer central processing units (CPU), graphics processing units (GPU), etc., or through interpretation, microcode execution, etc.

[0161] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, and the like.

[0162] Fig.11The components of the computer system (1100) shown are examples and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one or combination of components shown in the exemplary aspects of the computer system (1100).

[0163] The computer system (1100) may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to human conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, captured images obtained from a still image camera), and videos (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0164] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data gloves (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).

[0165] The computer system (1100) may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include tactile output devices (e.g., tactile feedback of a touch screen (1110), a data glove (not shown), or a joystick (1105), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (1109), headphones (not depicted)), visual output devices (e.g., screens (1110) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs exceeding three dimensions through devices such as stereoscopic image output, virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), and printers (not depicted).

[0166] The computer system (1100) may also include human-machine accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with CD / DVD and other media (1121), thumb drives (1122), removable hard drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security software dogs (not depicted), etc.

[0167] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transitory signals.

[0168] The computer system (1100) may also include an interface (1154) to one or more communication networks (1155). The network may be, for example, a wireless network, a wired network, an optical network. The network may further be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CANBus, etc. Some networks typically require an external network interface adapter attached to some common data port or peripheral bus (1149) (e.g., a USB port of the computer system (1100)); other network interfaces are typically integrated into the kernel of the computer system (1100) by attaching to a system bus as described below (e.g., connected to an Ethernet interface in a PC computer system or connected to a cellular network interface in a smartphone computer system). The computer system (1100) can use any of these networks to communicate with other entities. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus connected to certain CANBus devices), or bidirectional, for example, using a LAN or WAN digital network to connect to other computer systems. Certain protocols and protocol stacks may be used on each of those networks and network interfaces as described above.

[0169] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the kernel ( 1140 ) of the computer system ( 1100 ).

[0170] The kernel (1140) may include one or more central processing units (CPUs) (1141), graphics processing units (GPUs) (1142), dedicated programmable processing units in the form of field programmable gate areas (FPGAs) (1143), hardware accelerators (1144) for certain tasks, graphics adapters (1150), etc. These devices, as well as read-only memory (ROM) (1145), random access memory (1146), internal mass storage (1147) such as internal non-user accessible hard drives, SSDs, etc., may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly to the kernel's system bus (1148) or to the kernel's system bus (1148) via a peripheral bus (1149). In one example, a screen (1110) may be connected to a graphics adapter (1150). The architecture of the peripheral bus includes PCI, USB, etc.

[0171] The CPU (1141), GPU (1142), FPGA (1143) and accelerator (1144) can execute certain instructions, which can be combined to form the computer code mentioned above. The computer code can be stored in ROM (1145) or RAM (1146). Transition data can also be stored in RAM (1146), while permanent data can be stored, for example, in internal mass storage (1147). Fast storage and retrieval to any storage device can be performed by using a cache, which can be closely associated with one or more CPUs (1141), GPUs (1142), mass storage (1147), ROM (1145), RAM (1146), etc.

[0172] The computer readable medium may have thereon computer codes for performing various computer-implemented operations. The medium and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer codes may be of a type well known and available to those skilled in the art of computer software.

[0173] As an example, and not by way of limitation, a computer system (1100) having an architecture, and in particular a kernel (1140), can provide functionality because one or more processors (including CPUs, GPUs, FPGAs, accelerators, etc.) execute software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as some non-temporary kernel (1140) memories, such as a kernel internal mass storage (1147) or ROM (1145). Software implementing various aspects of the present disclosure can be stored in such devices and executed by the kernel (1140). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can cause the kernel (1140), in particular the processors therein (including CPUs, GPUs, FPGAs, etc.) to perform specific processes described herein or to perform specific parts of specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, a computer system may provide functionality due to logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1144)) that may replace or operate in conjunction with software to perform specific processes described herein or specific portions of specific processes described herein. Where appropriate, references to portions of software may include logic and vice versa. Where appropriate, references to portions of computer-readable media may include circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.

[0174] As used in this disclosure, "at least one of" or "one of" is intended to include any one or combination of the listed elements. For example, references to at least one of A, B, or C, at least one of A, B, and C, at least one of A, B, and / or C, and at least one of A to C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B, and one of A and B are intended to include A or B or (A and B). Where applicable, the use of "one of" does not exclude any combination of the listed elements, such as when the elements are not mutually exclusive.

[0175] Although the present disclosure has described several examples of various aspects, there are changes, permutations, and various replacement equivalents that fall within the scope of the present disclosure. Therefore, it should be appreciated that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.

Claims

1. A method for video decoding, comprising: Receiving a bitstream of coding information for a picture sequence, the coding information indicating that a formula-based inter-frame prediction technique is used to perform inter-frame prediction on a current block in a current picture; The formula-based inter-frame prediction technique generates prediction samples of the current block based on a formula by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, wherein the formula includes one or more parameters derived based on a current template of the current block and a reference template of the reference block; constructing a candidate list, the candidate list comprising one or more coding blocks associated with the current block, wherein coding blocks in the one or more coding blocks are candidates for providing formula-based inter-frame prediction information of the formula-based inter-frame prediction technique; Selecting a first coding block from the candidate list, the first coding block being encoded using first formula-based inter-frame prediction information of the formula-based inter-frame prediction technique; Determining current formula-based inter-frame prediction information of the current block based on the first formula-based inter-frame prediction information; deriving, based on the current formula-based inter-frame prediction information, one or more parameter values ​​for applying the formula-based inter-frame prediction technique to the one or more parameters of the current block; as well as At least one prediction sample of the current block is generated based on the formula including the one or more parameters set according to the one or more parameter values.

2. The method according to claim 1, wherein: The formula-based inter prediction technique is local illumination compensation (LIC), and the one or more coding blocks include at least one of a neighboring spatially neighboring coding block, a non-neighboring spatially neighboring block, and / or a temporally co-located coding block of the current block in the reference picture.

3. The method according to claim 1, wherein: Constructing the candidate list further includes: The candidate list is constructed according to a first-in-first-out (FIFO) queue, which stores historical formula-based inter-frame prediction information of one or more historical encoded blocks encoded using the formula-based inter-frame prediction technique before the current block.

4. The method according to claim 3, further comprising: The current formula-based inter prediction information applied to the current block is stored in the FIFO queue in a FIFO manner.

5. The method according to claim 3, further comprising at least one of the following: Resetting the FIFO queue at the beginning of a frame; resetting the FIFO queue at the beginning of a row of a coding tree unit (CTU); resetting the FIFO queue at the beginning of a row of a super block (SB); Resetting the FIFO queue at the beginning of a slice; resetting the FIFO queue at the start of a tile; and / or The FIFO queue is reset at the beginning of a segment.

6. The method according to claim 1, wherein: Selecting the first coding block includes: Decoding a first flag from the code stream; When the first flag is true, decoding an index from the code stream, the index indicating the first coding block in the candidate list; and When the first flag is false, the current formula-based inter prediction information is determined without using the candidate list.

7. The method according to claim 6, wherein: When the first flag is false, the method includes at least one of the following: Decoding the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique from the bitstream; and / or The current formula-based inter prediction information of the formula-based inter prediction technique is derived in the absence of additional signaling in the bitstream.

8. The method according to claim 1, wherein: Constructing the candidate list further includes: The candidate list is constructed based on a scanning order of at least one adjacent spatially adjacent coding block, a first-in-first-out (FIFO) queue storing previous coding blocks by the formula-based inter-frame prediction technique, at least one non-adjacent spatially adjacent block and / or at least one of the temporally co-located coding blocks of the current block in the reference picture.

9. The method according to claim 8, wherein: Constructing the candidate list further includes: checking whether the adjacent spatially adjacent coding block is encoded using the formula-based inter-frame prediction technique; When the adjacent spatially adjacent coding block is encoded using the formula-based inter-frame prediction technique, inserting the adjacent spatially adjacent coding block into the candidate list; Check whether the FIFO queue is empty; When the FIFO queue is not empty, inserting the FIFO queue into the candidate list; checking whether the non-adjacent spatial neighboring block is encoded using the formula-based inter-frame prediction technique; When the non-adjacent spatial neighboring block is encoded using the formula-based inter-frame prediction technique, inserting the non-adjacent spatial neighboring block into the candidate list; checking whether the temporally co-located coded block of the current block in the reference picture is encoded using the formula-based inter prediction technique; and When the temporally co-located coded block is encoded using the formula-based inter-frame prediction technique, the temporally co-located coded block is inserted into the candidate list.

10. The method according to claim 1, wherein: Constructing the candidate list further includes: respectively calculating the template matching costs for the one or more coding blocks; and The one or more coding blocks in the candidate list are reordered according to the template matching cost.

11. The method according to claim 1, wherein: The candidate list includes a first candidate located at a start position of the candidate list and the one or more coding blocks located after the first candidate, the first candidate indicating the current formula-based inter prediction information using derivation and / or direct signaling, and selecting the first coding block includes: Decoding an index from the codestream; and When the index indicates the first candidate, the current formula-based inter prediction information is determined without using the formula-based inter prediction information of the one or more coding blocks.

12. The method according to claim 11, wherein: When the index indicates the first candidate, the method includes at least one of the following: Decoding the current formula-based inter-frame prediction information of the formula-based inter-frame prediction technique from the bitstream; and / or The current formula-based inter prediction information of the formula-based inter prediction technique is derived without additional signaling from the codestream.

13. The method according to claim 11, further comprising: Calculating template matching costs for the one or more coding blocks respectively; as well as The one or more coding blocks located after the first candidate in the candidate list are reordered according to the template matching cost.

14. The method according to claim 1, wherein: The formula-based inter prediction information includes at least one of a template type and a formula type.

15. A method for video encoding, comprising: Determining to encode a current block in a current picture by inter prediction with potential use of a formula-based inter prediction technique; The formula-based inter-frame prediction technique generates prediction samples of the current block based on a formula by inputting one or more reconstructed samples of a reference block in a reference picture into the formula, wherein the formula includes one or more parameters derived based on a current template of the current block and a reference template of the reference block; constructing a candidate list, the candidate list including one or more coding blocks associated with the current block, wherein coding blocks of the one or more coding blocks are candidates for providing formula-based inter-frame prediction information of the formula-based inter-frame prediction technique for the inter-frame prediction of the current block; Determining current formula-based inter-frame prediction information of the current block based on the candidate list; as well as The current block is encoded based on the current formula-based inter-frame prediction information.

16. The method according to claim 15, wherein: The formula-based inter prediction technique is local illumination compensation (LIC), and the one or more coding blocks include at least one of a neighboring spatially neighboring coding block, a non-neighboring spatially neighboring block, and / or a temporally co-located coding block of the current block in the reference picture.

17. The method according to claim 15, wherein: Constructing the candidate list further includes: The candidate list is constructed according to a first-in-first-out (FIFO) queue, which stores historical formula-based inter-frame prediction information of one or more historical encoded blocks encoded using the formula-based inter-frame prediction technique before the current block.

18. The method according to claim 17, further comprising: The current formula-based inter prediction information applied to the current block is stored in the FIFO queue in a FIFO manner.

19. The method according to claim 17, further comprising at least one of the following: Resetting the FIFO queue at the beginning of a frame; resetting the FIFO queue at the beginning of a row of a coding tree unit (CTU); resetting the FIFO queue at the beginning of a row of a super block (SB); Resetting the FIFO queue at the beginning of a slice; resetting the FIFO queue at the start of a tile; and / or The FIFO queue is reset at the beginning of a segment.

20. A method of processing visual media data, the method comprising: Processing the code stream of visual media data according to format rules, where: The code stream includes encoding information of a picture sequence, the encoding information indicating that a formula-based inter-frame prediction technique is used to perform inter-frame prediction on a current block in a current picture; the formula-based inter-frame prediction technique generates a prediction sample of the current block by inputting one or more reconstructed samples of a reference block in a reference picture into the formula based on a formula, the formula including one or more parameters derived based on a current template of the current block and a reference template of the reference block; and The format rules specify: constructing a candidate list, the candidate list comprising one or more coding blocks associated with the current block, wherein coding blocks in the one or more coding blocks are candidates for providing formula-based inter-frame prediction information of the formula-based inter-frame prediction technique; Selecting a first coding block from the candidate list, the first coding block being encoded using first formula-based inter-frame prediction information of the formula-based inter-frame prediction technique; Determining current formula-based inter-frame prediction information of the current block based on the first formula-based inter-frame prediction information; determining, based on the current formula-based inter-frame prediction information, one or more parameter values ​​for applying the formula-based inter-frame prediction technique to the one or more parameters of the current block; and At least one prediction sample of the current block is generated based on the formula including the one or more parameters set according to the one or more parameter values.