Inter prediction adjustment method using linear prediction model

By adopting the formula-based inter prediction method in the video encoding and decoding technology, the parameters in the linear formula are adjusted to generate the adjusted linear formula, the problem of low inter prediction efficiency in the prior art is solved, and more efficient video encoding and decoding is achieved.

CN120077637APending Publication Date: 2025-05-30TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480004414.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2024-07-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are difficult to effectively adjust the linear prediction model in inter-frame prediction, resulting in low encoding efficiency.

Method used

The inter prediction method based on formula is adopted, and the slope parameters and offset parameters in the linear formula are adjusted, and the adjusted factor is used to generate the adjusted linear formula to improve the accuracy of the prediction sample.

Benefits of technology

The encoding efficiency of inter-prediction encoding is improved, and the quality and performance of video decoding are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077637A_ABST
    Figure CN120077637A_ABST
Patent Text Reader

Abstract

The method comprises the steps that a code stream is received, the code stream comprises coded information of a picture sequence, and the coded information indicates that inter-frame prediction is carried out on a current block in a current picture on the basis of a reference block in a reference picture, and an adjustment factor indicating a linear formula used when the inter prediction is a formula-based inter prediction, the formula-based inter prediction generating a prediction sample of the current block based on the linear formula, and one or more reconstructed samples of the reference block being input to the linear formula; the linear formula comprises one or more parameters, and the one or more parameters are derived based on a current template of the current block and a reference template of the reference block; applying the adjustment factor to the linear formula to generate an adjusted linear formula; and determining at least one reconstructed sample of the current block according to the adjusted linear formula.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims priority to U.S. Application No. 18 / 769,332, filed on July 10, 2024, with the title "Inter-Frame Prediction Adjustment Method Using a Linear Prediction Model", and to U.S. Provisional Application No. 63 / 526,166, filed on July 11, 2023, with the title "Inter-Frame Prediction Adjustment Method Using a Linear Prediction Model", the entire contents of which are incorporated herein by reference. Technical Field

[0003] Embodiments of this application relate to video encoding and decoding. Background Art

[0004] The background description provided herein is intended to present the background of the application as a whole. The extent to which the work of the presently named inventors, which is described in the background art section and in various aspects of this specification, was carried out does not indicate that it was prior art at the time of filing of this application, and it has never been expressly or implicitly admitted to be prior art of this application.

[0005] Image / video compression can help transfer image / video files between different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress an image based on spatial redundancy. For example, intra-frame prediction can use reference data from the current picture in reconstruction to perform sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress an image based on temporal redundancy. For example, inter-frame prediction can utilize motion compensation to predict samples in the current picture from a previously reconstructed picture. Motion compensation is typically indicated by a motion vector (MV). Summary of the Invention

[0006] Aspects of this application provide video encoding / decoding methods and apparatuses. In some examples, a video decoding apparatus includes a processing circuit.

[0007] Some aspects of this application provide a video decoding method, including:

[0008] receiving a bitstream, the bitstream including encoded information of a picture sequence, the encoded information indicating that for a current block in a current picture, inter-frame prediction is performed on the current block based on a reference block in a reference picture, and indicating an adjustment factor of a linear formula used when the inter-frame prediction is formula-based inter-frame prediction, wherein,

[0009] The formula-based inter-frame prediction generates prediction samples for the current block based on the linear formula, where

[0010] one or more reconstructed samples of the reference block are input into the linear formula;

[0011] the linear formula includes one or more parameters, and the one or more parameters are derived based on the current template of the current block and the reference template of the reference block;

[0012] applying the adjustment factor to the linear formula to generate an adjusted linear formula; and

[0013] determining at least one reconstructed sample of the current block according to the adjusted linear formula.

[0014] In some embodiments, applying the adjustment factor to the linear formula includes:

[0015] deriving the linear formula based on the current template of the current block and the reference template of the reference block, and the linear formula has an initial slope parameter;

[0016] applying the adjustment factor to the initial slope parameter to determine an adjusted slope parameter;

[0017] calculating an adjusted offset parameter according to the adjusted slope parameter, the reference template of the reference block, and the current template of the current block.

[0018] In some embodiments, calculating the adjusted offset parameter further includes:

[0019] calculating a first weighted average of a plurality of first samples in the reference template of the reference block;

[0020] calculating a second weighted average of a plurality of second samples in the current template of the current block;

[0021] scaling the first weighted average according to the adjusted slope parameter to generate a scaled first weighted average;

[0022] calculating the adjusted offset parameter based on the difference between the second weighted average and the scaled first weighted average.

[0023] In some embodiments, the plurality of first samples include each sample in the reference template, and the plurality of second samples include each sample in the current template.

[0024] In some embodiments, the plurality of first samples includes a first sample subset in the reference template, and the plurality of second samples includes a second sample subset in the current template, wherein, based on the motion information of the current block, the first sample subset corresponds to the second sample subset. In some embodiments, the first sample subset includes a plurality of subsampled samples in the reference template based on a subsampling pattern. In some embodiments, the first sample subset includes a plurality of corner samples and / or center samples in the reference template.

[0025] In some embodiments, applying the adjustment factor to the linear formula includes:

[0026] Deriving the linear formula based on the current template of the current block and the reference template of the reference block, the linear formula having an initial slope parameter and an initial offset parameter;

[0027] Applying the adjustment factor to the initial slope parameter to determine an adjusted slope parameter;

[0028] Calculating an adjusted offset parameter based on the adjusted slope parameter, the reference template of the reference block, and the current template of the current block;

[0029] Calculating a combined offset parameter according to a combination between the initial offset parameter and the adjusted offset parameter, wherein the adjusted offset parameter and the combined offset parameter are used to form the adjusted linear formula.

[0030] In some embodiments, the method further includes:

[0031] Decoding a first syntax element from the bitstream;

[0032] Determining a first adjustment factor according to the first syntax element, wherein the inter-frame prediction uses a plurality of color components and a plurality of linear formulas, and the first adjustment factor is applied to at least a first color component among the plurality of color components and / or at least a first linear formula among the plurality of linear formulas.

[0033] In some embodiments, the first syntax element indicates an entry in a lookup table, and the entry includes the value of the first adjustment factor.

[0034] In some embodiments, when the value of the first adjustment factor is applied to the first color component, the method further includes at least one of the following:

[0035] Determining that the first color component is predefined; and / or

[0036] Decode a high-level syntax element that indicates the first color component, where the high-level syntax element is at least one of a sequence header, a picture header, a slice header, and a frame header.

[0037] In some embodiments, the value of the first adjustment factor is applied to all color components.

[0038] In some embodiments, the first syntax element indicates an entry in a lookup table, where the entry includes at least the value of the first adjustment factor applied to the first color component and the value of a second adjustment factor applied to a second color component.

[0039] In some embodiments, the method further includes:

[0040] Decode at least a first syntax element and a second syntax element from the bitstream;

[0041] Determine a first adjustment factor for the first color component according to the first syntax element;

[0042] Determine a second adjustment factor for the second color component according to the second syntax element.

[0043] In some embodiments, the method further includes:

[0044] Decode one or more control flags that indicate whether to apply the adjustment factors to one or more color components respectively.

[0045] In some embodiments, the method further includes:

[0046] Decode at least a first control flag and a second control flag, where the first control flag indicates whether to apply the adjustment factor to the luminance component, and the second control flag indicates whether to apply the adjustment factor to the chrominance component.

[0047] In some embodiments, the method further includes:

[0048] Decode a first control flag that indicates whether to apply the adjustment factor to the luminance component;

[0049] When the first control flag is true, decode a second control flag that indicates whether to apply the adjustment factor to the chrominance component.

[0050] In some embodiments, the decoding of the first syntax element further includes:

[0051] When the current block is encoded with uni-directional prediction, decode the first syntax element from the bitstream;

[0052] When the current block is encoded with bi - directional prediction, set the value of the first adjustment factor to zero.

[0053] Some aspects of the present application provide a video encoding method, including:

[0054] Determine to encode a current block in a current picture according to inter - frame prediction, where the inter - frame prediction is formula - based inter - frame prediction, and the formula - based inter - frame prediction has an adjustment factor, where,

[0055] The formula - based inter - frame prediction generates prediction samples of the current block based on a linear formula, where,

[0056] Input one or more reconstructed samples of a reference block in a reference picture into the linear formula;

[0057] The linear formula includes one or more parameters, and the one or more parameters are derived based on the current template of the current block and the reference template of the reference block;

[0058] Determine the adjustment factor applied to the linear formula; and,

[0059] Encode the current block according to the adjustment factor to form a bitstream.

[0060] Some aspects of the present application also provide a visual media data processing method, including:

[0061] Process the bitstream of visual media data according to format rules, where:

[0062] The bitstream includes encoded information of a picture sequence, and the encoded information indicates that for a current block in a current picture, inter - frame prediction is performed for the current block based on a reference block in a reference picture, and indicates an adjustment factor of a linear formula used when the inter - frame prediction is formula - based inter - frame prediction, where,

[0063] The formula - based inter - frame prediction generates prediction samples of the current block based on the linear formula, where,

[0064] Input one or more reconstructed samples in the reference block into the linear formula;

[0065] The linear formula includes one or more parameters, and the one or more parameters are derived based on the current template of the current block and the reference template of the reference block;

[0066] The format rules specify:

[0067] Apply the adjustment factor to the linear formula to generate an adjusted linear formula; and,

[0068] Determine at least one reconstructed sample of the current block according to the adjusted linear formula.

[0069] Some aspects of the present application also provide a video encoding or decoding device. The video encoding or decoding device includes a processing circuit configured to implement any one of the above video encoding methods or video decoding methods.

[0070] Some aspects of the present application also provide a non - volatile computer - readable storage medium, on which instructions are stored. When the instructions are executed by a computer, the computer is caused to execute any one of the above video encoding methods or video decoding methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0072] Figure 1 is a schematic diagram of an exemplary block diagram of a communication system (100);

[0073] Figure 2 is a schematic diagram of an exemplary block diagram of a decoder;

[0074] Figure 3 is a schematic diagram of an exemplary block diagram of an encoder;

[0075] Figure 4 shows the positions of spatial merge candidates according to an embodiment of the present application;

[0076] Figure 5 shows candidate pairs considered for redundancy check of spatial merge candidates according to an embodiment of the present application;

[0077] Figure 6 shows an example motion vector scaling for temporal merge candidates;

[0078] Figure 7 shows an example candidate position of a temporal merge candidate of the current block;

[0079] Figure 8 shows a diagram of a template in some examples;

[0080] Figure 9 shows a flowchart of a method of a decoding process according to an embodiment of the present application;

[0081] Figure 10 shows a flowchart of a method of an encoding process according to an embodiment of the present application;

[0082] Figure 11 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION

[0083] Figure 1 FIG. shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example application of the subject matter disclosed in this application, a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0084] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, for creating an uncompressed video picture stream (102). In one example, the video picture stream (102) includes samples taken by the digital camera. Compared with the encoded video data (104) (or encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high-data-volume video picture stream. The video picture stream (102) can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or encoded video bitstream (104)) is depicted as a thin line to emphasize the lower-data-volume encoded video data (104) (or encoded video bitstream (104)), which can be stored on a streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1 the client subsystem (106) and the client subsystem (108) in, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes an incoming copy (107) of the encoded video data and produces an output video picture stream (411) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), (107), and (109) (such as a video bitstream) may be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In one example, a video coding standard under development is informally referred to as Next Generation Video Coding (Versatile Video Coding, VVC), and this application can be used in the context of the VVC standard.

[0085] Note that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may further include a video encoder (not shown).

[0086] Figure 2 An exemplary block diagram of a video decoder (210) is shown. The video decoder (210) may be disposed in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 the video decoder (110) in the example.

[0087] The receiver (231) may receive at least one encoded video sequence to be decoded by the video decoder (210); in the same embodiment or another embodiment, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be disposed outside the video decoder (210) (not labeled). In other cases, a buffer memory (not labeled) is disposed outside the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle the playback timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may have an adaptive size, and may be at least partially implemented in an operating system or a similar element (not labeled) outside the video decoder (210).

[0088] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), etc. The display device is not a component of the electronic device (230), but may be coupled to the electronic device (230), as Figure 2 shown. The control information for the display device may be a Supplemental Enhancement Information (SEI message) or a parameter set segment (not labeled) of Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence may be performed according to a video coding technology or standard, and may follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0089] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0090] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved may be controlled by subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and the multiple units below are not described.

[0091] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0092] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block including sample values, and the sample values can be input into the aggregator (255).

[0093] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0094] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from the address in the reference picture memory (257) can be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, and so on.

[0095] The output samples of the aggregator (255) can be adopted by various loop filtering techniques in the loop filter unit (256). Video compression techniques may include in-loop filter techniques, which are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used in the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, video compression techniques may also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.

[0096] The output of the loop filter unit (256) can be a sample stream, which can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.

[0097] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0098] The video decoder (210) can perform decoding operations according to, for example, a predetermined video compression technique in the ITU-T H.265 standard. In the sense that an encoded video sequence conforms to the syntax specified by the video compression technique or standard used, the encoded video sequence can comply with the syntax of the video compression technique or standard. Specifically, a profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0099] In one example, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by the video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.

[0100] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used in place of Figure 1 the video encoder (103) in the embodiments.

[0101] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the embodiments), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).

[0102] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.301 Y CrCB, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as a spatial pixel array, and each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.

[0103] According to an embodiment, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of the controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of simplicity, the couplings are not labeled in the figures. The parameters set by the controller (350) can include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) can be used for other suitable functions that relate to the video encoder (303) optimized for a certain system design.

[0104] In some embodiments, the video encoder (303) operates in an encoding loop. As a simple description, in one example, the encoding loop can include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on the input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data (since in the video compression techniques considered in this application, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream (sample data) is input into the reference picture memory (334). Since the decoding of the symbol stream produces a bit-exact result regardless of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples that the prediction part of the encoder "sees" are exactly the same as the sample values that the decoder will "see" when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs in cases where synchronization cannot be maintained, e.g., due to channel errors) is also used in some related technologies.

[0105] The operation of the "local" decoder (333) can be the same as that of the "remote" decoder, which has been described in detail above in connection with Figure 2 the video decoder (210). However, briefly referring additionally to Figure 2 , when the symbols are available and the entropy encoder (345) and the parser (220) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0106] In one embodiment, any decoder technology other than parsing / entropy decoding that exists in the decoder exists in the corresponding encoder in the same or substantially the same functional form. For this reason, this application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is the reverse of the decoder technology described in detail. A more detailed description is only needed in certain areas and is provided below.

[0107] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. With reference to at least one previously encoded picture designated as a "reference picture" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (332) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.

[0108] The local video decoder (333) may decode the encoded video data that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the coding engine (332) may be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 3 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture cache (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has the same content as the reconstructed reference picture to be obtained by the remote video decoder (without transmission errors).

[0109] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate on a sample block-by-block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory (334).

[0110] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0111] The outputs of all the above functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0112] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (660), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0113] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can typically be assigned to any of the following picture types:

[0114] An intra picture (I picture), which can be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of the variations of I pictures and their corresponding applications and characteristics.

[0115] A predictive picture (P picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most one motion vector and a reference index to predict the sample values of each block.

[0116] A bi - predictive picture (B picture), which can be a picture that can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0117] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block-wise. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding pictures applied to the blocks. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra-frame prediction). Pixel blocks of a P picture can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0118] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0119] In one example, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures, and slices, SEI messages, VUI parameter set segments, etc.

[0120] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra-picture prediction (often abbreviated as intra-frame prediction) exploits the spatial correlation within a given picture, while inter-picture prediction exploits the (temporal or other) correlation between pictures. In one example, a particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension that identifies the reference picture.

[0121] In some embodiments, bidirectional prediction techniques can be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be past and future respectively in display order). A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.

[0122] In addition, the merge mode technique can be used in inter-picture prediction to improve the encoding and decoding efficiency.

[0123] According to some embodiments disclosed in the present application, predictions such as inter-picture prediction and intra-picture prediction are performed on a per-block basis. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. The CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into at least one coding unit (CU) by a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In one example, each CU is analyzed to determine the prediction type for the CU, such as an inter-prediction type or an intra-prediction type. In addition, depending on the temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one example, the prediction operation in encoding (encoding / decoding) is performed on a per-prediction block basis. Taking the luminance prediction block as the prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and so on.

[0124] Note that any suitable technology can be used to implement the video encoders (103), (303) and video decoders (110), (210). In one example, at least one integrated circuit can be used to implement the video encoders (103), (303) and video decoders (110), (210). In another embodiment, at least one processor executing software instructions can be used to implement the video encoders (103), (303) and video decoders (110), (210).

[0125] This application provides techniques for an inter-frame prediction adjustment method using a linear prediction model in inter-frame prediction coding and decoding.

[0126] Various inter-frame prediction modes can be used in video coding. For example, in VVC, for an inter-frame prediction CU, the motion parameters can include one or more MVs, one or more reference picture indices, a reference picture list usage index, and additional information about certain coding features to be used for generating inter-frame prediction samples. The motion parameters can be signaled explicitly or implicitly. When a CU is encoded in skip mode, the CU can be associated with a PU and can not have valid residual coefficients, no encoded motion vector difference or MV difference (e.g., MVD) or reference picture index. A merge mode can be specified, where the motion parameters of the current CU are obtained from one or more neighboring CUs, including spatial and / or temporal candidates, and optionally additional information such as introduced in VVC. The merge mode can be applied to inter-frame prediction CUs, not just for skip mode. In an example, an alternative to the merge mode is the explicit transmission of motion parameters, where each CU explicitly signals one or more MVs, the corresponding reference picture indices for each reference picture list and a reference picture list usage flag, and other information.

[0127] In an embodiment, such as in VVC, the VVC Test Model (VTM) reference software includes one or more modified inter-frame prediction coding tools, which include: extended merge prediction, merged motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, sub-block based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bi-directional prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter-frame and intra-frame prediction (CIIP), geometric partitioning mode (GPM), etc. Inter-frame prediction and related methods are described in detail below.

[0128] In some examples, extended merge prediction can be used. In an example, such as in VTM4, a merge candidate list is constructed by sequentially including the following five types of candidates: one or more spatial motion vector predictors (MVPs) from one or more spatially adjacent CUs, one or more temporal MVPs from one or more co-located CUs, one or more history-based MVPs (HMVP) from a first-in-first-out (FIFO) table, one or more pairwise average MVPs, and one or more zero MVs.

[0129] The size of the merge candidate list can be signaled in the slice header. In an example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in merge mode, a truncated unary binarization (TU) can be used to encode the index of the best merge candidate (e.g., the merge index). The first binary bit of the merge index can be encoded with context (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other binary bits.

[0130] Some examples of the generation process of merge candidates for each category are provided below. In an embodiment, one or more spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In an example, up to four merge candidates are selected from the candidates at the positions depicted in Figure 4 the positions shown.

[0131] Figure 4 The positions of spatial merge candidates according to an embodiment of the present application are shown. Referring to Figure 4 , the derivation order is B1, A1, B0, A0, and B2. Only when any of the CUs at positions A0, B0, B1, and A1 are unavailable (e.g., because the CU belongs to another slice or another tile) or are intra-coded, is position B2 considered. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates to ensure that candidates with the same motion information are excluded from the candidate list, thereby improving the coding efficiency.

[0132] To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs connected by the arrows in Figure 5 are considered, and a candidate is added to the candidate list only if the corresponding candidates used for the redundancy check do not have the same motion information.

[0133] Figure 5 The candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present application are shown. Referring to Figure 5, the pairs connected by corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0134] In an embodiment, one or more temporal candidates are derived as follows. In the example, only one temporal merge candidate is added to the candidate list. Figure 6 An example motion vector scaling for temporal merge candidates is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) can be derived based on a collocated CU (612) belonging to a collocated reference picture (604) (e.g., shown by the Figure 6 dashed line in). The reference picture list for deriving the collocated CU (612) can be signaled explicitly in the slice header. As shown by the Figure 6 dashed line in, the scaled MV (621) for the temporal merge candidate can be obtained. The scaled MV (621) can be scaled from the MV of the collocated CU (612) using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the collocated reference picture (604) of the collocated picture (603) and the collocated picture (603). The reference picture index of the temporal merge candidate can be set to zero. The collocated picture is a reference picture that serves as a source picture for temporal motion information derivation. The collocated picture can be identified in one of two lists, called list0 or list1. In some examples, the encoder can use appropriate syntax techniques to determine the collocated picture and concurrently signal the collocated picture.

[0135] Figure 7 An example candidate position (e.g., C0 and C1) for the temporal merge candidate of the current CU is shown. The position for the temporal merge candidate can be selected from candidate positions C0 and C1. Candidate position C0 is at the lower right corner of the collocated CU (710) of the current CU. Candidate position C1 is at the center of the collocated CU (710) of the current CU. If the CU at candidate position C0 is unavailable, intra-coded, or outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, inter-coded, and within the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0136] According to some aspects of the present application, a formula-based prediction method can be used in inter-frame prediction or intra-frame prediction. For inter-frame prediction, the formula-based prediction method can use a formula to generate samples of a current block in a current picture based on reference samples of a reference block in a reference picture. For intra-frame prediction, the formula-based prediction method can use a formula to generate a first color component of a current block based on a second color component of the current block. In some examples, the parameters in the formula for the formula-based prediction method can be derived based on a template of the current block.

[0137] In some examples, local illumination compensation (LIC) is used as an inter-frame prediction technique to model local illumination changes between a current block and a predicted block (also referred to as a reference block) of the current block using a linear function. The predicted block is in a reference picture and can be pointed to by a motion vector (MV). The parameters of the linear formula can include a scaling factor α and an offset β, and the linear formula can be represented by α×p[x,y]+β to compensate for illumination changes, where p[x,y] represents a reference sample at position [x,y] in the reference block (also referred to as the predicted block), and the reference block is pointed to by the MV from the current block. In some examples, the scaling factor α and the offset β can be derived based on a template of the current block and a corresponding reference template of the reference block using the least squares method, so no signaling overhead is required, except that an LIC flag can be signaled to indicate the use of LIC. The scaling factor α and the offset β derived based on the template of the current block can be referred to as a template-based parameter set.

[0138] In some examples, LIC is used for uni-directional prediction between CUs. In some examples, intra-frame adjacent samples of the current block (adjacent samples predicted using intra-frame prediction) can be used for LIC parameter derivation. In some examples, LIC is disabled for blocks having less than 32 luma samples. In some examples, for non-sub-block modes (e.g., non-affine modes), LIC parameter derivation is performed based on the modulo-block samples of the current CU rather than the partial modulo-block samples of the top-left 16×16 unit. In some examples, LIC parameter derivation is performed based on partial modulo-block samples (such as partial modulo-block samples for the top-left 16×16 unit). In some examples, the template samples of the reference block are determined by using motion compensation (MC) with the MV of the block without rounding it to integer pixel precision.

[0139] In some examples, cross-component prediction can be used as an intra-frame prediction technique. Cross-component prediction can include a first technique called cross-component linear model (CCLM), a second technique called multi-model linear model (MMLM), a third technique called convolutional cross-component model (CCCM), and a fourth technique called gradient linear model (GLM).

[0140] For example, the first technique CCLM is used to reduce cross-component redundancy. In CCLM, by using a linear model (also referred to as a linear formula), such as using the following equation (1), chrominance samples are predicted based on the reconstructed luminance samples of the same CU:

[0141] pred C (i,j) = a·rec L ′(i,j) + b Equation (1)

[0142] where pred C (i,j) represents the predicted chrominance sample in the CU, and rec L ′(i,j) represents the downsampled reconstructed luminance sample of the same CU. The CCLM linear model includes parameters (a and b) that can be derived from at most four adjacent chrominance samples and their corresponding downsampled luminance samples in the example. In the example, at most four adjacent chrominance samples and their corresponding downsampled luminance samples are referred to as the template of the CU.

[0143] In some examples, based on the positions of adjacent chrominance samples, CCLM can include different modes called LM_T (LM top mode or upper mode LM_A), LM_L (LM left mode), and LM_LT (LM upper left corner mode or left upper mode LM_LA or just LM mode). For example, if the size of the current chrominance block is W×H, then W' and H' can be set for various modes in CCLM. When applying the LM mode (also referred to as LM_LT or LM_LA), W' = W, H' = H; when applying the LM-A mode, W' = W + H; when applying the LM-L mode, H' = H + W.

[0144] Note that MMLM, CCCM, and GLM also use functions for prediction. The parameters of the functions can be derived based on the template.

[0145] Note that the following description uses inter prediction to illustrate the techniques for deriving the encoded information for formula-based prediction methods, and these techniques can be appropriately used to derive the encoded information for intra prediction.

[0146] In some aspects of the present application, some inter - frame prediction techniques are designed to minimize the distortion between a current block in a corresponding reference picture and its predicted block. For example, an inter - frame prediction technique (also referred to as a first method of an inter - frame prediction method, a formula - based inter - frame prediction technique, a function - based inter - frame prediction technique, or a model - based inter - frame prediction technique) can apply formulas such as non - linear formulas, linear formulas, etc., using the original predicted block in the reference picture as the input of the formula to generate the current block in the current picture. For example, an inter - frame prediction technique can generate a prediction of samples in the current block based on a formula that takes one or more predicted samples in the reference picture as input. The formula can include linear terms or non - linear terms and can include one or more parameters that can be derived. Note that LIC is one of such inter - frame prediction techniques.

[0147] In some examples, the formula is a linear formula and can be represented by where n is a non - negative integer and p(x i ,y i ) is the predicted sample at position (x i ,y i ) in the reference picture, and this predicted sample is pointed to based on the MV associated with the current block. In addition, a set of predicted samples represented by p(x i ,y i ) (where i = 0, ……, n) can be a set of predicted samples around the corresponding sample in the reference samples pointed to by the MV based on the current sample to be predicted. In some examples, the parameters α i and β can be derived by minimizing the difference between the current block template and its predicted block template (e.g., by using the least - squares method) based on the template of the current block (also referred to as the current block template) and the template of the predicted block of the current block (also referred to as the predicted block template). The template of the current block consists of spatially adjacent reconstructed samples of the current block, and the template of the predicted block consists of spatially adjacent reconstructed samples of the predicted block.

[0148] Figure 8 FIG. shows templates in some examples. For example, the template (810) is called an L - shaped template T L , and includes adjacent samples at the upper row, left column, and upper - left corner of the current block (also referred to as the current coding block); the template (820) is called an upper and left template T a+l , and includes adjacent samples in the upper row and left column of the current block; the template (830) is called an upper template T a , and includes adjacent samples in the upper row of the current block; and the template (840) is called a left template T l and includes adjacent samples in the left column of the current block. It should be noted that the template can include adjacent samples of other suitable shapes not shown in Figure 8 .

[0149] In some examples, multiple candidate template types may also be supported, and one candidate template type is selected to derive the parameters of the linear formula. Syntax (also referred to as second syntax in some examples) may be signaled in the bitstream (e.g., at the block level) to indicate which candidate template type is selected.

[0150] In some examples, a control flag associated with the formula-based inter prediction technique may be signaled in the bitstream (e.g., at the block level) to indicate whether the formula-based inter prediction technique is applied to the current block. Alternatively, the value of the control flag may also be inherited from another coded block (e.g., set to the same). More specifically, a first control flag of the formula-based inter prediction technique associated with the current block is inherited from a second control flag of the formula-based inter prediction technique associated with another or more coded blocks. Additionally, the control flag may be derived at the coded block level to adaptively determine whether to apply the formula-based inter prediction technique.

[0151] In some examples, a first control flag of the formula-based inter prediction technique associated with the current block is inherited from a second control flag of the formula-based inter prediction technique associated with another or more coded blocks. In some examples, the coded information of the formula-based inter prediction technique may be derived from adjacent coded blocks, non-adjacent coded blocks, or coded blocks that store the coded information in a buffer.

[0152] Note that in this application, without loss of generality, in the examples, the term parameter refers to parameters ɑ i and β, which are used to determine the linear formula for deriving the prediction block. The term template type refers to different template shapes, such as but not limited to Figure 8 one of the template types for parameter derivation in the non-linear formula or the linear formula as shown in

[0153] Some aspects of the present application provide adjustment techniques for inter prediction using a linear prediction model (e.g., using a linear formula). The adjustment techniques may improve the coding efficiency of inter prediction coding.

[0154] In some examples, a linear formula (also referred to as a linear prediction model) is used in formula-based inter prediction techniques. The linear prediction model can be represented as α×p(x,y)+β, where p(x,y) is the reconstructed sample (also referred to as a reference sample) at position (x,y) in the reference picture, ɑ is referred to as the slope parameter, and β is referred to as the offset parameter. In some examples, the reconstructed sample corresponds to the current sample in the current block for prediction and is pointed to by the MV based on the current sample. The parameters α and β are derived based on the current template of the current block and the reference template of the reference block (also referred to as the prediction block) in the reference picture, for example, by using the least squares method. In some examples, in the case of bi-directional prediction, two linear prediction models can be derived and applied to the reference blocks (also referred to as the prediction blocks) of the two reference lists respectively.

[0155] According to some aspects of the present application, in order to improve the accuracy of the linear prediction model, a technique called an adjustment technique is used to allow the linear prediction model to have an adjustment factor. For example, the encoder / decoder can apply the adjustment factor to the linear formula to generate an adjusted linear formula and perform encoding / decoding according to the adjusted linear formula. In some examples, the adjustment factor can be determined according to a look-up table and syntax (also referred to as the third syntax). The look-up table includes multiple potential values of the adjustment factor that are indexed. The syntax can indicate the index of the entry in the look-up table and can signal the syntax to indicate which value of the adjustment factor is applied to the linear formula.

[0156] According to one aspect of the present application, a single adjustment factor μ is used to adjust the derived parameters of the linear prediction model, and the syntax (also referred to as the third syntax) is signaled to indicate the index of the entry in the look-up table. In some examples, the look-up table stores the entries of the potential values of the adjustment factor, and the potential value in the entry pointed to by the index is applied as the single adjustment factor μ to the linear prediction model.

[0157] In some embodiments, the initial linear prediction model includes a slope parameter but does not include an offset parameter. The initial value of the slope parameter can be determined based on the current template of the current block and the reference template of the reference block (also referred to as the prediction block) in the reference picture, for example, by using the least squares method. In addition, the adjustment factor μ is combined with the initial linear prediction model (e.g., with the initial value of the slope parameter) to generate an adjusted linear prediction model, which includes an adjusted slope parameter and an adjusted offset parameter. For example, when the adjustment factor μ is included in the linear prediction model and the adjusted linear prediction model is modified, as shown in Equation (2):

[0158] (ɑ+μ)×p(x,y)+(avgT cur -(α+μ)×avgT ref ) Equation (2)

[0159] where avgT cur represents the weighted average of the first one or more samples in the current template of the current block, and avgT ref represents the weighted average of the second one or more samples in the reference template. According to the motion information of the current block for inter-frame prediction, the second one or more samples correspond to the first one or more samples.

[0160] In some examples, avgT cur and avgT ref are respectively the weighted averages of all samples within the templates of the current block and the predicted block. For example, avgT cur represents the weighted average of all first samples in the current template of the current block, and avgT ref represents the weighted average of all second samples in the reference template. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0161] In some examples, avgT cur and avgT ref are respectively the weighted averages of some specific samples within the templates of the current block and the predicted block. In the example, the weighted average is calculated by weighted averaging the samples at the four corners of the template and one or more samples at the center position.

[0162] In some examples, avgT cur represents the weighted average of the first samples in the current template of the current block according to a subsampling pattern (such as a checkerboard pattern, etc.), and avgT ref represents the weighted average of the second samples in the reference template according to the subsampling pattern. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0163] In some examples, avgT cur represents the weighted average of the first samples at specific positions (e.g., the four corners and / or the center position) in the current template of the current block, and avgT ref represents the weighted average of the second samples at specific positions in the reference template. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0164] In some embodiments, the initial linear prediction model includes an (initial) slope parameter and an (initial) offset parameter. The initial values of the slope parameter and the offset parameter can be determined based on the current template of the current block and the reference template of the reference block (also referred to as the prediction block) in the reference picture, for example, by using the least squares method. In addition, an adjustment factor μ is combined with the initial linear prediction model (e.g., the initial slope parameter) to generate an adjusted linear prediction model, which includes an adjusted slope parameter (e.g., α + μ) and a combined offset parameter (the initial offset parameter β and the combined adjusted offset parameter (avgT cur -(α + μ) × avgT ref ). In some examples, the adjustment factor μ is included in the linear prediction model, and the adjusted linear prediction model is modified as shown by the adjusted linear formula in Equation (3):

[0165] (α + μ) × p(x, y) + (β - (avgT cur -(α + μ) × avgT ref )) Equation (3)

[0166] where avgT cur represents the weighted average of the first one or more samples in the current template of the current block, and avgT ref represents the weighted average of the second one or more samples in the reference template. According to the motion information of the current block for inter-frame prediction, the second one or more samples correspond to the first one or more samples.

[0167] In some examples, avgT cur and avgT ref are respectively the weighted averages of all the samples within the templates of the current block and the prediction block. For example, avgT cur represents the weighted average of all the first samples in the current template of the current block, and avgT ref represents the weighted average of all the second samples in the reference template. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0168] In some examples, avgT cur and avgT ref are respectively the weighted averages of some specific samples within the templates of the current block and the prediction block. In an example, the weighted average is calculated by weighted averaging the samples at the four corners of the template and one or more samples at the center position.

[0169] In some examples, avgT currepresents the weighted average of the first samples in the current template of the current block according to a subsampling pattern (such as a checkerboard pattern, etc.), and avgT ref represents the weighted average of the second samples in the reference template according to the subsampling pattern. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0170] In some examples, avgT cur represents the weighted average of the first samples at specific positions (e.g., the four corners and / or the center position) in the current template of the current block, and avgT ref represents the weighted average of the second samples at specific positions in the reference template. According to the motion information of the current block for inter-frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0171] In some embodiments, the initial linear prediction model includes a slope parameter and an offset parameter. The initial values of the slope parameter and the offset parameter can be determined based on the current template of the current block and the reference template of the reference block (also referred to as the prediction block) in the reference picture, for example, by using the least squares method. In addition, an adjustment factor μ is combined with the initial linear prediction model (e.g., the initial slope parameter) to generate an adjusted linear prediction model, which includes an adjusted slope parameter (e.g., α + μ) and the initial offset parameter (e.g., β) and the combined offset parameter of the adjusted offset parameter (e.g., (β + (avgT cur -(α + μ)×avgT ref ))). In some examples, the adjustment factor μ is included in the linear prediction model, and the adjusted linear prediction model is modified as shown by the adjusted linear formula in Equation (4):

[0172] (α + μ)×p(x,y)+(β + (avgT cur -(α + μ)×avgT ref )) Equation (4)

[0173] where avgT cur represents the weighted average of the first one or more samples in the current template of the current block, and avgT ref represents the weighted average of the second one or more samples in the reference template. According to the motion information for inter-frame prediction, the second one or more samples correspond to the first one or more samples.

[0174] In some examples, avgT cur and avgT ref are respectively the weighted averages of all samples within the templates of the current block and the prediction block. For example, avgTcur represents the weighted average of all first samples in the current template of the current block, and avgT ref represents the weighted average of all second samples in the reference template. According to the motion information of the current block for inter - frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0175] In some examples, avgT cur and avgT ref are respectively the weighted averages of some specific samples within the templates of the current block and the predicted block. In the example, the weighted average is calculated by weighted - averaging the samples at the four corners of the template and one or more samples at the center position.

[0176] In some examples, avgT cur represents the weighted average of the first samples in the current template of the current block according to a subsampling pattern (such as a checkerboard pattern, etc.), and avgT ref represents the weighted average of the second samples in the reference template according to the subsampling pattern. According to the motion information of the current block for inter - frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0177] In some examples, avgT cur represents the weighted average of the first samples at specific positions (e.g., the four corners and / or the center position) in the current template of the current block, and avgT ref represents the weighted average of the second samples at specific positions in the reference template. According to the motion information of the current block for inter - frame prediction, the second samples in the reference template correspond to the first samples in the current template.

[0178] According to some aspects of the present application, an adjustment factor can be signaled to be applied to at least one of all color components and / or applied to at least one of all linear prediction models. In some examples, the color components include a luminance component and two chrominance components. In some examples, multiple linear prediction models can be applied to the current block, for example, applied to different samples at different positions in the current block.

[0179] In some embodiments, a single index is signaled and used to apply an associated adjustment factor to a linear prediction model in at least one of all color components. For example, the single index is signaled and points to an entry in a lookup table that includes the value of the adjustment factor. In an embodiment, the signaled adjustment factor is applied to only one specific color component. The specific color component is a predefined color component. In some examples, the specific color component may be signaled in a high-level syntax (such as a sequence header, a frame header, a picture header, a slice header, etc.). In an example, the signaled adjustment factor is applied to only the luminance component.

[0180] In another embodiment, the signaled adjustment factor is applied to all color components.

[0181] In some embodiments, a single index is signaled and used to apply an associated adjustment factor to a linear prediction model in at least one of all color components. In some examples, the single index points to an entry in a lookup table that includes the value of the adjustment factor. In an example, the value of the adjustment factor may be applied to only the luminance component. In another example, the value of the adjustment factor may be applied to the luminance component and the chrominance component.

[0182] In some examples, the single index points to an entry in the lookup table that includes a combination of a first value and a second value of the adjustment factor. In an example, the first value of the adjustment factor may be applied to the luminance component, and the second value of the adjustment factor may be applied to the chrominance component.

[0183] In another embodiment, multiple indices are signaled and used to indicate the associated adjustment factors for different color components. In an example, a first index and a second index are signaled. The first index points to the first entry in the lookup table, and the second index points to the second entry in the lookup table. The first entry includes the first value of the adjustment factor applied to the luminance component. The second entry includes the second value of the adjustment factor applied to the chrominance component.

[0184] In some embodiments, control flags regarding the application of the adjustment factor to different color components may be signaled separately. For example, one or more control flags may be used to signal separately whether to apply the adjustment to the luminance component or the chrominance component.

[0185] In an embodiment, a first control flag is signaled to indicate whether to apply the adjustment factor for the luminance component, and a second control flag is signaled to indicate whether to apply the adjustment factor for one or more chrominance components.

[0186] In another embodiment, a first control flag is signaled to indicate whether an adjustment to the luminance component is applied. When the first control flag indicates that the adjustment to the luminance component is not applied, the adjustment to the chrominance component is also not applied. Otherwise, when the first flag indicates that the adjustment to the luminance component is applied, a second control flag is signaled to indicate whether one or more chrominance component adjustment factors are applied. In an example, when one or more adjustment factors are applied to the luminance component and the chrominance component, the luminance component and one or more chrominance components share the same index or share the same value of the adjustment factor. In another example, when one or more adjustment factors are applied to the luminance component and the chrominance component, the values of the adjustment factors applied to the luminance component and one or more chrominance components can be determined separately.

[0187] According to one aspect of the present application, the adjustment factor can be applied to unidirectional prediction but not to bidirectional prediction. In some embodiments, when the current block is encoded with unidirectional prediction, at least one index is signaled to indicate the associated adjustment factor. When the current block is encoded with bidirectional prediction, no index is signaled, and the adjustment factor of one or more linear prediction models is inferred to be 0.

[0188] According to another aspect of the present application, the adjustment factor can be applied to both unidirectional prediction and bidirectional prediction. In some embodiments, when the current block is encoded with bidirectional prediction, a single index is signaled and used to apply the associated adjustment factor to the linear prediction models of two prediction blocks. In an example, a single index is signaled, and the single index points to an entry in a lookup table storing the value of the adjustment factor. In an example, when the current block is encoded with bidirectional prediction, the same value is applied as the adjustment factor to the linear formulas of the first reference picture (e.g., from L0) and the second reference picture (e.g., from L1). In another example, when the current block is encoded with bidirectional prediction, the value is applied as the adjustment factor to the first linear formula of the first reference block in the first reference picture (e.g., from L0), and the negative value of the value is applied as the adjustment factor to the second linear formula of the second reference block in the second reference picture (e.g., from L1).

[0189] In some embodiments, when the current block is encoded with bidirectional prediction, more than one index is used to apply multiple associated adjustment factors to the linear prediction models of multiple prediction blocks. In some examples, when the current block is encoded with bidirectional prediction, a first index is signaled and used to apply a first value of the adjustment factor to the first linear formula (also referred to as the first linear prediction model) of the first reference block in the first reference picture, and a second index is signaled and used to apply a second value of the adjustment factor to the second linear formula (also referred to as the second linear prediction model) of the second reference block in the second reference picture.

[0190] In some embodiments, a first flag is signaled to indicate whether an adjustment factor is zero (zero means the adjustment factor is not applied and the derived model parameters are used). When the first flag is signaled and indicates that the adjustment factor is not zero, then the magnitude of the adjustment factor and a sign value indicating whether the adjustment factor is positive or negative are further signaled.

[0191] Figure 9 A flowchart illustrating a process (900) according to one aspect of the present application is shown. The process (900) can be used in a video decoder. In various aspects, the process (900) is executed by processing circuitry, such as processing circuitry that performs the functions of video decoder (110), processing circuitry that performs the functions of video decoder (210), etc. In some aspects, the process (900) is implemented as software instructions, so that when the processing circuitry executes the software instructions, the processing circuitry executes the process (900). The process begins at (S901) and proceeds to (S910).

[0192] At (S910), a bitstream is received, the bitstream including encoded information of a picture sequence, the encoded information indicating that for a current block in a current picture, inter - frame prediction is performed for the current block based on a reference block in a reference picture, and an adjustment factor of a linear formula used when the inter - frame prediction is formula - based inter - frame prediction, wherein, for the formula - based inter - frame prediction, prediction samples of the current block are generated based on the linear formula, and one or more reconstructed samples of the reference block are input into the linear formula; the linear formula includes one or more parameters, and the one or more parameters are derived based on a current template of the current block and a reference template of the reference block.

[0193] At (S920), the adjustment factor is applied to the linear formula to generate an adjusted linear formula.

[0194] At (S930), at least one reconstructed sample of the current block is determined according to the adjusted linear formula.

[0195] In some embodiments, to apply the adjustment factor, a linear formula with a slope parameter is initially derived based on a current template of the current block and a reference template of the reference block. In some examples, the initial offset parameter is set to zero. The adjustment factor is applied to the slope parameter to determine an adjusted slope parameter (e.g., (ɑ + μ) in Equation (2)). Further, according to the adjusted slope parameter, the reference template of the reference block, and the current template of the current block, an adjusted offset parameter (e.g., (avgT cur -(α + μ)×avgT ref )) is calculated, as shown in Equation (2) for example.

[0196] In some embodiments, to calculate the adjusted offset parameter, a first weighted average of a plurality of first samples in a reference template of a reference block is calculated, and a second weighted average of a plurality of second samples in a current template of a current block is calculated. Additionally, based on the adjusted slope parameter, the first weighted average is scaled to generate a scaled first weighted average, and based on the difference between the second weighted average and the scaled first weighted average, the adjusted offset parameter is calculated, as shown in Equation (2).

[0197] In some examples, the plurality of first samples includes all samples in the reference template, and the plurality of second samples includes all samples in the current template.

[0198] In some examples, the plurality of first samples includes a first subset of samples in the reference template, and the plurality of second samples includes a second subset of samples in the current template, and based on the motion information (e.g., motion vector) of the current block, the first subset of samples corresponds to the second subset of samples.

[0199] In some examples, the first subset of samples includes a plurality of subsampled samples in the reference template based on a subsampling pattern.

[0200] In some examples, the first subset of samples includes samples at specific locations, such as a plurality of corner samples and / or center samples in the reference template.

[0201] In some embodiments, to apply an adjustment factor, a linear formula with an initial slope parameter (e.g., α) and an initial offset parameter (e.g., β) is derived based on the current template of the current block and the reference template of the reference block. The adjustment factor is applied to the initial slope parameter to determine the adjusted slope parameter (e.g., α + μ). Based on the adjusted slope parameter, the reference template of the reference block, and the current template of the current block, the adjusted offset parameter (e.g., (avgT cur -(α + μ)×avgT ref )) is calculated, and based on the combination between the initial offset parameter and the adjusted offset parameter, the combined offset parameter is calculated, as shown by Equation (3) and Equation (4). The adjusted offset parameter and the combined offset parameter are used to form an adjusted linear formula.

[0202] In some examples, to calculate the adjusted offset parameter, a first weighted average of a plurality of first samples in a reference template of a reference block is calculated, and a second weighted average of a plurality of second samples in a current template of a current block is calculated. The first weighted average is scaled according to the adjusted slope parameter to generate a scaled first weighted average. The adjusted offset parameter is calculated as the difference between the second weighted average and the scaled first weighted average.

[0203] In some examples, the multiple first samples include all samples in a reference template, and the multiple second samples include all samples in a current template.

[0204] In some examples, the multiple first samples include a first sample subset in a reference template, and the multiple second samples include a second sample subset in a current template, and based on the motion information (e.g., motion vectors) of a current block, the first sample subset corresponds to the second sample subset.

[0205] In some examples, the first sample subset includes multiple subsampled samples in a reference template based on a subsampling pattern.

[0206] In some examples, the first sample subset includes samples at specific positions, such as multiple corner samples and / or center samples in a reference template.

[0207] According to one aspect of the present application, a first syntax element is decoded from a bitstream, and based on the first syntax element, a first adjustment factor is determined. Inter-frame prediction uses multiple color components and multiple linear formulas, and the first adjustment factor is applied to at least a first color component among the multiple color components, and / or at least a first linear formula among the multiple linear formulas.

[0208] In some embodiments, the first syntax element indicates an entry in a lookup table, and the entry includes the value of the first adjustment factor.

[0209] In some examples, the value of the first adjustment factor is only applied to the first color component, such as being applied to the linear formula for inter-frame prediction of the first color component. In an example, the first color component is predefined. In another example, the first color component is signaled by a high-level syntax element, such as a sequence header, a frame header, a picture header, a slice header, etc.

[0210] In some examples, the value of the first adjustment factor is applied to all color components, such as in multiple linear formulas respectively used in inter-frame prediction of multiple color components.

[0211] In some examples, the first syntax element indicates an entry in a lookup table, and the entry at least includes the value of the first adjustment factor applied to the first color component and the value of a second adjustment factor applied to a second color component.

[0212] In some examples, at least a first syntax element and a second syntax element are decoded from a bitstream. Based on the first syntax element, a first adjustment factor for a first color component is determined, and based on the second syntax element, a second adjustment factor for a second color component is determined.

[0213] In some embodiments, one or more control flags are used to indicate whether an adjustment factor is applied to corresponding color components. In some examples, one or more control flags are decoded, and the one or more control flags indicate whether the adjustment factor is applied to one or more color components respectively. In an example, at least a first control flag and a second control flag are decoded, where the first control flag indicates whether the adjustment factor is applied to a luminance component, and the second control flag indicates whether the adjustment factor is applied to a chrominance component. In another example, a first control flag is decoded, and the first control flag indicates whether the adjustment factor is applied to a luminance component; when the first control flag is true, a second control flag is decoded, and the second control flag indicates whether the adjustment factor is applied to a chrominance component. In an example, when the adjustment factor is applied to both the luminance component and the chrominance component, the same value of the adjustment factor is applied to the luminance component and the chrominance component.

[0214] In some embodiments, when the current block is encoded with uni-directional prediction, the first syntax element is decoded from the bitstream; when the current block is encoded with bi-directional prediction, the value of the first adjustment factor is set to zero.

[0215] In some embodiments, when the current block is encoded with bi-directional prediction, the value of a first adjustment factor and the value of a second adjustment factor are determined according to a first syntax element. The value of the first adjustment factor is applied to a first linear formula associated with a first reference picture, and the value of the second adjustment factor is applied to a second linear formula associated with a second reference picture. In an example, the value of the first adjustment factor and the value of the second adjustment factor are the same. In another example, the value of the second adjustment factor is the negative value of the value of the first adjustment factor.

[0216] In some embodiments, when the current block is encoded with bi-directional prediction, a first syntax element and a second syntax element are decoded from the bitstream. The value of the first adjustment factor is determined according to the first syntax element, and the value of the second adjustment factor is determined according to the second syntax element. The value of the first adjustment factor is applied to a first linear formula associated with a first reference picture, and the value of the second adjustment factor is applied to a second linear formula associated with a second reference picture.

[0217] In some embodiments, a flag indicating whether the adjustment factor is zero is decoded. When the flag indicates that the adjustment factor is non-zero, the magnitude and the sign of the adjustment factor are decoded according to one or more syntax elements in the bitstream.

[0218] Then, the process proceeds to (S999) and terminates.

[0219] The process (900) can be appropriately modified. One or more steps in the process (900) can be modified and / or omitted. One or more additional steps can be added. Any suitable order of implementation can be used.

[0220] Figure 10 A flowchart showing an overview of a process (1000) according to one aspect of the present application is presented. The process (1000) can be used in a video encoder. In various aspects, the process (1000) is executed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (1000) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (1000). The process starts at (S1001) and proceeds to (S1010).

[0221] At (S1010), it is determined to encode a current block in a current picture according to inter-frame prediction, where the inter-frame prediction is formula-based inter-frame prediction, the formula-based inter-frame prediction has an adjustment factor, and wherein the formula-based inter-frame prediction generates prediction samples of the current block based on a linear formula, and one or more reconstructed samples of a reference block in a reference picture are input into the linear formula; the linear formula includes one or more parameters, and the one or more parameters are derived based on the current template of the current block and the reference template of the reference block.

[0222] At (S1020), the adjustment factor applied to the linear formula is determined.

[0223] At (S1030), the current block is encoded according to the adjustment factor to form a bitstream.

[0224] In some embodiments, an initial linear formula with an initial slope parameter is derived based on the current template of the current block and the reference template of the reference block. In some examples, the initial offset parameter is set to zero. Then, possible values of the adjustment factor can be applied to the initial linear formula respectively to generate potential adjusted linear formulas, and the potential adjusted linear formulas can be evaluated (e.g., based on a cost function, based on rate distortion, etc.) to determine the final value of the adjustment factor.

[0225] In an example, for a potential value of the adjustment factor, the adjustment factor is applied to the initial slope parameter to determine the adjusted slope parameter. In addition, an adjusted offset parameter is calculated according to the adjusted slope parameter, the reference template of the reference block, and the current template of the current block, as shown in equation (2) for example.

[0226] In some embodiments, to calculate the adjusted offset parameter, a first weighted average of a plurality of first samples in a reference template of a reference block is calculated, and a second weighted average of a plurality of second samples in a current template of a current block is calculated. Additionally, based on the adjusted slope parameter, the first weighted average is scaled to generate a scaled first weighted average, and based on the difference between the second weighted average and the scaled first weighted average, the adjusted offset parameter is calculated, as shown in Equation (2).

[0227] In some examples, the plurality of first samples includes all samples in the reference template, and the plurality of second samples includes all samples in the current template.

[0228] In some examples, the plurality of first samples includes a first subset of samples in the reference template, and the plurality of second samples includes a second subset of samples in the current template, and based on the motion information (e.g., motion vector) of the current block, the first subset of samples corresponds to the second subset of samples.

[0229] In some examples, the first subset of samples includes a plurality of subsampled samples in the reference template based on a subsampling pattern.

[0230] In some examples, the first subset of samples includes samples at specific locations, such as a plurality of corner samples and / or center samples in the reference template.

[0231] In some embodiments, based on the current template of the current block and the reference template of the reference block, an initial linear formula with an initial slope parameter and an initial offset parameter is derived. Then, possible values of an adjustment factor can be applied to the initial linear formula respectively to generate potential adjusted linear formulas, and the potential adjusted linear formulas can be evaluated (e.g., based on a cost function, based on rate distortion, etc.) to determine the final value of the adjustment factor.

[0232] In an example, for a potential value of the adjustment factor, the adjustment factor is applied to the initial slope parameter to determine the adjusted slope parameter. Based on the adjusted slope parameter, the reference template of the reference block, and the current template of the current block, the adjusted offset parameter is calculated. Based on the combination between the initial offset parameter and the adjusted offset parameter, a combined offset parameter is calculated, as shown in Equation (3) and Equation (4). The adjusted offset parameter and the combined offset parameter are used to form the adjusted linear formula.

[0233] In some examples, to calculate the adjusted offset parameter, a first weighted average of a plurality of first samples in a reference template of a reference block is calculated, and a second weighted average of a plurality of second samples in a current template of a current block is calculated. The first weighted average is scaled according to the adjusted slope parameter to generate a scaled first weighted average. The adjusted offset parameter is calculated as the difference between the second weighted average and the scaled first weighted average.

[0234] In some examples, the plurality of first samples includes all samples in the reference template, and the plurality of second samples includes all samples in the current template.

[0235] In some examples, the plurality of first samples includes a first subset of samples in the reference template, and the plurality of second samples includes a second subset of samples in the current template, and based on the motion information (e.g., motion vector) of the current block, the first subset of samples corresponds to the second subset of samples.

[0236] In some examples, the first subset of samples includes a plurality of subsampled samples in the reference template based on a subsampling pattern.

[0237] In some examples, the first subset of samples includes samples at specific positions, such as a plurality of corner samples and / or center samples in the reference template.

[0238] According to one aspect of the present application, one or more syntax elements are encoded into a bitstream to indicate the final value of the adjustment factor. In some examples, the look-up table includes entries corresponding to potential values of the adjustment factor, and the index of the entry storing the final value of the adjustment factor can be encoded by one or more syntax elements in the bitstream.

[0239] In some examples, all color components share the same final value of the adjustment factor, and this final value can be signaled by a syntax element. In some examples, all linear models used in inter prediction share the same final value of the adjustment factor, and this final value can be signaled by a syntax element. In some examples, the adjustment factor is only applied to a specific color component. In an example, the specific color component can be predefined. In another example, the specific color component is signaled by high-level syntax.

[0240] In some examples, each entry in the look-up table includes a combination of values for different color components. Then, the index can be encoded to indicate the entry with the final value of the adjustment factor for different color components.

[0241] In some examples, at least a first syntax element and a second syntax element are encoded into the bitstream. The first syntax element indicates a first final value of the adjustment factor for a first color component, and the second syntax element indicates a second final value of the adjustment factor for a second color component.

[0242] In some embodiments, the encoder may signal one or more control flags in the bitstream to indicate whether an adjustment factor is applied to a corresponding color component. In an example, the encoder may signal a first control flag and a second control flag, where the first control flag indicates whether the adjustment factor is applied to the luminance component, and the second control flag indicates whether the adjustment factor is applied to the chrominance component.

[0243] In another example, the encoder may signal a first control flag indicating whether an adjustment factor is applied to the luminance component. When the first control flag is true, the encoder may signal a second control flag indicating whether the adjustment factor is applied to the chrominance component. When the first control flag is false, no further signaling is required for the second control flag.

[0244] In some embodiments, the adjustment factor is not used for bidirectional prediction. When the current block is encoded using unidirectional prediction, the encoder may encode syntax elements to indicate the value of the adjustment factor. When the current block is encoded using bidirectional prediction, no syntax elements for the adjustment factor need to be encoded.

[0245] In some embodiments, the adjustment factor may be used for bidirectional prediction. In some examples, the encoder encodes an index indicating the final value of the adjustment factor. In an example, the same value of the adjustment factor may be applied to a first linear formula associated with a first reference picture and a second linear formula associated with a second reference picture. In another example, opposite values of the adjustment factor may be applied to the first linear formula associated with the first reference picture and the second linear formula associated with the second reference picture, respectively.

[0246] In some embodiments, when the current block is encoded using bidirectional prediction, the encoder encodes a first syntax element and a second syntax element into the bitstream. The first syntax element indicates a first value of a first adjustment factor applied to a first linear formula associated with a first reference picture, and the second syntax element indicates a second value of a second adjustment factor applied to a second linear formula associated with a second reference picture.

[0247] In some embodiments, the encoder encodes a flag indicating whether the adjustment factor is zero. When the flag indicates that the adjustment factor is non-zero, the encoder encodes one or more syntax elements in the bitstream to indicate the magnitude and the sign of the adjustment factor.

[0248] Then, the process proceeds to (S1099) and terminates.

[0249] The process (1000) can be modified appropriately. One or more steps in the process (1000) can be modified and / or omitted. One or more additional steps can be added. Any suitable order of implementation can be used.

[0250] According to one aspect of the present application, a method for processing visual media data is provided. In this method, a bitstream of visual media data is processed according to formatting rules. For example, the bitstream can be a bitstream decoded / encoded by any decoding and / or encoding method described herein. The formatting rules can specify one or more constraints of the bitstream and / or one or more processes to be performed by a decoder and / or an encoder.

[0251] In an example, the bitstream includes encoded information of a picture sequence, the encoded information indicating inter-frame prediction of a current block in a current picture based on a reference block in a reference picture for the current block, and an adjustment factor of a linear formula used in formula-based inter-frame prediction. The formula-based inter-frame prediction generates prediction samples of the current block based on a linear formula, where one or more reconstructed samples in the reference block are input into the linear formula, and the linear formula includes one or more parameters derived based on a current template of the current block and a reference template of the reference block. The formatting rules stipulate that the adjustment factor is applied to the linear formula to generate an adjusted linear formula, and at least one reconstructed sample of the current block is determined according to the adjusted linear formula.

[0252] The above techniques can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable storage medium. For example, Figure 11 A computer system (1100) is shown, which is adapted to implement certain embodiments of the disclosed subject matter.

[0253] The computer software can be encoded by any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through methods such as decoding and microcode.

[0254] The instructions can be executed on various types of computers or their components, including for example personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0255] Figure 11 The components shown for the computer system (1100) are exemplary in nature and are not used to impose any limitation on the scope of use or functions of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one component or combination thereof shown in the exemplary embodiments of the computer system (1100).

[0256] A computer system (1100) may include certain human - machine interface input devices. Such human - machine interface input devices can respond to the input of at least one human user through tactile inputs (such as keyboard input, swiping, data glove movement), audio inputs (such as voice, applause), visual inputs (such as gestures), and olfactory inputs (not shown). The human - machine interface device can also be used to capture certain media, which need not be directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still - image camera), and video (e.g., two - dimensional video, three - dimensional video including stereoscopic video).

[0257] The human - machine interface input device may include at least one of the following (only one is shown): keyboard (1101), mouse (1102), touchpad (1103), touchscreen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).

[0258] The computer system (1100) may also include certain human - machine interface output devices. Such human - machine interface output devices can stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human - machine interface output devices may include tactile output devices (such as tactile feedback through the touchscreen (1110), data glove (not shown), or joystick (1105), but there can also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1109), headphones (not shown)), visual output devices (such as screens (1110) including cathode - ray tube screens, liquid - crystal screens, plasma screens, organic - light - emitting diode screens, each of which may or may not have touchscreen input functionality and may or may not have tactile feedback functionality - some of which can output two - dimensional visual output or output above three - dimensions through means such as stereoscopic picture output; virtual - reality glasses (not shown), holographic displays, and smoke - emitting boxes (not shown)), and printers (not shown).

[0259] The computer system (1100) may also include human - accessible storage devices and their associated media, such as optical media including high - density read - only / rewritable compact discs (CD / DVD ROM / RW) (1120) with CD / DVD or similar media (1121), thumb drives (1122), removable hard - disk drives or solid - state drives (1123), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0260] Those skilled in the art should also understand that the term "computer-readable storage medium" used in connection with the disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0261] The computer system (1100) may also include an interface (1154) to at least one communication network (1155). For example, the network can be wireless, wired, optical. The network can also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, etc. The network also includes local area networks such as Ethernet, wireless local area network, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television wired or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), etc. Some networks typically require an external network interface adapter for connection to certain common data ports or peripheral buses (1149) (e.g., the USB port of the computer system (1100)); other systems are typically integrated into the core of the computer system (1100) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (1100) can communicate with other entities. The communication can be one-way, only for receiving (e.g., wireless television), one-way only for sending (e.g., CAN bus to certain CAN bus devices), or two-way, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0262] The above-mentioned human-machine interface device, human-accessible storage device, and network interface can be connected to the core (1140) of the computer system (1100).

[0263] The core (1140) may include at least one central processing unit (CPU) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (1143), a hardware accelerator for specific tasks (1144), a graphics adapter (1150), etc. These devices, as well as a read-only memory (ROM) (1145), a random access memory (1146), an internal mass storage (such as an internal non-user-accessible hard disk drive, a solid-state drive, etc.) (1147), etc. may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessed in the form of at least one physical plug so as to be expandable via an additional central processing unit, a graphics processing unit, etc. Peripheral devices may be directly attached to the system bus (1148) of the core or connected via a peripheral bus (1149). In one example, a screen (1110) may be connected to the graphics adapter (1150). The architecture of the peripheral bus includes an external controller interface PCI, a universal serial bus USB, etc.

[0264] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) may execute certain instructions, which, when combined, may constitute the above-mentioned computer code. The computer code may be stored in the ROM (1145) or the RAM (1146). Transitional data may also be stored in the RAM (1146), while permanent data may be stored in, for example, the internal mass storage (1147). Fast storage and retrieval of any memory device may be achieved by using a cache memory, which may be closely associated with at least one CPU (1141), GPU (1142), mass storage (1147), ROM (1145), RAM (1146), etc.

[0265] The computer-readable storage medium may have computer code for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this application or may be well-known and available to those skilled in the art of computer software.

[0266] By way of example and not limitation, a computer system having an architecture (1100), particularly a core (1140), can provide the functionality of a processor (including a CPU, GPU, FPGA, accelerator, etc.) to execute software contained in at least one tangible computer-readable storage medium. Such a computer-readable storage medium can be the medium associated with the user-accessible mass storage described above, as well as a specific memory of the non-volatile core (1140), such as the core internal mass storage (1147) or ROM (1145). The software implementing the various embodiments of the present application can be stored in such a device and executed by the core (1140). Depending on specific needs, the computer-readable storage medium can include one or more storage devices or chips. The software can cause the core (1140), particularly the processors therein (including CPU, GPU, FPGA, etc.), to execute the specific processes or specific parts of the specific processes described herein, including defining data structures stored in RAM (1146) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality that is logically hardwired or otherwise included in a circuit (e.g., an accelerator (1144)), which can operate in place of or in conjunction with the software to execute the specific processes or specific parts of the specific processes described herein. In appropriate cases, references to software can include logic, and vice versa. In appropriate cases, references to computer-readable storage media can include circuits (such as integrated circuits (ICs)) that store and execute software, circuits that contain execution logic, or both. The present application encompasses any suitable combination of hardware and software.

[0267] The use of "at least one" in this application is intended to include any one or combination of the recited elements. For example, a reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof.

[0268] Although the present application has described multiple exemplary embodiments, various changes, permutations, and various equivalent substitutions of the embodiments are within the scope of the present application. Therefore, it should be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and are thus within the spirit and scope of the present application.

Claims

1. A video decoding method, characterized in that: include: A code stream is received, the code stream comprising encoded information of a picture sequence, the encoded information indicating that, for a current block in a current picture, inter-frame prediction is performed on the current block based on a reference block in a reference picture, and indicating an adjustment factor of a linear formula used when the inter-frame prediction is a formula-based inter-frame prediction, wherein: The formula-based inter-frame prediction generates prediction samples of the current block based on the linear formula, wherein: Inputting one or more reconstructed samples of the reference block into the linear formula; The linear formula includes one or more parameters, wherein the one or more parameters are derived based on the current template of the current block and the reference template of the reference block; applying the adjustment factor to the linear formula to generate an adjusted linear formula; and, At least one reconstructed sample of the current block is determined according to the adjusted linear formula.

2. The method according to claim 1, wherein: The applying the adjustment factor to the linear formula comprises: deriving the linear formula based on the current template of the current block and the reference template of the reference block, wherein the linear formula has an initial slope parameter; Applying the adjustment factor to the initial slope parameter to determine an adjusted slope parameter; An adjusted offset parameter is calculated according to the adjusted slope parameter, the reference template of the reference block and the current template of the current block.

3. The method according to claim 2, wherein: The calculating of the adjusted offset parameter further includes: Calculating a first weighted average of a plurality of first samples in the reference template of the reference block; Calculating a second weighted average of a plurality of second samples in the current template of the current block; Scaling the first weighted average value according to the adjusted slope parameter to generate a scaled first weighted average value; The adjusted offset parameter is calculated based on a difference between the second weighted average and the scaled first weighted average.

4. The method according to claim 3, wherein: The plurality of first samples includes each sample in the reference template, and the plurality of second samples includes each sample in the current template.

5. The method according to claim 3, wherein: The plurality of first samples include a first sample subset in the reference template, and the plurality of second samples include a second sample subset in the current template, wherein the first sample subset corresponds to the second sample subset based on motion information of the current block.

6. The method according to claim 5, wherein: The first sample subset includes a plurality of sub-sampled samples in the reference template based on a sub-sampling pattern.

7. The method according to claim 5, wherein: The first sample subset includes a plurality of corner samples and / or center samples in the reference template.

8. The method according to claim 1, wherein: The applying the adjustment factor to the linear formula comprises: deriving the linear formula based on the current template of the current block and the reference template of the reference block, wherein the linear formula has an initial slope parameter and an initial offset parameter; Applying the adjustment factor to the initial slope parameter to determine an adjusted slope parameter; Calculating an adjusted offset parameter based on the adjusted slope parameter, the reference template of the reference block, and the current template of the current block; According to the combination of the initial offset parameter and the adjusted offset parameter, a combined offset parameter is calculated, wherein the adjusted offset parameter and the combined offset parameter are used to form the adjusted linear formula.

9. The method according to claim 1, further comprising: Decoding a first syntax element from the codestream; A first adjustment factor is determined according to the first syntax element, wherein the inter-frame prediction uses multiple color components and multiple linear formulas, and the first adjustment factor is applied to at least a first color component of the multiple color components and / or at least a first linear formula of the multiple linear formulas.

10. The method according to claim 9, wherein: The first syntax element indicates an entry in a lookup table, the entry comprising a value of the first adjustment factor.

11. The method according to claim 10, wherein: The value of the first adjustment factor is applied to the first color component, the method further comprising at least one of the following: determining that the first color component is predefined; and / or A high-level syntax element indicating the first color component is decoded, where the high-level syntax element is at least one of a sequence header, a frame header, a picture header, and a slice header.

12. The method according to claim 10, wherein: The value of the first adjustment factor is applied to all color components.

13. The method according to claim 9, wherein: The first syntax element indicates an entry in a lookup table, the entry comprising at least a value of the first adjustment factor applied to the first color component and a value of a second adjustment factor applied to a second color component.

14. The method according to claim 1, further comprising: Decoding at least a first syntax element and a second syntax element from the codestream; determining, based on the first syntax element, a first adjustment factor for a first color component; Based on the second syntax element, a second adjustment factor for a second color component is determined.

15. The method according to claim 1, further comprising: One or more control flags are decoded, the one or more control flags indicating whether to apply the adjustment factor to one or more color components, respectively.

16. The method according to claim 15, further comprising: At least a first control flag is decoded, the first control flag indicating whether the adjustment factor is applied to a luma component and the second control flag indicating whether the adjustment factor is applied to a chroma component.

17. The method according to claim 15, further comprising: decoding a first control flag indicating whether the adjustment factor is applied to a luma component; When the first control flag is true, a second control flag is decoded, the second control flag indicating whether the adjustment factor is applied to chroma components.

18. The method according to claim 9, wherein: The decoding the first syntax element further comprises: When the current block is encoded in unidirectional prediction, decoding the first syntax element from the bitstream; When the current block is encoded using bidirectional prediction, the value of the first adjustment factor is set to zero.

19. A video encoding method, characterized in that: include: Determine to encode a current block in a current picture according to an inter-frame prediction, wherein the inter-frame prediction is a formula-based inter-frame prediction, and the formula-based inter-frame prediction has an adjustment factor, wherein, The formula-based inter-frame prediction generates prediction samples of the current block based on a linear formula, wherein: Inputting one or more reconstructed samples of a reference block in a reference picture into the linear formula; The linear formula includes one or more parameters, wherein the one or more parameters are derived based on the current template of the current block and the reference template of the reference block; determining the adjustment factor to be applied to the linear equation; and, The current block is encoded according to the adjustment factor to form a code stream.

20. A method for processing visual media data, characterized in that: include: Process the code stream of visual media data according to the format rules, where: The code stream includes encoded information of a picture sequence, the encoded information indicating that, for a current block in a current picture, inter-frame prediction is performed on the current block based on a reference block in a reference picture, and indicating an adjustment factor of a linear formula used when the inter-frame prediction is an inter-frame prediction based on a formula, wherein, The formula-based inter-frame prediction generates prediction samples of the current block based on the linear formula, wherein: Inputting one or more reconstructed samples in the reference block into the linear formula; The linear formula includes one or more parameters, and the one or more parameters are derived based on the current template of the current block and the reference template of the reference block; The format rules specify: applying the adjustment factor to the linear formula to generate an adjusted linear formula; and, At least one reconstructed sample of the current block is determined according to the adjusted linear formula.