Derivation and propagation of encoded information for inter prediction
By introducing formula-based inter prediction technology in video encoding and decoding, using the information of reference blocks to derive parameters, the prediction samples of the current block are generated, which solves the problem of low inter prediction efficiency in the prior art, and realizes more efficient video encoding and decoding.
Patent Information
- Application Number
- CN202480004367.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-07-10
- Filing Date
- 2024-07-11
- Publication Date
- 2025-05-27
AI Technical Summary
The existing video encoding and decoding technology is difficult to effectively utilize the information in the reference image in inter-frame prediction, resulting in low encoding efficiency.
By introducing formula-based inter prediction technology in video encoding and decoding, the template information of the current block and the reference block is used to derive parameters, and a prediction sample of the current block is generated. The technology includes receiving an encoded video code stream, determining whether the current block applies a formula-based inter prediction technique, and generating a reconstruction sample based on the derived parameters.
The encoding efficiency of inter-frame prediction is improved, and through more accurate prediction sample generation, the encoded data volume is reduced and the video decoding quality is improved.
Smart Images

Figure CN120051985A_ABST
Abstract
Description
Incorporation by Reference
[0001] This application claims the benefit of priority to U.S. Patent Application No. 18 / 769,323, filed on July 10, 2024, entitled "ON DERIVATION AND PROPAGATION OF CODED INFORMATION OF INTER PREDICTION", and to U.S. Provisional Application No. 63 / 526,155, filed on July 11, 2023, entitled "ON DERIVATION AND PROPAGATION OF CODED INFORMATION OF INTER PREDICTION". The disclosures of the prior applications are hereby incorporated by reference in their entirety. Technical Field
[0002] The description of embodiments of this application generally relates to aspects of video coding and decoding. Background Art
[0003] The background description provided herein is for the purpose of generally presenting the background of embodiments of this application. The work of the currently named inventors (to the extent that the work is described in this background art section), and aspects that may not constitute prior art as of the time of filing, are neither expressly nor impliedly admitted to be prior art to embodiments of this application.
[0004] Image / video compression can help transmit image / video data between different devices, storage devices, and networks while minimizing quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In an example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data in the currently reconstructed image to predict samples. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can predict samples in the current image based on previously reconstructed images through motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention
[0005] Aspects of embodiments of this application include bitstreams, methods, and apparatuses for video encoding / decoding. In some examples, an apparatus for video encoding / decoding includes processing circuitry.
[0006] Some aspects of embodiments of the present application provide a method for video decoding. The method includes receiving an encoded video bitstream, which includes first encoded information of a current block in a current image. The first encoded information indicates inter-frame prediction of the current block and the possibility of applying a formula-based inter-frame prediction technique in the inter-frame prediction of the current block. The inter-frame prediction of the current block is based on a reference block in a reference image of the current block. The formula-based inter-frame prediction technique generates prediction samples of the current block based on the formula by inputting at least one reconstructed sample in the reference block into the formula, and the formula includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The method further includes determining first formula-based inter-frame prediction information about the possibility of applying the formula-based inter-frame prediction technique to the current block based on at least one of the first encoded information of the current block and second encoded information of a second block reconstructed before the current block. The first formula-based inter-frame prediction information includes at least one of the following: a first control flag regarding the application possibility, a first template type for deriving at least one parameter, and a first formula type of the formula. The method further includes applying the formula-based inter-frame prediction technique to the current block to generate at least one reconstructed sample of the current block when the first control flag is true.
[0007] According to one aspect of embodiments of the present application, when the second encoded information of the second block indicates that a second control flag of the second block is true, the method includes determining the first control flag of the current block based on a comparison between a value of at least one parameter and a threshold, and the second control flag of the second block is used to indicate that the formula-based inter-frame prediction technique is applied to the second block. In some examples, the method includes deriving at least one parameter based on the current template of the current block and the reference template of the reference image, determining that the first control flag is false when a first derived parameter value of at least a first parameter among the at least one parameter is greater than an upper threshold of a first range of the first parameter or less than a lower threshold of the first range of the first parameter, and determining that the first control flag is true when the derived parameter value of the at least one parameter is within the respective corresponding ranges of the at least one parameter.
[0008] In some examples, the method includes deriving at least one parameter based on the current template of the current block and the reference template of the reference image, determining that the first control flag is false when each derived parameter value among the derived parameter values of the at least one parameter is greater than an upper threshold of an associated range or less than a lower threshold of the associated range, and determining that the first control flag is true when at least one of the derived parameter values of the at least one parameter is within the associated range.
[0009] In an example, the threshold is a constant value. In another example, the threshold is decoded from at least one of a sequence header, a slice header, a picture header, and a frame header in an encoded video bitstream. In another example, the threshold is determined based on at least one of a quantization parameter of a current block, a block size of the current block, a shape of the current block, and a temporal distance between the current block and a reference block.
[0010] In some embodiments, the method includes determining a first template type according to a default template type associated with a shape of a current block.
[0011] In some examples, when the shape of the current block is square and an upper template and a left template are available, the default template type is the upper and left template types.
[0012] In some examples, when the shape of the current block is rectangular, a ratio of a height to a width of the current block is greater than a predefined threshold, and a left template is available, the default template type is the left template type.
[0013] In some examples, when the shape of the current block is rectangular, a ratio of a width to a height of the current block is greater than a predefined threshold, and an upper template is available, the default template type is the upper template type.
[0014] In some embodiments, the method includes determining the first template type according to a second template type of a second block when a second control flag of the second block is true. In an example, the method includes determining that a first template type of the current block and the second template type of the second block are of the same type when it is determined that a first control flag of the current block and the second control flag of the second block are the same flag.
[0015] In some embodiments, the method includes determining that a first control flag of the current block and a second control flag of the second block are the same flag regardless of whether the second block uses a reference image for inter-frame prediction.
[0016] In some embodiments, the method includes: storing, in association with the current block, first formula-based inter-frame prediction information of a formula-based inter-frame prediction technique that has been applied to the current block, and reconstructing at least one third block in the current image based on the first formula-based inter-frame prediction information.
[0017] In some examples, the first formula-based inter-frame prediction information includes at least one of the following: a first control flag of the current block, a first template type of the current block, a first formula type of the current block, and / or a transform type of residual data of the formula-based inter-frame prediction technique.
[0018] In some embodiments, the first formula-based inter-frame prediction information is stored in units of m×n samples, where m and n are positive integers.
[0019] Some aspects of embodiments of the present application provide a method for video coding. The method includes: determining first formula-based inter prediction information of a first prediction candidate for a current block by using a formula-based inter prediction technique for the current block based on at least one of first encoded information of the current block in a current picture and second encoded information of a second block encoded before the current block. The formula-based inter prediction technique inputs at least one reconstructed sample of a reference block in a reference picture into a formula, and generates prediction samples for the current block based on the formula. The formula includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. The first formula-based inter prediction information includes at least one of the following: a first control flag for the first prediction candidate, a first template type for deriving at least one parameter, and a first formula type of the formula. When the first control flag is true, the method includes: applying the formula-based inter prediction technique to the current block to generate a first prediction candidate for the current block; and encoding the current block based on the first prediction candidate.
[0020] Aspects of embodiments of the present application also provide an apparatus for video coding or video decoding. The apparatus for a video coder / decoder includes processing circuitry configured to implement any of the methods for video coding or video decoding described above.
[0021] Aspects of embodiments of the present application also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the methods for video decoding / encoding described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0023] Figure 1 is a schematic diagram of an example block diagram of a communication system (100).
[0024] Figure 2 is a schematic diagram of an example block diagram of a decoder.
[0025] Figure 3 is a schematic diagram of an example block diagram of an encoder.
[0026] Figure 4 shows positions of spatial merge candidates of embodiments of the present application.
[0027] Figure 5 shows candidate pairs for redundancy check considered for spatial merge candidates of embodiments of the present application.
[0028] Figure 6 shows an example of motion vector scaling for temporal merge candidates.
[0029] Figure 7 Shows an example of a candidate position for time merging candidates of the current block.
[0030] Figure 8 Shows a diagram of a template in some examples.
[0031] Figure 9 Shows an overview flowchart of a decoding method for some aspects of an embodiment of the present application.
[0032] Figure 10 Shows an overview flowchart of an encoding method for some aspects of an embodiment of the present application.
[0033] Figure 11 Is a schematic diagram of a computer system in one aspect. Detailed Description
[0034] Figure 1 Shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an application example of the disclosed subject matter, a video encoder, and a video decoder in a streaming environment. The disclosed subject matter of the present application can be equivalently applied to other video-supported applications, including, for example, video conferencing, digital television, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0035] The video processing system (100) may include an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera, and the video source creates an uncompressed video image stream (102). In an embodiment, the video image stream (102) includes samples taken by a digital camera. Compared with the encoded video data (104) (or the encoded video bitstream), the video image stream (302) is depicted as a thick line to emphasize the high data volume of the video image stream. The video image stream (102) can be processed by an electronic device (120), and the electronic device (120) includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or enforce aspects of the disclosed subject matter described in more detail below. Compared with the video image stream (102), the encoded video data (104) (or the encoded video bitstream (104)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (104) (or the encoded video bitstream (104)), which can be stored on a streaming server (105) for future use. At least one streaming client subsystem, such as Figure 1The client subsystems (106) and client subsystems (108) therein can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video image stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (such as a video bitstream) can be encoded according to certain video coding / compression standards. Examples of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and the present application can be used in the context of the VVC standard.
[0036] It should be noted that the electronic device (120) and the electronic device (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown), and the electronic device (130) can further include a video encoder (not shown).
[0037] Figure 2 is a block diagram of an example of a video decoder (210). The video decoder (210) can be provided in an electronic device (230). The electronic device (230) can include a receiver (231) (such as a receiving circuit). The video decoder (210) can be used to replace Figure 1 the video decoder (110) in the embodiment.
[0038] A receiver (231) may receive at least one encoded video sequence to be decoded by a video decoder (210) (e.g., included in a video bitstream). In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequence may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data as well as other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not shown). The receiver (231) may separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and an entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided external to the video decoder (210) (not shown). In other cases, a buffer memory (not shown) is provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle playback timing. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (415), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not shown) external to the video decoder (210).
[0039] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), such as a display device that is not part of the electronic device (230) but may be coupled to the electronic device (230), such as Figure 2As shown. The control information for the display device may be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) can extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) can also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.
[0040] The parser (220) can perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0041] Depending on the type of the encoded video image or a part of the encoded video image (e.g., inter-frame image and intra-frame image, inter-frame block and intra-frame block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved can be controlled by the subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and multiple units below are not described.
[0042] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually divided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually divide into the functional units below.
[0043] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives the quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output a block including sample values, and the sample values can be input into the aggregator (255).
[0044] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to intra-coded blocks. An intra-coded block is a block that does not use predictive information from a previously reconstructed image, but may use predictive information from a previously reconstructed part of the current image. Such predictive information may be provided by the intra-image prediction unit (252). In some cases, the intra-image prediction unit (252) generates a surrounding block with the same size and shape as the block being reconstructed using the reconstructed information extracted from the current image buffer (258). For example, the current image buffer (258) buffers a partially reconstructed current image and / or a fully reconstructed current image. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.
[0045] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference image memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from an address in the reference image memory (257) can be controlled by a motion vector, and the motion vector is in the form of the symbol (221) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference image components. Motion compensation may also include interpolation of the sample values extracted from the reference image memory (257), a motion vector prediction mechanism, etc. when using sub-sample accurate motion vectors.
[0046] The output samples of the aggregator (255) can be employed by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques that are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be available to the loop filter unit (256) as symbols (221) from a parser (220). Video compression techniques can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded image or an encoded video sequence, and to previously reconstructed and loop-filtered sample values.
[0047] The output of the loop filter unit (256) can be a sample stream that can be output to a display device (212) and stored in a reference image memory (257) for subsequent inter-frame image prediction.
[0048] Once fully reconstructed, some encoded images can be used as reference images for future prediction. For example, once the encoded image corresponding to the current image is fully reconstructed and the encoded image is identified as a reference image (by, for example, a parser (220)), the current image buffer (258) can become part of the reference image memory (257), and a new current image buffer can be reallocated before starting to reconstruct subsequent encoded images.
[0049] The video decoder (210) can perform decoding operations according to a predetermined video compression technique or a standard such as ITU-T H.265. The encoded video sequence can conform to the syntax specified by the video compression technique or standard used, in the sense that the encoded video sequence follows the syntax of the video compression technique or standard and the profile recorded in the video compression technique or standard. Specifically, a profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum image size, maximum frame rate, maximum reconstruction sampling rate (measured, for example, in megasamples per second), maximum reference image size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.
[0050] In one aspect, a receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The additional data may be used by a video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, a temporal, spatial, or signal noise ratio (SNR) enhancement layer, redundant slices, redundant pictures, forward error correction codes, etc.
[0051] Figure 3 is a block diagram of an example of a video encoder (303). The video encoder (303) is provided in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used in place of Figure 1 the video encoder (103) in the embodiment.
[0052] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the embodiment), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).
[0053] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures, which are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include at least one sample depending on the sampling structure, color space, etc. used. Those skilled in the art can easily understand the relationship between pixels and samples. The following focuses on describing samples.
[0054] According to one aspect, a video encoder (303) may encode and compress images of a source video sequence into an encoded video sequence (343) in real time or under any other desired time constraint. Enforcing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls and is functionally coupled to other functional units as described below. For simplicity, couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used for other suitable functions related to the video encoder (303) optimized for a certain system design.
[0055] In some aspects, the video encoder (303) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input image to be encoded and reference images) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference image memory (334). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference image memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference image samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This reference image synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related technologies.
[0056] The operation of the “local” decoder (333) may be the same as that of the “remote” decoder, for example, which has been described in detail above in connection with Figure 2 the video decoder (210). However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0057] In one aspect, any decoder technique other than parsing / entropy decoding present in the decoder exists in the corresponding encoder in the same or substantially the same functional form. Accordingly, this application focuses on decoder operations. The description of encoder techniques can be simplified because the encoder techniques are inverse to the decoder techniques described comprehensively. More detailed descriptions are provided in certain areas below.
[0058] During operation, in some embodiments, the source encoder (330) may perform motion-compensated predictive coding. With reference to at least one previously encoded image designated as a "reference image" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input image. In this way, the coding engine (332) encodes the difference between the pixel blocks of the input image and the pixel blocks of the reference image, and the reference image can be selected as the prediction reference for the input image.
[0059] The local video decoder (333) may decode the encoded video data of the image that can be designated as a reference image based on the symbols created by the source encoder (330). The operation of the coding engine (332) may be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 3 not shown), the reconstructed video sequence is generally a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference image and may store the reconstructed reference image in the reference image memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference image, which has the same content (without transmission errors) as the reconstructed reference image to be obtained by the remote video decoder.
[0060] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new image to be encoded, the predictor (335) may search in the reference image memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference image motion vectors, block shapes, etc., that can be used as an appropriate prediction reference for the new image. The predictor (335) may operate on a per-pixel-block basis of the sample blocks to find a suitable prediction reference. In some cases, according to the search results obtained by the predictor (335), it can be determined that the input image may have a prediction reference obtained from multiple reference images stored in the reference image memory (334).
[0061] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding the video data.
[0062] The outputs of all the above functional units can be entropy - encoded in an entropy encoder (345). The entropy encoder (345) performs lossless compression on the symbols generated by various functional units according to techniques such as Huffman coding, variable - length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.
[0063] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) to prepare for transmission over a communication channel (360), which can be a hardware / software link leading to a storage device that will store the encoded video data. The transmitter (340) can merge the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).
[0064] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded image type to each encoded image, but this may affect the encoding technique applicable to the corresponding image. For example, an image can typically be assigned to any of the following image types:
[0065] An intra - frame image (I - image) is an image that can be encoded and decoded without using any other image in the sequence as a prediction source. Some video codecs allow different types of intra - frame images, including, for example, an Independent Decoder Refresh (IDR) image.
[0066] A predictive image (P - image) is an image that can be encoded and decoded using intra - frame prediction or inter - frame prediction, which uses motion vectors and reference indices to predict the sample values of each block.
[0067] A bi - directional predictive image (B - image) is an image that can be encoded and decoded using intra - frame prediction or inter - frame prediction, which uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive images can use more than two reference images and associated metadata for reconstructing a single block.
[0068] Source images can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be prediction-encoded with reference to other (already encoded) blocks, and the other blocks are determined according to the coding assignment of the corresponding image applied to the block. For example, blocks of an I image can be non-prediction-encoded, or the blocks can be prediction-encoded with reference to already encoded blocks of the same image (spatial prediction or intra-frame prediction). Pixel blocks of a P image can be prediction-encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference image. Blocks of a B image can be prediction-encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference images.
[0069] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (303) can perform various compression operations, including prediction coding operations that utilize the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.
[0070] In one aspect, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant images, and slices, SEI messages, VUI parameter set fragments, etc.
[0071] The captured video can be multiple source images (video images) in a time series. Intra-frame image prediction (often simplified to intra-frame prediction) utilizes the spatial correlation in a given image, while inter-frame image prediction utilizes the (temporal or other) correlation between images. In an embodiment, the particular image being encoded / decoded is segmented into blocks, and the particular image being encoded / decoded is referred to as the current image. When a block in the current image is similar to a reference block in a reference image that has been previously encoded and is still buffered in the video, the block in the current image can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference image, and in the case of using multiple reference images, the motion vector can have a third dimension that identifies the reference image.
[0072] In some aspects, bidirectional prediction techniques can be used in inter - frame image prediction. According to the bidirectional prediction technique, two reference images are used, for example, a first reference image and a second reference image that are both before the current image in the decoding order (but may be past and future respectively in the display order) in the video. A block in the current image can be encoded by a first motion vector pointing to a first reference block in the first reference image and a second motion vector pointing to a second reference block in the second reference image. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.
[0073] In addition, the merge mode technique can be used in inter - frame image prediction to improve coding efficiency.
[0074] According to some aspects disclosed in the present application, predictions such as inter - frame image prediction and intra - frame image prediction are performed in units of blocks. For example, according to the HEVC standard, an image in a video image sequence is segmented into coding tree units (CTUs) for compression. CTUs in the image have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into at least one coding unit (CU) by a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In one aspect, each CU is analyzed to determine the prediction type for the CU, such as an inter - frame prediction type or an intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into at least one prediction unit (PU). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In an embodiment, the prediction operation in encoding (encoding / decoding) is performed in units of prediction blocks. Taking the luminance prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.
[0075] Note that the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using any suitable technology. In one aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one integrated circuit. In another aspect, the video encoders (103) and (303) and the video decoders (110) and (210) can be implemented using at least one processor that executes software instructions.
[0076] Embodiments of the present application provide techniques for deriving and propagating encoded information in inter-frame prediction coding.
[0077] Various inter-frame prediction modes can be used in video coding and decoding. For example, in VVC, for a CU in inter-frame prediction, the motion parameters can include at least one MV, at least one reference image index, a reference image list usage index, and additional information related to certain coding features for generating inter-frame prediction samples. The motion parameters can be signaled explicitly or implicitly. When a CU is encoded in skip mode, the CU can be associated with a PU and may not have significant residual coefficients, no encoded motion vector delta or MV difference (e.g., MVD), or reference image index. A merge mode can be defined, where the motion parameters of the current CU are obtained from at least one neighboring CU, including spatial candidates and / or temporal candidates, and optionally (as introduced in VVC) additional information. The merge mode can be applied to CUs in inter-frame prediction, not just for skip mode. In an example, an alternative to the merge mode is the explicit transmission of motion parameters, where at least one MV, the corresponding reference image index for each reference image list, and a reference image list usage flag, and other information are signaled explicitly for each CU.
[0078] In an embodiment, such as in VVC, the VVC Test model (VTM) reference software includes at least one enhanced inter-prediction coding tool, which includes: extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8×8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter-prediction and related methods are described in detail below.
[0079] Extended merge prediction may be used in some examples. In an example, such as in VTM4, a merge candidate list is constructed by sequentially incorporating the following five types of candidates: at least one spatial motion vector predictor (MVP) from at least one spatially adjacent CU, at least one temporal MVP from at least one collocated CU, at least one history-based MVP (HMVP) from a first-in-first-out (FIFO) table, at least one pairwise-average MVP, and at least one zero MV.
[0080] The size of the merge candidate list can be signaled in the stripe header. In the example, in VTM4, the maximum allowed size of the merge candidate list is 6. For each CU encoded in merge mode, the index of the best merge candidate (e.g., the merge index) can be encoded using truncated unary binarization (TU). The first binary digit of the merge index can be encoded with context (e.g., context adaptive binary arithmetic coding, CABAC), and bypass coding can be used for the other binary digits.
[0081] Some examples of the generation process of merge candidates for each category are provided below. In an embodiment, at least one spatial candidate is derived in the following manner. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In the example, up to four merge candidates are selected from the candidates located at the Figure 4 positions shown.
[0082] Figure 4 The positions of the spatial merge candidates according to an embodiment of the present application are shown. Referring to Figure 4 , the derivation order is B1, A1, B0, A0, and B2. Only when any of the CUs located at A0, B0, B1, and A1 is unavailable (e.g., because the CU belongs to another stripe or another tile) or is intra-coded, the position B2 is considered. After adding the candidate located at A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving the coding efficiency.
[0083] To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the Figure 5 candidate pairs connected by arrows in are considered, and a candidate is added to the candidate list only when the candidate does not have the same motion information after the redundancy check.
[0084] Figure 5 The candidate pairs considered for the redundancy check of the spatial merge candidates according to an embodiment of the present application are shown. Referring to Figure 5 , the candidate pairs connected by the corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, the candidates at positions B1, A0, and / or B2 can be compared with the candidate at position A1, and the candidates at positions B0 and / or B2 can be compared with the candidate at position B1.
[0085] In an embodiment, at least one temporal candidate is derived in the following manner. In the example, only one temporal merge candidate is added to the candidate list.Figure 6 An example of a motion vector for temporal merge candidates after scaling is shown. To derive the temporal merge candidates for the current CU (611) in the current picture (601), the scaled MV (621) can be derived based on the co-located CU (612) in the co-located reference picture (604) (e.g., as shown by the dashed line in Figure 6 . The reference picture list for deriving the co-located CU (612) can be explicitly signaled in the slice header. The scaled MV (621) of the temporal merge candidates as shown by the dashed line in Figure 6 can be obtained. The scaled MV (621) can be the MV of the co-located CU (612) scaled using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the co-located reference picture (604) of the co-located picture (603) and the co-located picture (603). The reference picture index of the temporal merge candidates can be set to zero. The co-located picture is the reference picture that is used as the source picture for temporal motion information derivation. The co-located picture can be determined in one of two lists (referred to as list0 or list1). In some examples, the encoder can use appropriate syntax techniques to determine the co-located picture and signal the co-located picture.
[0086] Figure 7 Examples of candidate positions (e.g., C0 and C1) for the temporal merge candidates of the current CU are shown. The position of the temporal merge candidate can be selected from the candidate positions C0 and C1. The candidate position C0 is at the lower right corner of the co-located CU (710) of the current CU. The candidate position C1 is at the center of the co-located CU (710) of the current CU. If the CU at the candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, the candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at the candidate position C0 is available, is inter-coded, and is in the current row of the CTU, the candidate position C0 is used to derive the temporal merge candidate.
[0087] According to some aspects of embodiments of the present application, a formula-based prediction method can be used for inter prediction or intra prediction. For inter prediction, the formula-based prediction method can use a formula to generate samples of the current block in the current picture based on reference samples of a reference block in a reference picture. For intra prediction, the formula-based prediction method can use a formula to generate the first color component of the current block based on the second color component of the current block. In some examples, the parameters in the formula for the formula-based prediction method can be derived based on a template of the current block.
[0088] In some examples, local illumination compensation (LIC) is used as an inter-frame prediction technique to model the local illumination change between a current block and a predicted block (also referred to as a reference block) of the current block by using a linear function. The predicted block is in a reference image and can be pointed to by a motion vector (MV). The parameters of the linear formula can include a scale α and an offset β, and the linear formula can be expressed as α×p[x,y]+β to compensate for the illumination change, where p[x,y] represents the reference sample at position [x,y] in the reference block (also referred to as the predicted block), and the MV points from the current block to the reference block. In some examples, the scale α and the offset β can be derived based on the template of the current block and the corresponding reference template of the reference block by using the least squares method, so no signaling overhead is required, except that an LIC flag can be signaled to indicate the use of LIC. The scale α and the offset β derived based on the template of the current block can be referred to as a template-based parameter set.
[0089] In some examples, LIC is used for inter-frame CUs of uni-directional prediction. In some examples, the intra-frame neighboring samples of the current block (the neighboring samples predicted using intra-frame prediction) can be used for LIC parameter derivation. In some examples, for a block with fewer than 32 luma samples, LIC is disabled. In some examples, for non-sub-block modes (e.g., non-affine modes), LIC parameter derivation is performed based on the modulo-block samples of the current CU (instead of the partial modulo-block samples of the top-left 16×16 unit). In some examples, LIC parameter derivation is performed based on partial modulo-block samples (such as the partial modulo-block samples of the top-left 16×16 unit). In some examples, the template samples of the reference block are determined by performing motion compensation (MC) on the block using the MV of the block without rounding the MV of the block to integer pixel precision.
[0090] In some examples, cross-component prediction can be used as an intra-frame prediction technique. Cross-component prediction can include a first technique called cross component linear model (CCLM), a second technique called multi-model linear model (MMLM), a third technique called convolutional cross-component model (CCCM), and a fourth technique called gradient linear model (GLM).
[0091] For example, the first technique CCLM is used to reduce cross-component redundancy. In CCLM, chrominance samples are predicted based on the reconstructed luma samples of the same CU using a linear model (also known as a linear formula), such as using Equation (1): pred C (i,j) = a · rec′ L (i,j) + b Equation (1) where pred C (i,j) represents the predicted chrominance sample in the CU, and rec L (i,j) represents the downsampled reconstructed luma sample of the same CU. The CCLM linear model includes parameters (a and b), and in an example, these parameters can be derived using at most four adjacent chrominance samples and their corresponding downsampled luma samples. In an example, the at most four adjacent chrominance samples and their corresponding downsampled luma samples are referred to as the template of the CU.
[0092] In some examples, based on the positions of adjacent chrominance samples, CCLM can include different modes, which are referred to as LM_T (LM top mode or upper mode LM_A), LM_L (LM left mode), and LM_LT (LM left-top mode or left-upper mode LM_LA or just LM mode). For example, if the size of the current chrominance block is W×H, then W' and H' can be set for various modes in CCLM. When the LM mode (also known as LM_LT or LM_LA) is applied, W’ = W, H’ = H; when the LM_A mode is applied, W’ = W + H; when the LM_L mode is applied, H’ = H + W.
[0093] Note that MMLM, CCCM, and GLM also use functions for prediction. The parameters of these functions can be derived based on the template.
[0094] Note that the following description uses inter prediction to illustrate the techniques for deriving encoded information based on formula-based prediction methods, and these techniques can be applied to derive the encoded information of intra prediction.
[0095] Some aspects of the embodiments of the present application provide techniques for deriving and propagating encoded information in inter prediction coding. For example, an encoder / decoder can determine formula-based inter prediction information based on at least one of the encoded information of the current block and the encoded information of a second block that has been reconstructed before the current block, and the formula-based inter prediction information is used by a formula-based inter prediction technique that may be applied to the current block. The formula-based inter prediction information includes at least one of a control flag regarding the application possibility, a template type for deriving at least one parameter, and a formula type of the formula.
[0096] In some aspects of embodiments of the present application, some inter - frame prediction techniques are designed to minimize the distortion between a current block and its predicted block in a corresponding reference image. For example, an inter - frame prediction technique (also referred to as a first method of an inter - frame prediction method, a formula - based inter - frame prediction technique, a function - based inter - frame prediction technique, or a model - based inter - frame prediction technique) can use the original predicted block in the reference image as the input of a formula to generate the current block in the current image using a formula (such as a non - linear formula, a linear formula, etc.). For example, the inter - frame prediction technique can generate a prediction of the samples in the current block based on a formula that takes at least one predicted sample in the reference image as input. The formula can include linear terms or non - linear terms and can include at least one parameter that can be derived. Note that LIC is one of such inter - frame prediction techniques.
[0097] In some examples, the formula is a linear formula and can be represented by , where n is a non - negative integer, and p(x i , y i ) is the predicted sample at position (x i , y i ) in the reference image, and this predicted sample is pointed to based on the MV associated with the current block. Further, a set of predicted samples represented by p(x i , y i ) (where i = 0, ……, n) can be a set of predicted samples around the corresponding samples in the reference samples pointed to by the MV from the current sample to be predicted. In some examples, the parameters α i and β can be derived by minimizing the difference between the current block template and its predicted block template (for example, by using the least - squares method) based on the template of the current block (also referred to as the current block template) and the template of the predicted block of the current block (also referred to as the predicted block template). The template of the current block consists of the spatially adjacent reconstructed samples of the current block, and the template of the predicted block consists of the spatially adjacent reconstructed samples of the predicted block.
[0098] Figure 8 FIGs. show diagrams of templates in some examples. For example, the template (810) is called an L - shaped template T L , and includes the adjacent samples in the row above, the left - most column, and the upper - left corner of the current block (also referred to as the current coding block); the template (820) is called an above - and - left template T a+l , and includes the adjacent samples in the row above and the left - most column of the current block; the template (830) is called an above template T a , and includes the adjacent samples in the row above the current block; while the template (840) is called a left template T l , and includes the adjacent samples in the left - most column of the current block. Note that the template can include Figure 8Adjacent samples having other suitable shapes not shown.
[0099] In some examples, multiple candidate template types may also be supported, and one candidate template type may be selected to derive the parameters of the linear formula. Syntax may be signaled in the codestream (eg, at the block level) to indicate which candidate template type is selected.
[0100] In some examples, a control flag may be signaled in a bitstream (e.g., at a block level) in a manner associated with a formula-based inter-frame prediction technique to indicate whether the formula-based inter-frame prediction technique is applied to a current block. Alternatively, the value of the control flag may also be inherited from another encoded block. More specifically, a first control flag for a formula-based inter-frame prediction technique associated with a current block may be inherited from a second control flag for a formula-based inter-frame prediction technique associated with another encoded block or multiple encoded blocks. In addition, a control flag may be derived at a coding block level to adaptively determine whether to apply a formula-based inter-frame prediction technique.
[0101] In some examples, a first control flag for a formula-based inter-frame prediction technique associated with a current block is inherited from a second control flag for a formula-based inter-frame prediction technique associated with another coded block or multiple coded blocks. In some examples, the coded information of the formula-based inter-frame prediction technique can be derived from an adjacent coded block, a non-adjacent coded block, or a coded block storing coded information in a buffer.
[0102] Note that in the present embodiment, the generality is not limited. In the example, the term "parameter" refers to the parameter α used to determine the linear formula for deriving the prediction block. i and β. The term “template type” refers to different template shapes, such as but not limited to Figure 8 One of the template types shown for parameter derivation in nonlinear or linear formulas.
[0103] Some aspects of the embodiments of the present application provide a technique for deriving information for applying a formula-based inter-frame prediction technique to a current block (such as a control flag, template type, formula type, etc. for the formula-based inter-frame prediction technique for the current block) using encoded information (such as encoded information of a current block and / or encoded information of at least one other encoded block).
[0104] According to one aspect of an embodiment of the present application, when a control flag associated with another coded block (e.g., an adjacent neighboring block, a non-adjacent neighboring block, a temporally co-located block, etc.) is true, based on a derived parameter (e.g., a parameter α i The control flag of the current block is derived by comparing the value of β and β) with a predefined threshold.
[0105] In some embodiments, when one of these parameters is greater than and / or less than a corresponding threshold, the control flag is false; otherwise, the control flag is true. In some examples, the value of the control flag of another encoded block (e.g., an encoded block encoded using a formula-based inter prediction technique in a buffer) is true, and when the value of at least one of the parameters of the formula used to apply the formula-based inter prediction technique to the current block (e.g., based on template derivation) is out of range (such as greater than an upper threshold or less than a lower threshold), the control flag is false; otherwise, if all the parameters of the formula used to apply the formula-based inter prediction technique to the current block are within a predefined range, the control flag for applying the formula-based inter prediction technique is true.
[0106] In some embodiments, when all parameters are greater than and / or less than their respective corresponding thresholds, the control flag is false; otherwise, the control flag is true. In some examples, the value of the control flag of another encoded block (e.g., an encoded block encoded using a formula-based inter prediction technique in a buffer) is true, and when the value of each of the parameters of the formula used to apply the formula-based inter prediction technique to the current block (e.g., based on template derivation) is out of the specific range of that parameter (such as greater than the upper threshold of that parameter or less than the lower threshold of that parameter), the control flag is false; otherwise, if at least one parameter is within the specific range of each parameter, the control flag is true.
[0107] In some embodiments, at least one threshold may be a predefined constant value. In some examples, at least one constant value may be signaled in an encoded video bitstream (e.g., but not limited to, sequence header, slice header, picture header, frame header, etc.).
[0108] In some embodiments, the at least one threshold may be determined by other encoded information (including but not limited to: a function of quantization parameter, the size / shape of an encoded block (e.g., the current block), the temporal distance between a predicted block and the current block).
[0109] According to one aspect of the embodiments of the present application, a specific template type may be specified as the default template type for the current block. The specific template type may be, but is not limited to, Figure 8 one of the template types shown. In some embodiments, different block shapes may have different template types.
[0110] In some embodiments, when the top and left templates T a+l (such as Figure 8 the top and left templates (820) in a+l are both available, the top and left template T
[0111] In some embodiments, when the left template T l (like Figure 8 When the left template (840) in is available, the left template T l The default template type is specified for rectangular blocks where the ratio of block height to block width is greater than a predefined threshold. The predefined threshold is a positive value. In some examples, the predefined threshold can be a constant value, or the predefined threshold can be signaled in a high-level syntax (such as a sequence header, frame header, picture header, slice header, etc.) in the encoded video bitstream.
[0112] In some embodiments, when the top template T a (like Figure 8 When the upper template (830) in the a The default template type specified as a rectangular block where the ratio of block width to block height is greater than a predefined threshold. The predefined threshold is a positive value. In some examples, the predefined threshold can be a constant value, or the predefined threshold can be signaled in a high-level syntax (such as a sequence header, frame header, picture header, slice header, etc.) in the encoded video bitstream.
[0113] According to one aspect of an embodiment of the present application, when the control flag of another encoded block (such as an adjacent neighboring block, a non-adjacent neighboring block, a temporally co-located block, etc.) is true, the template type of the formula-based inter-frame prediction technology can be derived based on the template type selected for the other encoded block.
[0114] In some embodiments, when a control flag for a formula-based inter prediction technique applied to a current block is derived from another coded block, the template type of the current block inherits the template type selected for the formula-based inter prediction technique of the (same) other coded block.
[0115] According to one aspect of an embodiment of the present application, when a control flag associated with another coded block is true, the control flag may be inherited regardless of whether the other coded block is coded using the same reference image as the current block. In other words, the coded information of the other coded block may be partially merged, for example, still using a signal to represent the reference image, but inheriting the control flag, or vice versa.
[0116] Some aspects of embodiments of the present application also provide techniques for storing information related to formula-based inter prediction techniques for encoded blocks, and the stored information can be used to encode / decode subsequent blocks during the encoding / decoding process. The stored information includes, but is not limited to, control flags, template types, formula types, etc. In an example, the buffer for buffering adjacent samples of the current block also includes information on formula-based inter prediction techniques for the adjacent samples. In another example, the row buffer for buffering some samples in the above CTU also includes information related to formula-based inter prediction techniques for those samples in the above CTU.
[0117] In some embodiments, the signal-represented and / or derived template type of the formula-based inter prediction technique can be stored for subsequent blocks during the encoding or decoding process.
[0118] In some embodiments, the derived control flag of the formula-based inter prediction technique of the current block can be stored for subsequent blocks during the encoding or decoding process.
[0119] In some embodiments, the transform type of the derived residual signal associated with the formula-based inter prediction technique of the current block can be stored for subsequent blocks during the encoding or decoding process.
[0120] In some embodiments, the information related to the formula-based inter prediction technique is stored in m×n units, where m and n are non-zero positive integers. In an example, the information on the formula-based inter prediction technique is stored in units of 4×4 (e.g., stored once every 4×4 samples). In another example, the information on the formula-based inter prediction technique is stored in units of 8×8 (e.g., stored once every 8×8 samples).
[0121] Figure 9 An overview flowchart of a method (900) according to an aspect of embodiments of the present application is shown. The method (900) can be used in a video decoder. In various aspects, the method (900) is executed by a processing circuit (such as a processing circuit that executes the functions of the video decoder (110), a processing circuit that executes the functions of the video decoder (210), etc.). In some aspects, the method (900) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the method (900). The method starts at (S901) and proceeds to (S910).
[0122] At (S910), an encoded video bitstream is received, which includes first encoded information of a current block in a current picture. The first encoded information indicates inter prediction of the current block, and also indicates a possibility of applying a formula-based inter prediction technique in the inter prediction of the current block. The inter prediction of the current block is based on a reference block in a reference picture of the current block. The formula-based inter prediction technique generates prediction samples of the current block based on a formula, where at least one reconstructed sample in the reference block is input into the formula. The formula includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. A second block may be an adjacent neighboring block of the current block, a non-adjacent neighboring block of the current block, and a co-located block of the current block in a different picture.
[0123] At (S920), first formula-based inter prediction information regarding a possibility of applying the formula-based inter prediction technique to the current block is determined based on at least one of the first encoded information of the current block and second encoded information of a previously reconstructed second block before the current block. The first formula-based inter prediction information includes at least one of the following information regarding the possibility of applying the formula-based inter prediction technique: a first control flag, a first template type, and a first formula type.
[0124] At (S930), when the first control flag is true, the formula-based inter prediction technique is applied to the current block to generate at least one reconstructed sample of the current block.
[0125] According to an aspect of an embodiment of the present application, when the second encoded information of the second block indicates that a second control flag of the second block is true, the first control flag of the current block is determined based on a comparison between a value of at least one parameter and a threshold, and the second control flag of the second block is used to indicate that the formula-based inter prediction technique is applied to the second block.
[0126] In some examples, at least one parameter is derived based on the current template of the current block and the reference template of the reference picture. When a first derived parameter value of at least a first parameter among the at least one parameter is not within a first range defined for the first parameter (such as greater than an upper threshold of the first range of the first parameter or less than a lower threshold of the first range of the first parameter), the first control flag is determined to be false. When the derived parameter values of the at least one parameter are within the corresponding ranges of the at least one parameter, the first control flag is determined to be true.
[0127] In some embodiments, the at least one parameter is derived based on a current template of a current block and a reference template of a reference image. When each of the derived parameter values of the at least one parameter is not within an associated range defined for the corresponding parameter (such as greater than an upper threshold of the associated range or less than a lower threshold of the associated range), a first control flag is determined to be false. When at least one of the derived parameter values of the at least one parameter is within the associated range, the first control flag is determined to be true.
[0128] In some examples, the at least one threshold is a constant value, or the at least one threshold is decoded from at least one of a sequence header, a slice header, a picture header, and a frame header in an encoded video bitstream.
[0129] In some examples, the threshold is determined based on at least one of a quantization parameter of the current block, a block size of the current block, a shape of the current block, and a temporal distance between the current block and a reference block.
[0130] In some embodiments, a first template type is determined according to a default template type associated with a shape of a current block. In some examples, when the shape of the current block is square and an upper template and a left template are available, the default template type is the upper and left template types (e.g., Figure 8 the upper and left template types (820) in
[0131] In some examples, when the shape of the current block is rectangular, a ratio of a height to a width of the current block is greater than a predefined threshold, and a left template is available, the default template type is the left template type (e.g., Figure 8 the left template type (840) in
[0132] In some examples, when the shape of the current block is rectangular, a ratio of a width to a height of the current block is greater than a predefined threshold, and an upper template is available, the default template type is the upper template type (e.g., Figure 8 the upper template type (830) in
[0133] In some embodiments, when a second control flag of a second block is true, the first template type is determined according to a second template type of the second block.
[0134] In some embodiments, when it is determined that a first control flag of a current block and a second control flag of a second block are the same flag, a first template type of the current block and a second template type of the second block are the same type.
[0135] In some embodiments, it is determined that a first control flag of a current block and a second control flag of a second block are the same flag, regardless of whether the second block uses the same reference image as the first block for inter prediction.
[0136] In some embodiments, first formula-based inter prediction information is stored in association with a current block, the first formula-based inter prediction information being for a formula-based inter prediction technique that has been applied to the current block. The stored information can be used for reconstruction of subsequent blocks. For example, at least one third block in the current image is reconstructed based on the first formula-based inter prediction information. In an example, formula-based inter prediction information for a third block is determined according to the first formula-based inter prediction information.
[0137] Note that the first formula-based inter prediction information includes at least one of the following: a first control flag of the current block, a first template type of the current block, a first formula type of the current block, and / or a transform type of residual data of the formula-based inter prediction technique.
[0138] In some examples, the first formula-based inter prediction information is stored in units of m×n samples, where m and n are positive integers.
[0139] Then, the method proceeds to (S999) and terminates.
[0140] The method (900) can be appropriately adjusted. At least one step in the method (900) can be modified and / or omitted. At least one additional step can be added. Any suitable order of execution can be used.
[0141] Figure 10 An overview flowchart of a method (1000) according to an aspect of an embodiment of the present application is shown. The method (1000) can be used in a video encoder. In various aspects, the method (1000) is executed by a processing circuit (such as a processing circuit that executes the functions of the video encoder (103), a processing circuit that executes the functions of the video encoder (303), etc.). In some aspects, the method (1000) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the method (1000). The method starts at (S1001) and proceeds to (S1010).
[0142] At (S1010), based on at least one of the first encoded information of the current block in the current image and the second encoded information of the second block encoded before the current block, determine the first formula-based inter prediction information of the first prediction candidate, where the first prediction candidate uses the formula-based inter prediction technique for the current block. The formula-based inter prediction technique inputs at least one reconstructed sample of the reference block in the reference image into a formula, and generates the prediction sample of the current block based on the formula. Wherein the formula includes at least one parameter, and the at least one parameter is derived based on the current template of the current block and the reference template of the reference block. The first formula-based inter prediction information includes at least one of a first control flag, a first template type, and a first formula type for the formula-based inter prediction technique. The second block may be an adjacent neighboring block of the current block, may be a non-adjacent neighboring block of the current block, and may be a corresponding block of the current block in different images.
[0143] At (S1020), when the first control flag is true, apply the formula-based inter prediction technique to the current block to generate the first prediction candidate of the current block.
[0144] At (S1030), encode the current block based on the first prediction candidate. In some examples, at least two prediction candidates including the first prediction candidate may be generated. Cost values may be calculated for the at least two prediction candidates respectively. For example, calculate a first cost value and associate it with the first prediction candidate. Then, a prediction candidate may be selected from these prediction candidates based on the cost values associated with each prediction candidate. In an example, determine whether to apply the formula-based inter prediction technique to the current block based at least on the first cost value associated with the first prediction candidate. When the first prediction candidate has the lowest cost value, select the first prediction candidate, and further encode the current block according to the first prediction candidate.
[0145] Note that in some examples, among at least one candidate, the encoder may calculate the cost values associated with the at least one candidate, and select one candidate from these candidates based on these cost values. Then, encode the current block according to the selected candidate.
[0146] According to an aspect of the embodiments of the present application, when the second encoded information of the second block indicates that the second control flag of the second block is true, determine the first control flag of the current block based on the comparison between the value of at least one parameter and a threshold, and the second control flag of the second block is used to indicate that the formula-based inter prediction technique is applied to the second block.
[0147] In some examples, at least one parameter is derived based on a current template of a current block and a reference template of a reference image. When, among the at least one parameter, a first derived parameter value of at least a first parameter is not within a first range defined for the first parameter (such as greater than an upper threshold of the first range of the first parameter or less than a lower threshold of the first range of the first parameter), a first control flag is determined to be false. When the derived parameter values of the at least one parameter are within the respective ranges of the at least one parameter, the first control flag is determined to be true.
[0148] In some embodiments, the at least one parameter is derived based on a current template of a current block and a reference template of a reference image. When each of the derived parameter values of the at least one parameter is not within an associated range defined for the corresponding parameter (such as greater than an upper threshold of the associated range or less than a lower threshold of the associated range), the first control flag is derived to be false. When at least one of the derived parameter values of the at least one parameter is within the respective associated ranges of the at least one parameter, the first control flag is determined to be true.
[0149] In some examples, at least one threshold is a constant value, or the at least one threshold is encoded in at least one of the following in an encoded video bitstream: sequence header, slice header, picture header, and frame header.
[0150] In some examples, the threshold is determined based on at least one of the following: quantization parameter of the current block, block size of the current block, shape of the current block, temporal distance between the current block and a reference block.
[0151] In some embodiments, a first template type is determined according to a default template type associated with the shape of the current block. In some examples, when the shape of the current block is square and an upper template and a left template are available, the default template type is the upper and left template type (e.g., Figure 8 the upper and left template type (820) in
[0152] In some examples, when the shape of the current block is rectangular, the ratio of the height to the width of the current block is greater than a predefined threshold, and a left template is available, the default template type is the left template type (e.g., Figure 8 the left template type (840) in
[0153] In some examples, when the shape of the current block is rectangular, the ratio of the width to the height of the current block is greater than a predefined threshold, and an upper template is available, the default template type is the upper template type (e.g., Figure 8 the upper template type (830) in
[0154] In some embodiments, when the second control flag of the second block is true, the first template type is determined according to the second template type of the second block.
[0155] In some embodiments, when it is determined that the first control flag of the current block and the second control flag of the second block are the same flag, the first template type of the current block and the second template type of the second block are of the same type.
[0156] In some embodiments, it is determined that the first control flag of the current block and the second control flag of the second block are the same flag, regardless of whether the second block uses the same reference image as the first block for inter-frame prediction.
[0157] In some embodiments, first formula-based inter-frame prediction information is stored in association with the current block. The first formula-based inter-frame prediction information is for a formula-based inter-frame prediction technique that has been applied to the current block. The stored information can be used for encoding subsequent blocks. For example, at least one third block in the current image is encoded based on the first formula-based inter-frame prediction information. In the example, the formula-based inter-frame prediction information for the third block is determined according to the first formula-based inter-frame prediction information.
[0158] Note that the first formula-based inter-frame prediction information includes at least one of the following: the first control flag of the current block, the first template type of the current block, the first formula type of the current block, and / or the transform type of the residual data of the formula-based inter-frame prediction technique.
[0159] In some examples, the first formula-based inter-frame prediction information is stored in units of m×n samples, where m and n are positive integers.
[0160] Then, the method proceeds to (S1099) and terminates.
[0161] The method (1000) can be appropriately adjusted. At least one step in the method (1000) can be modified and / or omitted. At least one additional step can be added. Any suitable execution order can be used.
[0162] According to an aspect of an embodiment of the present application, a method for processing visual media data is provided. In this method, the bitstream of the visual media data is processed according to format rules. For example, the bitstream can be a bitstream decoded / encoded by any one of the decoding and / or encoding methods described herein. The format rules can specify at least one constraint for the bitstream and / or for at least one method performed by the decoder and / or encoder.
[0163] In an example, the bitstream includes encoded information of at least one image, and the at least one image includes a current block in a current image. The format rule specifies receiving a bitstream that includes first encoded information of a current block in the current image. The first encoded information indicates an inter-frame prediction of the current block based on a reference block in a reference image of the current block, and also indicates a possibility of applying a formula-based inter-frame prediction technique in the inter-frame prediction of the current block. The formula-based inter-frame prediction technique inputs at least one reconstructed sample in the reference block into a formula and generates predicted samples of the current block based on the formula. The formula includes at least one parameter derived based on a current template of the current block and a reference template of the reference block. In addition, the format rule specifies determining first formula-based inter-frame prediction information regarding a possibility of applying the formula-based inter-frame prediction technique to the current block based on at least one of the first encoded information of the current block and second encoded information of a second block reconstructed before the current block in the current image. The first formula-based inter-frame prediction information includes at least one of the following information regarding a possibility of applying the formula-based inter-frame prediction technique: a first control flag, a first template type, and a first formula type. In addition, the format rule specifies that when the first control flag is true, the formula-based inter-frame prediction technique is applied to the current block to generate at least one reconstructed sample of the current block.
[0164] The above techniques can be implemented as computer software using computer-readable instructions and physically stored in at least one computer-readable medium. For example, Figure 11 FIG. shows a computer system (1100) suitable for implementing certain aspects of the disclosed subject matter.
[0165] The computer software can be encoded using any suitable machine code or computer language that can be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0166] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.
[0167] Figure 11 The components shown for the computer system (1100) are examples and are not intended to impose any limitation on the scope of use or functionality of the computer software for implementing aspects of the embodiments of the present application. The configuration of the components should also not be construed as having any dependency or requirement related to any one or combination of the components illustrated in the example aspects of the computer system (1100).
[0168] A computer system (1100) may include certain human-machine interface input devices. Such human-machine interface input devices may respond to inputs generated by at least one human user through, for example, tactile inputs (such as keystrokes, swipes, data glove movements), audio inputs (such as speech, clapping), visual inputs (such as gestures), and olfactory inputs (not depicted). The human-machine interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (such as speech, music, ambient sounds), images (such as scanned images, photographic images obtained from a still-image camera), and video (such as two-dimensional video, three-dimensional video including stereoscopic video).
[0169] The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1101), mouse (1102), touchpad (1103), touch screen (1110), data glove (not shown), joystick (1105), microphone (1106), scanner (1107), camera (1108).
[0170] The computer system (1100) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen (1110), data glove (not shown), or joystick (1105), but there may also be tactile feedback devices that do not function as input devices), audio output devices (such as speakers (1109), headphones (not depicted)), visual output devices, and printers (not depicted). The visual output devices may be, for example, a screen (1110), virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted). The screen (1110) includes CRT screens, LCD screens, plasma screens, and OLED screens. Each screen may or may not have touch screen input capabilities, and each screen may or may not have tactile feedback capabilities. Some of these screens are capable of outputting two-dimensional visual output or more than three-dimensional output through, for example, stereoscopic output means.
[0171] The computer system (1100) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (1120) with CD / DVD or similar media (1121), thumb drives (1122), removable hard disk drives or solid-state drives (1123), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0172] Those skilled in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not cover a transmission medium, a carrier wave, or other transitory signals.
[0173] The computer system (1100) may also include an interface (1154) to at least one communication network (1155). The network can be, for example, wireless, wired, optical. The network can further be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, etc. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular and industrial networks including CANBus, etc. Certain networks typically require an external network interface adapter attached to certain common data ports or peripheral buses (1149) (e.g., the USB port of the computer system (1100)); other networks are typically integrated into the core of the computer system (1100) by attaching to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smartphone computer system). Using any of these networks, the computer system (1100) can communicate with other entities. Such communication can be unidirectional receive-only (e.g., broadcast TV), unidirectional transmit-only (e.g., CANbus to certain CANbus devices), or bidirectional (e.g., communicating with other computer systems using a local digital network or a wide area digital network). Certain protocols and protocol stacks can be used on each of the networks and network interfaces described above.
[0174] The above-described human-machine interface device, human-accessible storage device, and network interface can be attached to the core (1140) of the computer system (1100).
[0175] The kernel (1140) may include at least one central processing unit (CPU) (1141), a graphics processing unit (GPU) (1142), a dedicated programmable processing unit in the form of a field programmable gate array (FPGA) (1143), a hardware accelerator for certain tasks (1144), a graphics adapter (1150), etc. These devices, together with a read-only memory (ROM) (1145), a random access memory (1146), an internal mass storage device (such as an internal non-user-accessible hard disk drive), an SSD, etc. (1147), may be connected via a system bus (1148). In some computer systems, the system bus (1148) may be accessible in the form of at least one physical plug to enable expansion by attaching additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus (1148) of the kernel or attached to the system bus (1148) of the kernel via a peripheral bus (1149). In an example, a screen (1110) may be connected to the graphics adapter (1150). Architectures for peripheral buses include PCI, USB, etc.
[0176] The CPU (1141), GPU (1142), FPGA (1143), and accelerator (1144) may execute certain instructions that, when combined, may form the aforementioned computer code. The computer code may be stored in the ROM (1145) or the RAM (1146). Transitional data may also be stored in the RAM (1146), while permanent data may be stored, for example, in the internal mass storage device (1147). Fast storage and retrieval of any memory device may be achieved by using a cache memory that may be closely associated with at least one CPU (1141), GPU (1142), mass storage device (1147), ROM (1145), RAM (1146), etc.
[0177] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be media and computer code specially designed and constructed for the purposes of the embodiments of this application, or they may be of the type well-known and available to those skilled in the computer software art.
[0178] By way of example and not limitation, a computer system having an architecture (1100) (and in particular a kernel (1140)) can provide functionality due to at least one processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in at least one tangible, computer-readable medium. Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as introduced above, as well as certain storage devices of the kernel (1140) having a non-volatile nature, such as a kernel-internal mass storage device (1147) or a ROM (1145). Software implementing various aspects of the embodiments of the present application can be stored in such devices and executed by the kernel (1140). Depending on specific needs, the computer-readable medium can include at least one memory device or chip. The software can cause the kernel (1140) (and in particular the processors therein (including the CPU, GPU, FPGA, etc.)) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in the RAM (1146) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functionality through logic hard-wired or otherwise embodied in a circuit (e.g., an accelerator (1144)), which can operate instead of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. In appropriate cases, references to software can encompass logic and vice versa. In appropriate cases, references to a computer-readable medium can encompass a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit embodying logic for execution, or both. Embodiments of the present application encompass any suitable combination of hardware and software.
[0179] In embodiments of the present application, the use of "at least one of..." or "one of..." is intended to include any one or combination of the recited elements. For example, references to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). When applicable, such as when the elements are not mutually exclusive, the use of "one of..." does not exclude any combination of the recited elements.
[0180] Although embodiments of the present application have described several examples of aspects, there are changes, permutations, and various alternative equivalents that fall within the scope of embodiments of the present application. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods that, although not explicitly shown or described herein, embody the principles of embodiments of the present application and are thus within the spirit and scope of embodiments of the present application.
Claims
1. A video decoding method, characterized in that: include: Receive a code stream, the code stream comprising first encoded information of a current block in a current image; wherein the first encoded information indicates inter-frame prediction of the current block and possibility of applying a formula-based inter-frame prediction technique in the inter-frame prediction of the current block, the inter-frame prediction of the current block is based on a reference block in a reference image of the current block, the formula-based inter-frame prediction technique generates a prediction sample of the current block based on a formula by inputting at least one reconstructed sample in the reference block into the formula, the formula comprising at least one parameter derived based on a current template of the current block and a reference template of the reference block; Determine first formula-based inter-frame prediction information for the possibility of applying the formula-based inter-frame prediction technology to the current block based on at least one of the first encoded information of the current block and second encoded information of a second block reconstructed before the current block, the first formula-based inter-frame prediction information comprising at least one of the following: a first control flag regarding the possibility of application, a first template type for deriving the at least one parameter, and a first formula type of the formula; and When the first control flag is true, the formula-based inter-frame prediction technique is applied to the current block to generate at least one reconstructed sample of the current block.
2. The method according to claim 1, characterized in that Determining the first formula-based inter-frame prediction information includes: When the second encoded information of the second block indicates that the second control flag of the second block is true, thereby indicating that the formula-based inter-frame prediction technology is applied to the second block, the first control flag of the current block is determined based on the comparison of the value of the at least one parameter with a threshold.
3. The method according to claim 2, characterized in that Determining the first control flag further includes: deriving the at least one parameter based on the current template of the current block and the reference template of the reference image; When a first derived parameter value of at least a first parameter among the at least one parameter is greater than an upper threshold value of a first range of the first parameter or less than a lower threshold value of the first range of the first parameter, determining that the first control flag is false; and The first control flag is determined to be true when the derived parameter value of the at least one parameter is within a respective range of the at least one parameter.
4. The method according to claim 2, characterized in that: Determining the first control flag further includes: deriving the at least one parameter based on the current template of the current block and the reference template of the reference image; determining that the first control flag is false when each of the derived parameter values of the at least one parameter is greater than an upper threshold of an associated range or less than a lower threshold of the associated range; and The first control flag is determined to be true when at least one of the derived parameter values of the at least one parameter is within an associated range.
5. The method according to claim 2, characterized in that: The threshold is a constant value, or is decoded from at least one of a sequence header, a slice header, a picture header, and a frame header in the code stream.
6. The method according to claim 2, characterized in that The threshold is determined based on at least one of a quantization parameter of the current block, a block size of the current block, a shape of the current block, and a temporal distance between the current block and the reference block.
7. The method according to claim 1, characterized in that Determining the first formula-based inter-frame prediction information includes: The first template type is determined according to a default template type associated with a shape of the current block.
8. The method according to claim 7, characterized in that When the shape of the current block is a square, and an upper template and a left template are available, the default template type is an upper and left template type.
9. The method according to claim 7, characterized in that: When the shape of the current block is a rectangle, a ratio of the height to the width of the current block is greater than a predefined threshold, and a left template of the current block is available, the default template type is a left template type.
10. The method according to claim 7, characterized in that When the shape of the current block is a rectangle, a ratio of a width to a height of the current block is greater than a predefined threshold, and an upper template of the current block is available, the default template type is an upper template type.
11. The method according to claim 1, characterized in that: Determining the first formula-based inter-frame prediction information includes: When the second control flag of the second block is true, the first template type is determined according to the second template type of the second block.
12. The method according to claim 11, characterized in that Determining the first formula-based inter-frame prediction information includes: When it is determined that the first control flag of the current block and the second control flag of the second block are the same flag, it is determined that the first template type of the current block and the second template type of the second block are the same type.
13. The method according to claim 1, characterized in that Determining the first formula-based inter-frame prediction information includes: The first control flag of the current block and the second control flag of the second block are determined to be the same flag, regardless of whether the second block uses the reference image for inter-frame prediction.
14. The method according to claim 1, characterized in that Further including: storing, in association with the current block, the first formula-based inter-frame prediction information that has been applied to the current block; and At least one third block is reconstructed based on the first formula-based inter prediction information.
15. The method according to claim 14, characterized in that The first formula-based inter-frame prediction information includes at least one of the following: the first control flag of the current block; the first template type of the current block; The first formula type of the current block; and / or The transform type of the residual data of the formula-based inter-frame prediction technique.
16. The method according to claim 14, characterized in that The first formula-based inter-frame prediction information is stored in units of m×n samples, where m and n are positive integers.
17. A video encoding method, characterized in that: include: Based on at least one of first encoded information of a current block in a current image and second encoded information of a second block encoded before the current block, use a formula-based inter-frame prediction technique on the current block to determine first formula-based inter-frame prediction information of a first prediction candidate, the formula-based inter-frame prediction technique inputs at least one reconstructed sample of a reference block in a reference image into a formula, generates a prediction sample of the current block based on the formula, the formula includes at least one parameter derived based on a current template of the current block and a reference template of the reference block, the first formula-based inter-frame prediction information includes at least one of the following: a first control flag for the first prediction candidate, a first template type for deriving the at least one parameter, and a first formula type for the formula; When the first control flag is true, applying the formula-based inter-frame prediction technique to the current block to generate the first prediction candidate for the current block; and The current block is encoded based on the first prediction candidate.
18. The method according to claim 17, characterized in that Determining the first formula-based inter-frame prediction information includes: When the second encoded information of the second block indicates that the second control flag of the second block is true, the first control flag of the current block is determined based on a comparison of the value of the at least one parameter with a threshold.
19. The method according to claim 17, characterized in that Determining the first formula-based inter-frame prediction information includes: The first template type is determined according to a default template type associated with a shape of the current block.
20. A method for processing visual media data, characterized in that: The method comprises: Process the code stream of visual media data according to the format rules, where: The code stream includes first encoded information of a current block in a current image; wherein the first encoded information indicates inter-frame prediction of the current block and indicates the possibility of applying a formula-based inter-frame prediction technique in the inter-frame prediction of the current block, the inter-frame prediction of the current block is based on a reference block in a reference image of the current block, the formula-based inter-frame prediction technique generates a predicted sample of the current block based on a formula by inputting at least one reconstructed sample in the reference block into the formula, the formula including at least one parameter derived based on a current template of the current block and a reference template of the reference block; and The format rules specify: first formula-based inter-frame prediction information for determining the possibility of applying the formula-based inter-frame prediction technique to the current block based on at least one of the first encoded information of the current block and second encoded information of a second block reconstructed before the current block, the first formula-based inter-frame prediction information comprising at least one of: a first control flag regarding the possibility of application, a first template type for deriving the at least one parameter, and a first formula type of the formula; and When the first control flag is true, the formula-based inter-frame prediction technique is applied to the current block to generate at least one reconstructed sample of the current block.