Template-based prediction

By filtering and building linear models of the current block and reference block in video encoding, the problem of low video encoding efficiency in the existing technology is solved, and more efficient video compression and quality improvement is achieved.

CN120303943APending Publication Date: 2025-07-11TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480005096.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2024-07-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

When existing video encoding and decoding technologies compress video data, it is difficult to effectively utilize spatial and temporal redundancy, resulting in low encoding efficiency.

Method used

Using a template-based prediction method, filtering and linear model building of the templates of the current block and reference block is used to reconstruct the current block and improve the prediction accuracy.

Benefits of technology

It improves the efficiency and quality of video encoding, reduces the amount of data, and enhances the encoded video quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120303943A_ABST
    Figure CN120303943A_ABST
Patent Text Reader

Abstract

Aspects of the present disclosure include video decoding and video encoding methods and apparatuses, and a method of processing visual media data. And receiving the coded information in the code stream. The encoded information indicates whether filtering is to be applied to at least one of a current template of a current block and a reference template of a reference block in a current picture. The current block is predicted based on the reference block. When the encoded information indicates that filtering is to be applied, a plurality of samples within at least one of the current template and the reference template are filtered. Based on the filtered plurality of samples, a linear model between the current template and the reference template is determined. And reconstructing the current block based on the linear model and the reference block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Patent Application No. 18 / 769,339, titled "TEMPLATE BASED PREDICTION," filed on July 10, 2024, which claims priority to U.S. Provisional Patent Application No. 63 / 526,175, titled "METHOD AND APPARATUS FOR IMPROVEMENT ON TEMPLATE BASED PREDICTION," filed on July 11, 2023. The disclosures of these prior applications are hereby incorporated by reference in their entirety into this application. Technical Field

[0003] Aspects of the present disclosure generally relate to video coding and decoding. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. The work of the currently named inventors, to the extent it is described in the background art section, and aspects that may not be described in the background art and may not qualify as prior art at the time of the application, are neither expressly nor impliedly admitted as prior art against the present disclosure.

[0005] Image / video compression can help transmit image / video data across different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In an example, a video codec can use a technique called intra prediction, which can compress an image based on spatial redundancy. For example, intra prediction can use reference data from the current picture in reconstruction for sample prediction. In another example, a video codec can use a technique called inter prediction, which can compress an image based on temporal redundancy. For example, inter prediction can predict samples in the current picture from a previously reconstructed picture with motion compensation. Motion compensation can be indicated by a motion vector (MV). Summary of the Invention

[0006] Aspects of the present disclosure include video encoding / decoding methods and apparatuses.

[0007] According to one aspect of the present disclosure, a video decoding method includes: receiving encoded information in a bitstream. The encoded information indicates whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture. The current block is predicted based on a reference block of the current block. When the encoded information indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, the method includes: filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template. The method includes: reconstructing the current block based on the linear model and the reference block.

[0008] In an example, when the current block is predicted according to an IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture.

[0009] In an example, when the current block is predicted according to an inter prediction method, the reference block is in a reference picture different from the current picture.

[0010] The current template includes adjacent reconstructed samples of the current block. The reference template includes adjacent reconstructed samples of the reference block.

[0011] In an example, the method includes: applying the linear model to the reference block to determine a prediction signal, and reconstructing the current block based on the prediction signal. When the reference block is in the current picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block, a plurality of adjacent samples of the samples in the reference block, and a bias term. When the reference block is in the reference picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block and a bias term.

[0012] In an example, at least one of the current template and the reference template includes the current template and the reference template.

[0013] In an example, the method includes: filtering a first sample among the plurality of samples within the current template of the current block with a first filter, and filtering a second sample among the plurality of samples within the reference template of the reference block with a second filter different from the first filter.

[0014] In an example, at least one of the current template and the reference template consists of the current template or the reference template.

[0015] In an example, the method includes: using a filter to filter a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and the filter is

[0016] In an example, the encoded information in the bitstream includes a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at block level or a higher level than the block level.

[0017] In an example, filtering is applied to at least one of the current template and the reference template only when the current block is a luma block.

[0018] In an example, the current block is a luma block or a chroma block.

[0019] In an example, the current template includes at least one of the following: (i) a top template directly above the current block, a left template to the left of the current block, and an upper left template between the top template and the left template, and (ii) the top template, the left template, the upper left template between the top template and the left template, an upper right template above and to the right of the current block, and a lower left template below and to the left of the current block. The reference template has the same shape and the same size as the current template.

[0020] In an example, the plurality of samples is a subset of the samples within at least one of the current template of the current block and the reference template of the reference block.

[0021] In an example, the filter shape used for filtering depends on at least one of the position of the current template, the position of the reference template, the size of the current block, the shape of the current block, the size of the current template, and the shape of the current template.

[0022] In one aspect, a video coding method includes: when filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block in a current picture, filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template. The video coding method includes: encoding the current block based on the linear model and the reference block, and encoding a syntax element in the bitstream that indicates whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block.

[0023] In an example, when the current block is encoded according to the IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture. When the current block is encoded according to an inter prediction mode, the reference block is in a reference picture different from the current picture. The current template includes neighboring samples of the current block; and the reference template includes neighboring samples of the reference block.

[0024] In an example, at least one of the current template and the reference template includes the current template and the reference template.

[0025] In an example, the method includes: filtering a first sample among a plurality of samples within a current template of a current block using a first filter; and filtering a second sample among a plurality of samples within a reference template of a reference block using a second filter different from the first filter.

[0026] In an example, at least one of the current template and the reference template consists of the current template or the reference template.

[0027] In an example, a syntax element is a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at a block level or a higher level than the block level.

[0028] In one aspect, a method of processing visual media data includes: processing a bitstream of visual media data according to formatting rules. The bitstream includes a syntax element indicating whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture, and the current block is predicted based on a reference block of the current block. The formatting rules specify that when encoded information indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, filtering is performed on a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and a linear model between the current template and the reference template is determined based on the filtered plurality of samples within at least one of the current template and the reference template. The formatting rules provide that the current block is reconstructed based on the linear model and the reference block.

[0029] Aspects of the present disclosure also provide a video encoding apparatus. The video encoding apparatus includes processing circuitry configured to implement any of the described video encoding methods.

[0030] Aspects of the present disclosure also provide a video decoding apparatus. The video decoding apparatus includes processing circuitry configured to implement any of the described video encoding methods.

[0031] Aspects of the present disclosure also provide a non - volatile computer - readable medium storing instructions that, when executed by a computer, cause the computer to perform any of the described video decoding / encoding methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0033] Figure 1 is a schematic diagram of an example of a block diagram of a communication system (100).

[0034] Figure 2 is a schematic diagram of an example of a block diagram of a decoder.

[0035] Figure 3 It is a schematic diagram of an example of a block diagram of an encoder.

[0036] Figure 4 An example of performing intra prediction using reconstructed samples in a template according to one aspect of the present disclosure is shown.

[0037] Figure 5 An example of an intra-template matching prediction (IntraTMP) mode according to one aspect of the present disclosure is shown.

[0038] Figure 6A An example of constructing a linear model using a current template of a current block and a reference template of a reference block according to one aspect of the present disclosure is shown.

[0039] Figure 6B An example of applying a linear model to a reference block to determine a prediction signal of a current block according to one aspect of the present disclosure is shown.

[0040] Figure 7 An example of predicting a current block based on a reference block of the current block in a current picture according to one aspect of the present disclosure is shown.

[0041] Figure 8 An example of the shape of a 3×3 filter F according to one aspect of the present disclosure is shown.

[0042] Figure 9 An example of an extended current template according to one aspect of the present disclosure is shown.

[0043] Figure 10 An example of a current template according to one aspect of the present disclosure, where the filter is not applied to two samples at the ends of the middle row / column.

[0044] Figure 11 An example of when filtering two rows of samples of a current template according to one aspect of the present disclosure is shown.

[0045] Figure 12 A flowchart outlining a decoding process according to some aspects of the present disclosure is shown.

[0046] Figure 13 A flowchart outlining an encoding process according to some aspects of the present disclosure is shown.

[0047] Figure 14 It is a schematic diagram of a computer system according to one aspect. Detailed Description

[0048] Figure 1FIG. 0 shows a block diagram of a video processing system (100) in some examples. The video processing system (100) is an example of an application, a video encoder, and a video decoder of the subject matter disclosed in this application in a streaming environment. The subject matter disclosed in this application is equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0049] The video processing system (100) includes an acquisition subsystem (113), and the acquisition subsystem may include a video source (101) such as a digital camera, and the video source creates an uncompressed video picture stream (102). In an embodiment, the video picture stream (102) includes samples taken by a digital camera. Compared with the encoded video data (104) (or encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high-data-volume video picture stream. The video picture stream (102) can be processed by an electronic device (120), and the electronic device (120) includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the subject matter disclosed in more detail below. Compared with the video picture stream (102), the encoded video data (104) (or encoded video bitstream) is depicted as a thin line to emphasize the lower-data-volume encoded video data (104) (or encoded video bitstream), which can be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 the client subsystem (106) and the client subsystem (108) in, can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of the encoded video data and generates an output video picture stream (111) that can be presented on a display (112) (such as a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (such as a video bitstream) may be encoded according to certain video coding / compression standards. Embodiments of such standards include ITU-T H.265. In an embodiment, a video coding standard under development is informally referred to as Versatile Video Coding (VVC), and this application can be used in the context of the VVC standard.

[0050] Note that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may further include a video encoder (not shown).

[0051] Figure 2 is an example block diagram of a video decoder (210). The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 the video decoder (110) in the embodiment.

[0052] The receiver (231) may receive one or more encoded video sequences included in, for example, a bitstream to be decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, where the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (231) may receive the encoded video data and other data, e.g., encoded audio data and / or auxiliary data streams that may be forwarded to their respective using entities (not labeled). The receiver (231) may separate the encoded video sequences from the other data. To prevent network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as "parser (220)"). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) may be provided external to the video decoder (210) (not labeled). In yet other cases, a buffer memory (not labeled) is provided external to the video decoder (210) to, for example, prevent network jitter, and another buffer memory (215) may be configured inside the video decoder (210) to, for example, handle the playback timing. And when the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may also be possible not to configure the buffer memory (215), or the buffer memory may be made smaller. Of course, for use on a service packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may have an adaptive size, and may be implemented at least partially in an operating system or a similar element (not labeled) external to the video decoder (210).

[0053] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from an encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (210), and potential information for controlling a display device (212) (e.g., a display screen), such as a display device that is not part of the electronic device (230) but can be coupled to the electronic device (230), as Figure 2 shown. The control information for the display device can be a parameter set segment (not labeled) of Supplemental Enhancement Information (SEI message) or Video Usability Information (VUI). The parser (220) may perform parsing / entropy decoding on the received encoded video sequence. The encoding of the encoded video sequence can be performed according to video coding techniques or standards and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so on. The parser (220) may extract subgroup parameter sets for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the subgroup. The subgroup may include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), and so on. The parser (220) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on.

[0054] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).

[0055] Depending on the type of the encoded video picture or a part of the encoded video picture (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbols (221) may involve multiple different units. Which units are involved and the way they are involved can be controlled by subgroup control information parsed by the parser (220) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (220) and multiple units below are not described.

[0056] In addition to the functional blocks already mentioned, the video decoder (210) can conceptually be subdivided into several functional units as described below. In practical embodiments operating under commercial constraints, many of these units interact closely with each other and can be integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the functional units below.

[0057] The first unit is the scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantized transform coefficients as symbols (221) and control information from the parser (220), including which transform mode to use, block size, quantization factor, quantization scaling matrix, etc. The scaler / inverse transform unit (251) can output blocks including sample values, and the sample values can be input into the aggregator (255).

[0058] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to intra-coded blocks; that is, blocks that do not use predictive information from previously reconstructed pictures but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (255) adds the predictive information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) based on each sample.

[0059] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to inter-coded and potentially motion-compensated blocks. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to extract samples for prediction. After motion-compensating the extracted samples according to the symbol (221), these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251) (which is called the residual sample or residual signal in this case), thereby generating output sample information. The motion compensation prediction unit (253) obtaining the prediction samples from the addresses within the reference picture memory (257) can be controlled by motion vectors, and the motion vectors are in the form of the symbol (421) for use by the motion compensation prediction unit (253), and the symbol (221) includes, for example, X, Y, and reference picture components. Motion compensation can also include interpolation of sample values extracted from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.

[0060] The output samples of the aggregator (255) can be adopted by various loop filtering techniques in the loop filter unit (256). Video compression techniques can include in-loop filter techniques, which are controlled by parameters included in an encoded video sequence (also referred to as an encoded video bitstream), and the parameters can be used for the loop filter unit (256) as symbols (221) from the parser (220). However, in other embodiments, the video compression techniques can also respond to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or an encoded video sequence, and respond to previously reconstructed and loop-filtered sample values.

[0061] The output of the loop filter unit (256) can be a sample stream, which can be output to the display device (212) and stored in the reference picture memory (257) for subsequent inter-picture prediction.

[0062] Once fully reconstructed, some encoded pictures can be used as reference pictures for future prediction. For example, once the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (by, for example, the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before starting to reconstruct subsequent encoded pictures.

[0063] The video decoder (210) can perform decoding operations according to, for example, a predetermined video compression technique in the ITU-T H.265 standard. In the sense that the encoded video sequence conforms to the syntax specified by the video compression technique or standard and the profile recorded in the video compression technique or standard, the encoded video sequence can conform to the syntax of the video compression technique or standard used. Specifically, the profile can select certain tools from all the tools available in the video compression technique or standard as the only tools available under the profile. For compliance, it is also required that the complexity of the encoded video sequence be within the range defined by the level of the video compression technique or standard. In some cases, the level limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the level can be further defined by the Hypothetical Reference Decoder (HRD) specification and the metadata of the HRD buffer management signaled in the encoded video sequence.

[0064] In one aspect, a receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of an encoded video sequence. The additional data may be used by a video decoder (210) to decode the data appropriately and / or reconstruct the original video data more accurately. The additional data may be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, and the like.

[0065] Figure 3 FIG. is an example block diagram of a video encoder (303). The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) may be used to replace Figure 1 the video encoder (103) in the embodiment.

[0066] The video encoder (303) may receive video samples from a video source (301) (which is not Figure 3 part of the electronic device (320) in the embodiment), and the video source may capture video images to be encoded by the video encoder (303). In another embodiment, the video source (301) is part of the electronic device (320).

[0067] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), and the digital video sample stream may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits...), any color space (e.g., BT.601 Y CrCb, RGB...), and any suitable sampling structure (e.g., Y CrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing previously prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of individual pictures that are given motion when viewed in sequence. The pictures themselves may be constructed as spatial pixel arrays, where each pixel may include one or more samples depending on the sampling structure, color space, etc. used. The following focuses on the description of samples.

[0068] According to one aspect, a video encoder (303) may encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed is a function of a controller (350). In some aspects, the controller (350) controls other functional units as described below and is functionally coupled to these units. For the sake of brevity, the couplings are not labeled in the figures. Parameters set by the controller (350) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (350) may be used for other suitable functions that relate to the video encoder (303) optimized for a certain system design.

[0069] In some aspects, the video encoder (303) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols in a manner similar to how a (remote) decoder creates sample data to create sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream produces a bit-exact result independent of the decoder location (local or remote), the content in the reference picture memory (334) is also bit-exact corresponding between the local encoder and the remote encoder. In other words, the reference picture samples “seen” by the prediction part of the encoder are exactly the same as the sample values that the decoder will “see” when using the prediction during decoding. This reference picture synchronization principle (and the drift that occurs, for example, when the synchronization cannot be maintained due to channel errors) is also used in some related technologies.

[0070] The operation of the “local” decoder (333) may be the same as that of the “remote” decoder that has been described in detail above in connection with Figure 2 the video decoder (210). However, briefly referring additionally to Figure 2 , when symbols are available and the entropy encoder (345) and the parser (220) can encode / decode the symbols losslessly into the encoded video sequence, the entropy decoding part of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).

[0071] In one aspect, any decoder technology other than parsing / entropy decoding present in the decoder also exists in the corresponding encoder in the same or substantially the same functional form. Accordingly, the present application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is reciprocal to the decoder technology described comprehensively. A more detailed description in certain areas is provided below.

[0072] During operation, in some embodiments, the source encoder (330) may perform motion compensated predictive coding. Referencing one or more previously encoded pictures designated as "reference pictures" in the video sequence, the motion compensated predictive coding performs predictive coding on the input picture. In this way, the coding engine (332) encodes the difference between the pixel blocks of the input picture and the pixel blocks of the reference picture, and the reference picture can be selected as the prediction reference for the input picture.

[0073] The local video decoder (333) may decode the encoded video data of the picture that can be designated as a reference picture based on the symbols created by the source encoder (330). The operation of the coding engine (332) may be a lossy process. When the encoded video data can be decoded at a video decoder ( Figure 3 not shown), the reconstructed video sequence is typically a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture memory (334). In this way, the video encoder (303) can locally store a copy of the reconstructed reference picture that has the same content (in the absence of transmission errors) as the reconstructed reference picture that will be obtained by the remote video decoder.

[0074] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be encoded, the predictor (335) may search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or some metadata, such as reference picture motion vectors, block shapes, etc., that can serve as an appropriate prediction reference for the new picture. The predictor (335) may operate on a per-pixel block basis of the sample blocks to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input picture may have a prediction reference taken from multiple reference pictures stored in the reference picture memory (334).

[0075] The controller (350) may manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.

[0076] The outputs of all the above functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by the various functional units according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc., thereby converting the symbols into an encoded video sequence.

[0077] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0078] The controller (350) can manage the operation of the video encoder (303). During encoding, the controller (350) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types:

[0079] An intra picture (I picture), which can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.

[0080] A predictive picture (P picture), which can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses motion vectors and reference indices to predict the sample values of each block.

[0081] A bi - predictive picture (B picture), which can be encoded and decoded using intra prediction or inter prediction, where the intra prediction or inter prediction uses two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.

[0082] Source pictures can typically be spatially subdivided into multiple sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, which are determined according to the coding assignment of the corresponding picture applied to the block. For example, blocks of an I picture can be non-predictively encoded, or the blocks can be predictively encoded with reference to already encoded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P picture can be predictively encoded by spatial prediction or by temporal prediction with reference to a previously encoded reference picture. Blocks of a B picture can be predictively encoded by spatial prediction or by temporal prediction with reference to one or two previously encoded reference pictures.

[0083] The video encoder (303) can perform encoding operations according to a predetermined video coding technique or standard such as the ITU-T H.265 recommendation. In operation, the video encoder (303) can perform various compression operations, including predictive coding operations that exploit the temporal and spatial redundancies in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video coding technique or standard used.

[0084] In one aspect, the transmitter (340) can transmit additional data when transmitting the encoded video. The source encoder (330) can include such data as part of the encoded video sequence. The additional data can include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0085] The captured video can be a plurality of source pictures (video pictures) in a time series. Intra picture prediction (often abbreviated to intra prediction) exploits the spatial correlation within a given picture, while inter picture prediction exploits the (temporal or other) correlation between pictures. In an embodiment, the particular picture being encoded / decoded is segmented into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that has been previously encoded and is still buffered in the video, the block in the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block in the reference picture, and in the case of using multiple reference pictures, the motion vector can have a third dimension identifying the reference picture.

[0086] In one aspect, bidirectional prediction techniques can be used in inter - picture prediction. According to the bidirectional prediction technique, two reference pictures are used, for example, a first reference picture and a second reference picture that are both before the current picture in the decoding order (but may be past and future respectively in the display order) in the video. A block in the current picture can be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.

[0087] In addition, merge mode techniques can be used in inter - picture prediction to improve coding efficiency.

[0088] According to some aspects disclosed in the present application, predictions such as inter - picture prediction and intra - picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are segmented into coding tree units (CTUs) for compression. CTUs in a picture have the same size, such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), which are one luminance CTB and two chrominance CTBs. Further, each CTU can be split into one or more coding units (CUs) in a quadtree. For example, a 64×64 - pixel CTU can be split into a 64×64 - pixel CU, or 4 32×32 - pixel CUs, or 16 16×16 - pixel CUs. In an embodiment, each CU is analyzed to determine the prediction type for the CU, such as an inter - frame prediction type or an intra - frame prediction type. In addition, depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (PUs). Generally, each PU includes a luminance prediction block (PB) and two chrominance PBs. In one aspect, prediction operations in encoding (encoding / decoding) are performed in units of prediction blocks. Taking the luminance prediction block as the prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luminance values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0089] Note that any suitable technology may be used to implement video encoder (103) and video encoder (303), as well as video decoder (110) and video decoder (210). In one aspect, one or more integrated circuits may be used to implement video encoder (103) and video encoder (303), as well as video decoder (110) and video decoder (210). In another aspect, one or more processors executing software instructions may be used to implement video encoder (103) and video encoder (303), as well as video decoder (110) and video decoder (210).

[0090] Aspects of the present disclosure describe template-based prediction methods, including improvements to template-based prediction methods. In the present disclosure, a set of methods for video compression and / or image compression are described, including intra prediction mode coding and inter prediction mode coding.

[0091] Video coding has been widely used in many applications. Various video coding standards such as H264, H265, H266 (or VVC), AV1, and AVS have been widely adopted. In one aspect, a video codec may include multiple modules, including intra prediction / inter prediction, transform coding, quantization, entropy coding, loop filtering, etc. Intra prediction may be one of the main modules and may include signaling processing methods (e.g., signaling processing methods) and neural network-based methods.

[0092] The current coding block (also interchangeably referred to as the current block) and adjacent samples of the current block may share one or more similar texture characteristics. The template of the current block may include adjacent samples of the current block (e.g., adjacent reconstructed samples). The template of the current block may be interchangeably referred to as the current template of the current block. The current template of the current block may be used to predict the current block or improve the prediction signal of the current block.

[0093] The current template may be used to predict the current block. Figure 4 An example of performing intra prediction using reconstructed samples within a template according to one aspect of the present disclosure is shown. In Figure 4 , the current block (431) in the current picture (also interchangeably referred to as the current frame) (410) is being encoded, e.g., the current block (431) is being reconstructed. The current picture (410) may include a region (411) (gray). In the example, the region (411) has been reconstructed and is referred to as the reconstructed region. The current template (401) of the current block (431) may include adjacent samples of the current block (431). In the example, the current template (401) may be used to determine (e.g., find) the best-matched reconstructed block of the current block (431) in the current picture (410).

[0094] By comparing the current template (401) of the current block (431) with the template of the corresponding reference block (e.g., the reference template), for example, if the error between the reference template (402) and the current template (401) (e.g., a cost function based on the sum of absolute differences (SAD)) is the minimum among the errors between each reference template and the current template (401), then the best - matched reference template (e.g., Figure 4 the reference template (402) shown in Figure 4 may be found in the reconstructed region (411). After that, the reference block with the minimum error (e.g.,

[0095] the reference block (432) shown in

[0096] can be used to predict the current block (431). In an example, the reference template (402) and the current template (401) may have the same shape and the same size. The position of the reference template (402) relative to the reference block (432) may be the same as the position of the current template (401) relative to the current block (431).

[0095] In an example, the values of the reference samples in the reference block (432) are copied as the prediction signal for the current block (431). For example, the reference block (432) is the prediction signal, and the predicted values of the samples in the current block (431) are respectively the values of the reference samples in the reference block (432).

[0096] In an example, Figure 4 the method described in Figure 5 is the Intra - Template Matching Prediction (IntraTMP) mode. In an example, the IntraTMP mode is a special intra - prediction mode for intra - prediction. Figure 5 shows an example of the IntraTMP mode according to one aspect of the present disclosure. Referring to Figure 5 , in an example of the IntraTMP mode, the prediction block (521), such as the best - prediction block from the reconstructed part of the current picture (or current frame), can be copied. The template (520) (such as the L - shaped template of the prediction block (521)) can match the current template (530) of the current block (531) in the current picture. For a predefined search range, the encoder can search for the template most similar to the current template (530) in the reconstructed part of the current frame and can use the corresponding block (521) as the prediction block (also called the reference block). In an example, the encoder then signals to use the IntraTMP mode, and the same prediction operation is performed on the decoder side.

[0097] Referring to Figure 5, a prediction signal can be generated by matching the L-shaped template (530) of the current block (531) with another template in a predefined search area. In an example, the predefined search area includes: R1, R2, R3, and R4, where R1 is the current CTU, R2 is the upper left CTU of the current CTU, R3 is the upper CTU of the current CTU, and R4 is the left CTU of the current CTU. In an example, SAD is used as the cost function.

[0098] Within each region, the decoder can search for the template with the minimum cost (e.g., minimum SAD) relative to the current template and can use the block corresponding to the minimum cost (e.g., the reference block) as the prediction block.

[0099] The dimensions (SearchRange_w, SearchRange_h) of all regions can be set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. In an example, SearchRange_w = a × BlkW, SearchRange_h = a × BlkH, where "a" is a constant that controls the trade-off between gain and complexity. In an example, "a" is equal to 5.

[0100] To accelerate the template matching process, in some examples, the search ranges of all search regions are subsampled by a factor of 2, for example, which results in a 4-fold reduction in the template matching search. After finding the best match, a refinement process can be performed. The refinement process can be performed by performing a second template matching search with a narrowed range around the best match. In an example, the narrowed range is defined as min(BlkW, BlkH) / 2.

[0101] The Intra Block Copy (IBC) mode can be used to predict a block in the current picture based on a reference block (e.g., a reconstructed block) in the current picture. The IBC mode can be implemented as a block-level coding mode, and block matching (BM) can be performed at the encoder to find the optimal block vector (BV) for each CU. In the IBC mode, the BV can be used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture.

[0102] On the encoder side, in an example, hash-based motion estimation can be performed for the IBC mode. The encoder can perform rate-distortion (RD) checks on blocks with a width or height not greater than 16 luminance samples. For non-merging modes, block vector search can be first performed using hash-based search. If the hash search does not return a valid candidate, local search based on block matching can be performed.

[0103] In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block can be extended to all allowed block sizes. In the current picture, the hash key calculation for each position can be based on 4×4 sub-blocks. For a current block of a larger size, when all the hash keys of all 4×4 sub-blocks match the hash keys at the corresponding reference positions, it can be determined that the hash key of the current block matches the hash key of the reference block. If multiple reference block hash keys are found to match the current block hash key, the block vector cost for each matching reference can be calculated, and the matching reference with the minimum cost can be selected.

[0104] In block matching search, the search range can be set to include both the previous CTU and the current CTU (e.g., the previously reconstructed CTU and the current CTU).

[0105] In one aspect, for the return reference Figure 4 , a linear model can be constructed between the current template and the reference template (such as between the current template (401) and the reference template (402)). After modeling (e.g., after determining the linear model), the reference samples in the reference block (432) can be used as the input of the linear model to generate the predicted output of the current block (431). In the example, the current block (431) can be predicted based on the predicted output instead of directly based on the reference block (432).

[0106] Figure 6A An example of constructing a linear model using the current template (401) and the reference template (402) is shown according to one aspect of the present disclosure. In Figure 6A it is redrawn Figure 4 the current template (401), the reference template (402), the current block (431), the reference block (432), and the current picture (410) described in

[0107] In Figure 6A the example, the template sizes of the current template (401) and the reference template (402) have four rows (e.g., four rows and / or four columns) of samples, and the current template (401) and the reference template (402) have an "L" shape. The description of this example can be applicable to or can be appropriately adapted to any template size (e.g., greater than 4 rows or less than 4 rows (e.g., 3 rows)) and any suitable shape (e.g., "L" shape, only left template, only top template, etc.) of the current template (401) and the reference template (402).

[0108] In one aspect, the linear model shown in equation (1) can be used to represent (e.g., predict) each sample in the current template (401) (e.g., Figure 6A the sample (610) shown in

[0109] predVal = c0 × valC + c1 × valN + c2 × valS + c3 × valE + c4 × valW + c5B Equation (1)

[0110] predVal can be the predicted value of a sample (610) within the current template (401). valC, valN, valS, valE, and valW can be the sample values of the corresponding samples C, N, S, E, and W in the reference template (402). Refer to Figure 6A , C is the central (C) sample. C can correspond to the sample (610) to be predicted in the current template (401). In an example, the relative position of C in the reference template (402) is the same as the relative position of the sample (610) in the current template (401). N, S, E, and W are the upper neighbor / northern neighbor (N) of C, the lower neighbor / southern neighbor (S) of C, the right neighbor / eastern neighbor (E) of C, and the left neighbor / western neighbor (W) of C. c0 to c5 are the coefficients (also interchangeably referred to as parameters) of the linear model. Coefficients c0 to c4 correspond to samples C, N, S, E, and W. B is the bias term. Coefficient c5 corresponds to the bias term B.

[0111] The reference template (402) can be surrounded by samples in the region (620) (gray shaded). Refer to Figure 6A , the region (620) surrounds the reference template (402). In an example, when one of the samples N, S, E, and W is outside the reference template (402) (e.g., one of the samples N, S, E, and W is in the region (620)), one of the samples of N, S, E, and W in the region (620) can be used as an input sample (e.g., a special dependent input sample) of the linear model.

[0112] In an example, the linear model described in Equation (1) is a filter (e.g., a linear filter) defined based on samples C, N, S, E, and W. In Figure 6A the example shown, the filter shape of the linear model is a cross.

[0113] To learn (e.g., determine) the parameters (e.g., c0 to c5) of the linear model, methods such as LDL decomposition, Cholesky decomposition, variants of LDL decomposition, etc. can be used. In an example of LDL decomposition, the matrix A can be decomposed into A = LDL T . L is a lower unit triangular (single triangular) matrix, D is a diagonal matrix, and L T is the transpose of L. In an example, coefficients c0 to c5 can be derived based on minimizing the difference between the values of the reconstructed samples within the current template (401) and the predicted values of the reconstructed samples within the current template (401) (e.g., using Equation (1)) (e.g., via a regression-based minimization technique).

[0114] After determining the parameters of the linear model (e.g., c0 to c5), a predicted signal for the current block (431) can be derived based on the sample values of the reference block (432) and the linear model. Figure 6B An example is shown of applying a linear model with determined parameters to a reference block (432) to determine a predicted signal for the current block (431). In Figure 6B is redrawn in Figure 4 and / or Figure 6A the current template (401), reference template (402), current block (431), reference block (432), current picture (410), and region (620) described in

[0115] In one aspect, a linear model such as that shown in Equation (2) can be used to determine a predicted value predVal' for samples (e.g., Figure 6B the samples (650) shown in

[0116] predVal' = c0 × valC' + c1 × valN' + c2 × valS' + c3 × valE' + c4 × valW' + c5B Equation (2)

[0117] valC', valN', valS', valE', and valW' can be the sample values of the corresponding samples C', N', S', E', and W' in the reference block (432). Refer to Figure 6B , C' is the center (C') sample. C' can correspond to the sample (650) to be predicted in the current block (431). In the example, the relative position of C' in the reference block (432) is the same as the relative position of the sample (650) in the current block (431). N', S', E', and W' are the upper neighbor / north neighbor (N') of C', the lower neighbor / south neighbor (S') of C', the right neighbor / east neighbor (E') of C', and the left neighbor / west neighbor (W') of C'. c0 to c5 are the same as the coefficients of the linear model determined above with reference to Figure 6A

[0118] In one aspect, refer to Figures 6A to 6B , when determining the reference block (432) using the IntraTMP mode or IBC mode, Figures 6A to 6B the method described in Figures 6A to 6B ​As shown, up to 4 rows / 4 columns of samples above and to the left of the current block (431) can be applied to derive coefficients (e.g., c0 to c5).

[0119] In the method described in the reference Figures 6A to 6B the correlation between the current template (401) and the reference template (402) is determined and can be indicated by a linear model. The current template (401) and the reference template (402) are respectively close to (e.g., spatially close to) the current block (431) and the reference block (432). Thus, the learned correlation can be applied to the current block (431) and the reference block (432). In an example, from the perspective of machine learning, the construction of a linear model such as Figure 6A and as shown in Equation (1) can be referred to as the training phase, and the prediction of the current block (431) such as Figure 6B and as shown in Equation (2) can be referred to as the inference phase. Figures 6A to 6B The method described in

[0120] can be referred to as a training-inference method. Figures 6A to 6B The training-inference method (e.g., described in

[0121] and the FIBC mode) can be used to predict the current block (e.g., the current block (431)) within the same picture (or the same frame) as the reference block (e.g., the reference block (432)), and thus is (intra-frame) prediction. Figures 6A to 6B In some examples, the training-inference method using templates such as the current template and the reference template is not limited to intra-frame prediction, but can be applied to inter-frame prediction and improve inter-frame prediction. In inter-frame prediction, the reference template and the current template can be located in different pictures. In inter-frame prediction, the prediction signal of the current block (e.g., the inter-frame prediction signal) can be generated by any suitable inter-frame prediction method (e.g., using motion compensation), because the inter-frame prediction method can be an effective way to utilize temporal redundancy. The template can be used to compensate the prediction signal of the current block (e.g., the inter-frame prediction signal) using the training-inference method. In the case of inter-frame prediction, a linear model can be determined based on the reference template and the current template located in different pictures, and then the linear model can be applied to update (e.g., correct) the prediction signal. The current block can be predicted based on the updated prediction signal. The linear model in inter-frame prediction can be the same as or different from the linear model in Figures 6A to 6B In an example, the linear model in inter-frame prediction is different from the linear model in

[0122] In one aspect, the linear model shown in Equation (3) can be used to predict (or correct) each sample in the current template based on the sample value valC” of sample C” within the reference template.

[0123] predVal” = c6 × valC” + c7B” Equation (3)

[0124] predVal” can be the predicted value of the sample in the current template. In an example, the relative position of C” in the reference template is the same as the relative position of the sample in the current template. c6 and c7 are coefficients (or parameters) of the linear model. Coefficient c6 corresponds to sample C”. B” is the bias term. Coefficient c7 corresponds to the bias term B”.

[0125] Similar to Figures 6A to 6B As described in, after determining the linear model parameters (e.g., c6 and c7) in Equation (3), the predicted signal of the current block can be derived based on the sample values in the reference block and the linear model in Equation (3). In some examples, if the reference block is considered as the predicted signal, the predicted signal can be considered as the updated predicted signal. In one aspect, a linear model such as that shown in Equation (4) can be used to determine the predicted value predVal’” of the samples in the current block based on the sample value valC’” of sample C” in the reference block.

[0126] predVal’” = c6 × valC’” + c7B” Equation (4)

[0127] A common feature of the above-described template methods (such as those described in Equations (1) and (3)) is that the samples (e.g., the reconstructed samples) in the current template and the reference template (e.g., current template (401) and reference template (402)) used to determine the linear model (e.g., the linear model in Equations (1) to (2) or the linear model in Equations (3) to (4)) are directly used (e.g., without modification). In some examples, directly using the reconstructed samples in the current template and the reference template may not result in an optimal model (e.g., an optimal linear model) that fully utilizes the correlation between adjacent blocks. In some examples, a modified version of the current template and / or a modified version of the reference template can be used to represent the correlation in a more accurate manner.

[0128] The methods, aspects, and examples described in this disclosure can be used alone or in any combination. The term “IBC mode” can refer to the IBC mode or variants described in this disclosure. The term “IntraTMP mode” can refer to the IntraTMP mode or variants described in this disclosure. The term “FIBC mode” can refer to the FIBC mode or variants described in this disclosure.

[0129] The current block in the current picture can be predicted based on a reference block. In one aspect, the reference block can be in the current picture and can be determined, for example, using the method shown in Figure 4 In some examples, the reference block is in the current picture and is determined using the IntraTMP mode. In some examples, the reference block is in the current picture and is determined using the IBC mode. In one aspect, the reference block can be determined using inter prediction and can be in a reference picture different from the current picture.

[0130] According to one aspect of the present disclosure, the reference block is not directly copied to predict the current block. In one aspect, samples in the reference block (also referred to as reference samples) can be input into a linear model to generate a prediction signal for the current block. Thus, the linear model is applied to the reference block to determine the prediction signal for the current block. Equations (1) to (4) show some examples of the linear model. Examples of applying the linear model to generate the prediction signal for the current block are described using Equations (2) and (4).

[0131] In various examples, a block and a template of the block (e.g., including neighboring samples of the block) can share similar texture characteristics, e.g., because the block and the template are spatially close to each other. Thus, the linear model can be determined based on the reference template of the reference block and the current template of the current block.

[0132] Equations (1) to (4) show some examples of the linear model. The linear model can be any suitable linear model and can have any suitable shape and / or any suitable size (e.g., 5 samples, more than 5 samples, or less than 5 samples). Equations (1) to (2) and Figures 6A to 6B show an example of a cross shape. Refer to Figure 6B , the cross shape includes 5 samples C’, N’, S’, E’, and W’. Sample C’ is the central sample corresponding to the sample (650) being predicted in the current block (431). Other linear models can be used with five samples including C' and four other samples different from N’, S’, E’, and W’. Another linear model can include less than five samples or more than five samples. In the example shown in Equation (4), the linear model includes only sample C’’’ and a bias term (e.g., offset).

[0133] According to one aspect of the present disclosure, the current template and / or the reference template can be filtered (e.g., modified) before determining the linear model. In some examples, the filtered current template and / or the filtered reference template can be used to determine the linear model, and the linear model can be applied to the reference block to predict the current block.

[0134] Figure 7An example of predicting a current block (831) in a current picture (811) based on a reference block (832) of the current block (831) is shown. The reference block (832) is in a reference picture (812). According to one aspect of the present disclosure, a plurality of samples within at least one of a current template (801) of the current block (831) and a reference template (802) of the reference block (832) may be filtered. A linear model between the current template (801) and the reference template (802) may be determined based on the current template (801) and the reference template (802), wherein the plurality of samples within at least one of the current template (801) and the reference template (802) are filtered. For example, a linear model between the current template (801) and the reference template (802) may be determined based on the filtered plurality of samples within at least one of the current template (801) and the reference template (802). The current block (831) may be reconstructed based on the linear model and the reference block (832).

[0135] In an example, the reference picture (812) is the current picture (811), and the reference block (832) is in the current picture (811). In the example, the current block (831) is predicted according to the method described in Figure 4 In the example, the current block (831) is predicted according to the IntraTMP mode, such as Figure 5 as described in. In the example, the current block (831) is predicted according to the IBC mode.

[0136] In an example, the current picture (811) is different from the reference picture (812), and the reference block (832) is in a reference picture (812) different from the current picture (811). The current block (831) may be predicted according to any suitable inter prediction method.

[0137] In one aspect, the current template (801) may include adjacent samples (e.g., adjacent reconstructed samples) of the current block (831). The reference template (802) may include adjacent samples (e.g., adjacent reconstructed samples) of the reference block (832). The current template (801) and the reference template (802) may have any suitable shape and any suitable size. The current template (801) and the reference template (802) may have the same shape and may have the same size. In Figure 8 the example shown, the current template (801) and the reference template (802) include 4 rows of reconstructed samples adjacent to the current block (831) and the reference block (832), respectively. In Figure 8 the example shown, the current template (801) and the reference template (802) have an L shape.

[0138] In one aspect, it may be similar to Figure 6ADetermine the linear model between the current template (801) and the reference template (802) in the manner described, except that multiple samples within at least one of the current template (801) and the reference template (802) are filtered. In the example, the reference block (832) is in the current picture (811), and the current block (831) is predicted according to Figures 4 to 5 the method described. The linear pattern can be described using equations (1) to (2). The coefficients c0 to c5 can be determined using the current template (801) and the reference template (802), where multiple samples within at least one of the current template (801) and the reference template (802) are filtered.

[0139] In the example, the reference block (832) is in a reference picture (812) different from the current picture (811), and the current block (831) is predicted according to the inter-frame prediction method. The linear pattern can be described using equations (3) to (4). The coefficients c6 to c7 can be determined using the current template (801) and the reference template (802), where multiple samples within at least one of the current template (801) and the reference template (802) are filtered.

[0140] In one aspect, the current block (831) can be reconstructed based on the linear model and the reference block (832). The linear model can be applied to the reference block (832) to determine the prediction signal of the current block (831). For example, the prediction signal of the current block (831) includes the predicted values of the samples within the current block (831), similar to Figure 6B as described in and equation (2), or similar to that described in equation (4). The current block (831) can be reconstructed based on the predicted values of the samples.

[0141] Compared with the case of directly copying the reference block (832) to predict the current block (831) (for example, the unmodified reference block (832) is used as the prediction signal of the current block (831)), the prediction signal of the current block (831) determined based on the linear pattern (for example, as described in equation (2) or equation (4)) can be referred to as the updated (or corrected) prediction signal of the current block (831).

[0142] In the example, the reference block (832) is in the current picture (811), and according to the linear model such as Figure 6B and described in equation (2), the predicted values of the samples in the current block (831) (for example, (sample (850))) are the weighted sum of the sample C' in the reference block (832), multiple adjacent samples (for example, N', S', E', and W') of the sample C' in the reference block (832), and the bias term.

[0143] In an example, the reference block (832) is in a reference picture (812) different from the current picture (811), and according to a linear model such as that described in Equation (3), the predicted value of a sample (e.g., sample (850)) in the current block (831) is a weighted sum of the sample C’ in the reference block (832) and a bias term.

[0144] In one aspect, at least one of the current template (801) and the reference template (802) includes the current template (801) and the reference template (802), and a filter can be applied to a plurality of samples within the current template (801) and the reference template (802). The filter can be referred to as a unified filter. A modified version of the current template (801) (e.g., the filtered current template (801)) and a modified version of the reference template (802) (e.g., the filtered reference template (802)) can be used to determine a linear model, e.g., similar to that described in Equation (1) or (3). For example, referring to Equation (1), valC, valN, valS, valE, and valW in Equation (1) can be the filtered sample values of the corresponding samples C, N, S, E, and W within the filtered reference template (802). In addition, the coefficients of the linear model (e.g., c0 to c5 in Equation (1) or c6 to c7 in Equation (3)) can be derived based on the minimization of the difference between the values of the reconstructed samples within the filtered current template (801) and the predicted value of the current template (801), where the predicted value of the current template (801) is obtained, for example, using Equation (1) or Equation (3) based on the filtered reference template (802). Thus, the modified version of the current template (801) and the modified version of the reference template (802) can be used to train the correlation between adjacent blocks.

[0145] In one aspect, the filters applied to the current template (801) and the reference template (802) can be different. Thus, a first filter is applied to the samples within the current template (801), and a second filter is applied to the reference template (802). The second filter is different from the first filter. For example, a first sample within the current template (801) can be filtered with the first filter, and a second sample within the reference template (802) can be filtered with the second filter. The plurality of samples can include the first sample within the current template (801) and the second sample within the reference template (802).

[0146] In one aspect, only one of the current template (801) and the reference template (802) is filtered, while the other of the current template (801) and the reference template (802) is not filtered. For example, at least one of the current template (801) and the reference template (802) consists of the current template (801) or the reference template (802). In an example, at least one of the current template (801) and the reference template (802) consists of the current template (801), a filter is applied to the samples in the current template (801), and the reference template (802) is not modified (e.g., the reference template (802) is not filtered). The modified version of the current template (801) (e.g., the filtered current template (801)) and the unfiltered reference template (802) can be used to determine a linear model, similar to that described in Equation (1) or (3).

[0147] In one aspect, a filter is applied to the samples in the reference template (802), while the current template (801) is not modified. For example, at least one of the current template (801) and the reference template (802) consists of the reference template (802), a filter is applied to the samples in the reference template (802), and the current template (801) is not modified (e.g., the current template (801) is not filtered). The modified version of the reference template (802) (e.g., the filtered reference template (802)) and the unfiltered current template (801) can be used to determine a linear model, similar to that described in Equation (1) or (3).

[0148] Any suitable one or more filters can be used to filter at least one of the current template (801) and the reference template (802). The one or more filters can have any suitable shape and any suitable size. In one aspect, a 3×3 filter F is used to filter the samples in the template, such as filtering a plurality of samples within at least one of the current template (801) and the reference template (802). An example of the 3×3 filter F is given in Equation (5). In an example, the 3×3 filter F can be used to filter a plurality of samples within at least one of the current template (801) and the reference template (802).

[0149]

[0150] Figure 8 An example of the shape of the 3×3 filter F according to one aspect of the present disclosure is shown. Refer to Figure 8 and Equation (5), the central coefficient in the 3×3 filter F (e.g., having a value of 10) can correspond to Figure 8 the central position C in, and the coefficients of the adjacent samples (e.g., NW, N, NE, W, E, SW, S, and SE) of the central position C are -1.

[0151] In one aspect, a syntax element (e.g., a flag) can be signaled in the bitstream to indicate whether a filter is to be applied to multiple samples within at least one of a current template (801) and a reference template (802). For example, the flag can indicate whether to apply filtering to at least one of the current template (801) of the current block (831) and the reference template (802) of the reference block (832). One or more templates to be filtered can include the reference template (802) and / or the current template (801).

[0152] In one aspect, the syntax element (e.g., the flag) can be signaled at the block level or at a higher level above the block level. For example, the flag can be signaled at the block level or in a high-level syntax including but not limited to slice headers, picture headers, sequence headers, etc.

[0153] In one aspect, filtering of one or more templates (e.g., at least one of the current template (801) and the reference template (802)) can be applied only to the luminance component. For example, filtering is applied to at least one of the current template (801) and the reference template (802) only when the current block (831) is a luminance block, and no filtering is applied when the current block (831) is a chrominance block.

[0154] In one aspect, filtering of one or more templates (e.g., at least one of the current template (801) and the reference template (802)) can be applied to the luminance component and one or more chrominance components. In an example, the one or more chrominance components include Cb and Cr.

[0155] In an example, when the current block (831) is a luminance block or a chrominance block, filtering of one or more templates (e.g., at least one of the current template (801) and the reference template (802)) can be applied.

[0156] As described above, the current template (801) can have any suitable size and any suitable shape. In Figure 7 the example shown, the current template (801) has an L shape and includes 4 rows of reconstructed samples above and / or to the left of the current block (831). For example, referring Figure 7 to, the current template (801) includes a top template (862) directly above the current block (831), a left template (863) to the left of the current block (831), and a top-left template (861) between the top template (862) and the left template (863).

[0157] In an example, the current template (801) may include only the top template (862). In an example, the current template (801) may include only the left template (863).

[0158] In one aspect, the template size of a template (e.g., the current template (801) or the reference template (802)) may be extended, for example, to the upper right and lower left.

[0159] Figure 9 An example of the current template (911) of the extended current block (831) according to one aspect of the present disclosure is shown. In Figure 9 the example shown, the current template (911) includes 3 rows of reconstructed samples above the current block (831) and to the left of the current block (831). Referring Figure 9 to, the current template (911) includes a top template (952) directly above the current block (831), a left template (953) to the left of the current block (831), and a top left template (951) between the top template (952) and the left template (953). Figure 9 The current template (911) in is extended to the upper right and lower left of the current block (831) to form an extended current template (901). The extended current template (901) of the current block (831) includes a top template (952) directly above the current block (831), a left template (953) to the left of the current block (831), a top left template (951), an upper right template (922) above and to the right of the current block (831), and a lower left template (923) below and to the left of the current block (831). A reference template (not shown) corresponding to the extended current template (901) may have the same shape and the same size as the extended current template (901).

[0160] In an example, one or more extended sizes of the upper right template (922) and the lower left template (923) may be based on the width W and height H of the current block (831). For example, the width of the upper right template (922) is 2W. The height of the lower left template (923) is 2H.

[0161] In one aspect, a plurality of samples is a subset of the samples within at least one of the current template of the current block and the reference template of the reference block. In one aspect, the filter is not applied to all the samples within the template (e.g., the current template or the reference template), but to a subset of the samples within the template. In an example, referring Figure 9 to, only a row of samples within the template (e.g., the current template (911)) is filtered, for example, only one row / one column (e.g., the middle row of black dots) is filtered. In an example, referring Figure 9, when filtering the current template (911), only the black dots in the middle row are filtered. Therefore, multiple samples include the black dots and do not include other samples within the current template (911). In the example, when filtering the reference template corresponding to the current template (911), only the middle row samples within the reference template are filtered.

[0162] In one aspect, the filter may not be applied to an entire row of samples within the template (such as an entire row including an entire row and an entire column), but rather to a subset of the entire row of samples. Figure 10 An example of the current template (911) is shown, where the filter is not applied to the two samples at the end of the middle row / middle column. The current template (911) includes three rows, such as row 1, row 2, and row 3. As Figure 9 described, Figure 10 the current template (911) in Figure 10 includes a top template (952) directly above the current block (831), a left template (953) to the left of the current block (831), and a top left template (951) between the top template (952) and the left template (953). When filtering the Figure 10 current template (911) in Figure 8 the filter is only applied to samples (732) to (742) (

[0163] the black dots in Figure 11 An example is shown according to one aspect of the present disclosure when filtering two rows of samples (such as two rows / two columns) ( Figure 11 the black dots in Figure 11 the current template (801). Referring to

[0164] In one aspect, referring to Figure 7, the filter shape used for filtering at least one of the current template (801) and the reference template (802) can depend on at least one of the position of the current template (801), the position of the reference template (802), the size of the current block (831), the shape of the current block (831), the size of the current template (801), the shape of the current template (801), the size of the reference template (802), the shape of the reference template (802), etc. For example, the filter can have one of a variety of shapes. The variety of shapes can include a square (e.g., the 3×3 filter shown in Figure 8 ), a rectangle (e.g., a 3×2 filter or a 2×3 filter), etc. Which filter shape to select can depend on the encoded information or known information, including but not limited to the block position relative to the picture (e.g., the block position of the reference block relative to the reference picture, or the block position of the current block relative to the current picture), the block size or block shape of the current block, the block size or block shape of the reference block, the reference (adjacent sample) row, etc. Refer to Figure 11 , for the first reference row (e.g., L1), a first filter shape can be used. For the second reference row (e.g., L2), a second filter shape can be used.

[0165] In one aspect, if the reference template and / or the current template fall on a boundary such as a picture boundary, a strip boundary, etc., the filter can have a different shape. When a part of the reference template is on the boundary, the reference template can be on the boundary. When a part of the current template is on the boundary, the current template can be on the boundary. For example, when the current template falls on the boundary, a first filter shape can be used, and when the current template does not fall on the boundary, a second filter shape can be used.

[0166] Figure 12 FIG. shows a flowchart outlining a process (1200) according to one aspect of the present disclosure. The process (1200) can be used in a device such as a video decoder. In various aspects, the process (1200) is executed by a processing circuit, such as a processing circuit that executes the functions of the video decoder (110), a processing circuit that executes the functions of the video decoder (210), etc. In some aspects, the process (1200) is implemented by software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (1200). The process starts at (S1201) and proceeds to (S1210).

[0167] At (S1210), encoded information in the bitstream is received. The encoded information indicates whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block in the current picture. The current block is predicted based on the reference block of the current block.

[0168] In an example, the encoded information in the bitstream includes a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at the block level or a higher level above the block level.

[0169] In an example, the current template includes the neighboring reconstructed samples of the current block, and the reference template includes the neighboring reconstructed samples of the reference block.

[0170] In an example, when the current block is predicted according to the IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture.

[0171] In an example, when the current block is predicted according to an inter prediction method, the reference block is in a reference picture different from the current picture.

[0172] In an example, the current template includes a top template directly above the current block, a left template to the left of the current block, and a top-left template between the top template and the left template, such as Figure 7 the current template (801) shown in

[0173] In an example, the current template includes a top template, a left template, a top-left template between the top template and the left template, a top-right template above and to the right of the current block, and a bottom-left template below and to the left of the current block, such as Figure 9 the extended current template (901) shown in

[0174] The reference template may have the same shape and the same size as the current template.

[0175] At (S1220), when the encoded information indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, filter a plurality of samples within at least one of the current template of the current block and the reference template of the reference block. Based on the filtered plurality of samples within at least one of the current template and the reference template, determine a linear model between the current template and the reference template.

[0176] In an example, at least one of the current template and the reference template includes the current template and the reference template.

[0177] In an example, filter a first sample among the plurality of samples within the current template of the current block with a first filter, and filter a second sample among the plurality of samples within the reference template of the reference block with a second filter different from the first filter.

[0178] In an example, at least one of the current template and the reference template consists of the current template or the reference template.

[0179] In an example, a filter is used to filter a plurality of samples within at least one of a current template of a current block and a reference template of a reference block, and the filter is

[0180] In the example, filtering is applied to at least one of the current template and the reference template only when the current block is a luminance block.

[0181] In the example, the current block is a luminance block or a chrominance block.

[0182] In the example, the plurality of samples is a subset of the samples within at least one of the current template of the current block and the reference template of the reference block, such as Figures 9 to 11 as described in

[0183] In the example, the filter shape used in filtering depends on at least one of the position of the current template, the position of the reference template, the size of the current block, the shape of the current block, the size of the current template, and the shape of the current template.

[0184] At (S1230), the current block is reconstructed based on the linear model and the reference block.

[0185] In the example, a linear model is applied to the reference block to determine a prediction signal of the current block. The current block can be reconstructed based on the prediction signal. When the reference block is in the current picture, according to a linear model such as described in Figures 6A to 6B and equations (1) to (2), the sample value in the prediction signal is a weighted sum of the samples in the reference block, a plurality of adjacent samples of the samples in the reference block, and a bias term. When the reference block is in a reference picture, according to a linear model such as described in equations (3) to (4), the sample value in the prediction signal is a weighted sum of the samples in the reference block and a bias term.

[0186] Then, the process proceeds to (S1299) and terminates.

[0187] The process (1200) can be appropriately modified. One or more steps in the process (1200) can be modified and / or omitted. One or more other steps can be added. Any suitable order of implementation manners can be used.

[0188] Figure 13A flowchart showing a process (1300) according to an aspect of the present disclosure is shown. The process (1300) can be used in a video encoder. In various aspects, the process (1300) is executed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (1300) is implemented as software instructions, so when the processing circuit executes the software instructions, the processing circuit executes the process (1300). The process starts at (S1301) and proceeds to (S1310).

[0189] At (S1310), when filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block in the current picture, multiple samples within at least one of the current template of the current block and the reference template of the reference block are filtered. Based on the filtered multiple samples within at least one of the current template and the reference template, a linear model between the current template and the reference template is determined.

[0190] In an example, when the current block is encoded according to the IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture. When the current block is encoded according to an inter prediction mode, the reference block is in a reference picture different from the current picture. The current template includes adjacent samples of the current block. The reference template includes adjacent samples of the reference block.

[0191] In an example, at least one of the current template and the reference template includes the current template and the reference template.

[0192] In an example, a first sample among the multiple samples within the current template of the current block is filtered with a first filter, and a second sample among the multiple samples within the reference template of the reference block is filtered with a second filter different from the first filter.

[0193] In an example, at least one of the current template and the reference template consists of the current template or the reference template.

[0194] At (S1320), the current block is encoded based on the linear model and the reference block.

[0195] At (S1330), a syntax element is encoded in the bitstream, which indicates whether filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block.

[0196] In an example, the syntax element is a flag indicating whether filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at the block level or a higher level above the block level.

[0197] Then, the process proceeds to (S1399) and terminates.

[0198] The process (1300) can be modified appropriately. One or more steps in the process (1300) can be modified and / or omitted. One or more additional steps can be added. Any suitable order of implementation can be used.

[0199] Although the decoding and encoding processes are provided in separate flowcharts for purposes of description, it should be noted that aspects of the decoding and encoding processes can be used in combination. For example, a decoding process such as that described in process (1200) can include all or a portion of process (1300). In another example, an encoding process such as that described in process (1300) can be combined with process (1200).

[0200] In one aspect, a method of processing visual media data is disclosed. The method includes processing a bitstream of visual media data according to formatting rules. The bitstream includes a syntax element indicating whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture, where the current block is predicted based on a reference block of the current block. The formatting rules specify that when encoded information indicates that filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block, filtering is applied to a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and a linear model between the current template and the reference template is determined based on the filtered plurality of samples within at least one of the current template and the reference template. The formatting rules specify reconstructing the current block based on the linear model and the reference block.

[0201] Aspects and / or examples in the present disclosure can be used alone or in any combination. For example, some aspects and / or examples performed by a decoder can be performed by an encoder and vice versa. Each of the method, aspects, examples, encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non - volatile computer - readable medium.

[0202] The techniques described above can be implemented as computer software using computer - readable instructions and physically stored in one or more computer - readable media. For example, Figure 14 A computer system (1400) is shown that is suitable for implementing certain aspects of the disclosed subject matter.

[0203] The computer software can be encoded using any suitable machine code or computer language, which can be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that can be executed directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0204] The instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0205] Figure 14 The components shown for the computer system (1400) are examples and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence on or requirement for any one or combination of the components shown in the example aspects of the computer system (1400).

[0206] The computer system (1400) may include certain human - machine interface input devices. Such human - machine interface input devices can respond to inputs by one or more human users through, for example, tactile inputs (e.g., key presses, swipes, data glove movements), audio inputs (e.g., voice, taps), visual inputs (e.g., gestures), olfactory inputs (not depicted). The human - machine interface devices can also be used to capture certain media not directly related to conscious human input, such as audio (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still - image camera), video (e.g., two - dimensional video, three - dimensional video including stereoscopic video).

[0207] The input human - machine interface devices can include one or more of the following (each depicted only one): keyboard (1401), mouse (1402), trackpad (1403), touch screen (1410), data glove (not shown), joystick (1405), microphone (1406), scanner (1407), camera (1408).

[0208] The computer system (1400) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., the tactile feedback of a touch screen (1410), a data glove (not shown), or a joystick (1405), but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (1409), headphones (not depicted)), visual output devices (e.g., a screen (1410), including a cathode ray tube (CRT) screen, a liquid crystal display (LCD) screen, a plasma screen, an organic light-emitting diode (OLED) screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three-dimensional through, for example, stereoscopic flat painting output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).

[0209] The computer system (1400) may also include human-accessible storage devices and associated media of the storage devices, such as optical media, including CD / DVD ROM / RW (1420) with media such as CD / DVD (1421), thumb drives (1422), removable hard disk drives, or solid state drives (1423), legacy magnetic media such as tapes and floppy disks (not depicted), ROM / based application-specific integrated circuit (ASIC) / programmable logic device (PLD)-based special devices, such as security protection devices (not depicted), and so on.

[0210] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the currently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0211] The computer system (1400) may also include an interface (1454) to one or more communication networks (1455). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, latency-tolerant, etc. Examples of networks include, for example, Ethernet, local area networks of wireless LANs, cellular networks including Global System for Mobile Communications (GSM), third generation (3G), fourth generation (4G), fifth generation (5G), Long Term Evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular networks and industrial networks including Controller Area Network Bus (CANBus), etc. Some networks typically require an external network interface adapter attached to certain general-purpose data ports or peripheral buses (1449) (e.g., Universal Serial Bus (USB) ports of the computer system (1400)); other networks are typically integrated into the core of the computer system (1400) by attaching to the system bus as described below (e.g., integrated into a PC computer system through an Ethernet interface, or integrated into a smartphone computer system through a cellular network interface). By using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be only one-way reception (e.g., broadcast TV), only one-way transmission (e.g., CANBus connected to certain CANBus devices), or two-way, for example, using a local digital network or a wide area digital network to connect to other computer systems. Certain protocols and protocol stacks can be used on each of the networks and network interfaces as described above.

[0212] The above-described human-machine interface device, human-accessible storage device, and network interface may be attached to the core (1440) of the computer system (1400).

[0213] The core (1440) may include one or more central processing units (CPUs) (1441), a graphics processing unit (GPU) (1442), a dedicated programmable processing unit in the form of field programmable gate areas (FPGAs) (1443), a hardware accelerator (1444) for certain tasks, a graphics adapter (1450), and so on. These devices, together with a read-only memory (ROM) (1445), a random access memory (1446), and internal mass storage devices such as internal hard disk drives, solid state drives (SSDs), etc. that are not accessible by users internally (1447), can be connected via a system bus (1448). In some computer systems, the system bus (1448) can be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly or via a peripheral bus (1449) to the system bus (1448) of the core. In an example, a screen (1410) can be connected to the graphics adapter (1450). Architectures for peripheral buses include Peripheral Component Interconnect (PCI), USB, and so on.

[0214] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute certain instructions that, when combined, can constitute the above-mentioned computer code. The computer code can be stored in the ROM (1445) or the RAM (1446). Transitional data can also be stored in the RAM (1446), while permanent data can be stored, for example, in the internal mass storage device (1447). Fast storage and retrieval of any memory device can be achieved by using a cache memory that can be closely associated with one or more CPUs (1441), GPUs (1442), mass storage devices (1447), ROM (1445), RAM (1446), etc.

[0215] Computer code for performing various computer-implemented operations can be present on a computer-readable medium. The medium and the computer code can be those designed and constructed specifically for the purposes of this application, or can be of the kinds well-known and available to those skilled in the field of computer software.

[0216] By way of example and not limitation, a computer system having an architecture (1400) and in particular a core (1440) can provide functions resulting from software executed by processors (including CPUs, GPUs, FPGAs, accelerators, etc.) embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices and certain non-transitory storage devices of the core (1440) introduced above (e.g., the on-core mass storage device (1447) or ROM (1445) within the core). The software implementing aspects of the present application can be stored in such devices and executed by the core (1440). Depending on specific requirements, the computer-readable media can include one or more memory devices or chips. The software can cause the core (1440) and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.) to execute specific processes or specific parts of specific processes described herein, including defining data structures stored in the RAM (1446) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system can provide functions resulting from logic hardwired or otherwise embodied in circuitry (e.g., accelerator (1444)), which can operate instead of or in conjunction with the software to execute specific processes or specific parts of specific processes described herein. Where appropriate, references to software can encompass logic and vice versa. Where appropriate, references to computer-readable media can encompass circuitry (e.g., an integrated circuit (IC)) storing software for execution, circuitry embodying logic for execution, or both types of circuitry. The present application encompasses any suitable combination of hardware and software.

[0217] As used in this disclosure, "at least one" or "one of" is intended to include any one or combination of the recited elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include only A, only B, only C, or any combination thereof. Reference to one of A or B and one of A and B is intended to include A or B or (A and B). Use of "one of" does not exclude any combination of the recited elements in cases where applicable, such as when the elements are not mutually exclusive.

[0218] Although the present application describes several examples, within the scope of the present application, there can be various alterations, permutations, and various alternative equivalents. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, can embody the principles of the present application.

[0219] The above disclosure also encompasses the features recited below. These features can be combined in various ways and are not limited to the features recited below.

[0220] (1) A video decoding method, the method comprising: receiving encoded information in a bitstream, the encoded information indicating whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture, the current block being predicted based on a reference block of the current block; when the encoded information indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block; determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; and reconstructing the current block based on the linear model and the reference block.

[0221] (2) The method according to feature (1), wherein when the current block is predicted according to the IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture; when the current block is predicted according to an inter prediction method, the reference block is in a reference picture different from the current picture; the current template includes adjacent reconstructed samples of the current block; and the reference template includes adjacent reconstructed samples of the reference block.

[0222] (3) The method according to feature (2), wherein the reconstruction includes: applying the linear model to the reference block to determine a prediction signal of the current block; and reconstructing the current block based on the prediction signal; when the reference block is in the current picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block, a plurality of adjacent samples of the samples in the reference block, and a bias term; when the reference block is in the reference picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block and the bias term.

[0223] (4) The method according to any one of features (1) to (3), wherein at least one of the current template and the reference template includes the current template and the reference template.

[0224] (5) The method according to feature (4), wherein the filtering includes: filtering a first sample among the plurality of samples within the current template of the current block with a first filter; and filtering a second sample among the plurality of samples within the reference template of the reference block with a second filter different from the first filter.

[0225] (6) The method according to any one of features (1) to (3), wherein at least one of the current template and the reference template consists of the current template or the reference template.

[0226] (7) The method according to any one of features (1) to (6), wherein the filtering includes: using a filter to filter a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, the filter being

[0227] (8) The method according to any one of features (1) to (7), wherein the encoded information in the bitstream includes a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at the block level or a higher level than the block level.

[0228] (9) The method according to any one of features (1) to (8), wherein filtering is applied to at least one of the current template and the reference template only when the current block is a luminance block.

[0229] (10) The method according to any one of features (1) to (8), wherein the current block is a luminance block or a chrominance block.

[0230] (11) The method according to any one of features (1) to (10), wherein the current template includes at least one of the following: a top template directly above the current block, a left template to the left of the current block, and an upper left template between the top template and the left template; and the top template, the left template, the upper left template between the top template and the left template, an upper right template above and to the right of the current block, and a lower left template below and to the left of the current block; and the reference template has the same shape and the same size as the current template.

[0231] (12) The method according to any one of features (1) to (11), wherein the plurality of samples is a subset of the samples within at least one of the current template of the current block and the reference template of the reference block.

[0232] (13) The method according to any one of features (1) to (12), wherein the filter shape used for filtering depends on at least one of the position of the current template, the position of the reference template, the size of the current block, the shape of the current block, the size of the current template, and the shape of the current template.

[0233] (14) A video coding method, the method comprising: when filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block in the current picture, filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block; determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; encoding the current block based on the linear model and the reference block; and encoding a syntax element in the bitstream indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block.

[0234] (15) The method according to feature (14), wherein when the current block is encoded according to the IntraTMP (Intra Template Matching Prediction) mode, the reference block is in the current picture; when the current block is encoded according to the inter prediction mode, the reference block is in a reference picture different from the current picture; the current template includes adjacent samples of the current block; and the reference template includes adjacent samples of the reference block.

[0235] (16) The method according to any one of features (14) to (15), wherein at least one of the current template and the reference template includes the current template and the reference template.

[0236] (17) The method according to any one of features (14) to (16), wherein filtering includes: filtering a first sample among a plurality of samples within the current template of the current block with a first filter; and filtering a second sample among a plurality of samples within the reference template of the reference block with a second filter different from the first filter.

[0237] (18) The method according to any one of features (14) to (15), wherein at least one of the current template and the reference template consists of the current template or the reference template.

[0238] (19) The method according to any one of features (14) to (18), wherein the syntax element is a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at the block level or a higher level than the block level.

[0239] (20) A method for processing visual media data, the method including: processing a bitstream of visual media data according to format rules, wherein the bitstream includes a syntax element indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block in the current picture, the current block being predicted based on a reference block of the current block; and the format rules specify that when the syntax element indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; and reconstructing the current block based on the linear model and the reference block.

[0240] (21) A video decoding apparatus, including a processing circuit for performing the method according to any one of features (1) to (13).

[0241] (22) A video encoding apparatus, including a processing circuit for performing the method according to any one of features (14) to (19).

[0242] (23) A non-volatile computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method according to any one of features (1) to (19).

Claims

1. A video decoding method, characterized in that, The method includes: Receiving encoded information in a bitstream, the encoded information indicating whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture, where the current block is predicted based on the reference block of the current block; When the encoded information indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, Filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block; Determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; and Reconstructing the current block based on the linear model and the reference block.

2. The method according to claim 1, wherein When the current block is predicted according to the IntraTMP (Intra Template Matching) mode, the reference block is in the current picture; When the current block is predicted according to an inter prediction method, the reference block is in a reference picture different from the current picture; The current template includes adjacent reconstructed samples of the current block; The reference template includes adjacent reconstructed samples of the reference block.

3. The method according to claim 2, wherein The reconstruction includes: Applying the linear model to the reference block to determine a prediction signal of the current block; and Reconstructing the current block based on the prediction signal; When the reference block is in the current picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block, a plurality of adjacent samples of the samples in the reference block, and a bias term; When the reference block is in the reference picture, according to the linear model, the sample value in the prediction signal is a weighted sum of samples in the reference block and the bias term.

4. The method according to claim 1, wherein At least one of the current template and the reference template includes the current template and the reference template.

5. The method according to claim 4, characterized in that, The filtering includes: Filtering a first sample among a plurality of samples within the current template of the current block with a first filter; and Filtering a second sample among a plurality of samples within the reference template of the reference block with a second filter different from the first filter.

6. The method according to claim 1, characterized in that, At least one of the current template and the reference template consists of the current template or the reference template.

7. The method according to claim 1, wherein The filtering includes: using a filter to filter a plurality of samples in at least one of the current template of the current block and the reference template of the reference block, and the filter is 8. The method according to claim 1, characterized in that, The encoded information in the bitstream includes a flag indicating whether to apply filtering to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at the block level or a higher level.

9. The method according to claim 1, characterized in that Filtering is applied to at least one of the current template and the reference template only when the current block is a luminance block.

10. The method according to claim 1, characterized in that The current block is a luminance block or a chrominance block.

11. The method according to claim 1, wherein The current template includes at least one of the following: a top template directly above the current block, a left template to the left of the current block, and a top-left template between the top template and the left template; and the top template, the left template, the top-left template between the top template and the left template, a top-right template above and to the right of the current block, and a bottom-left template below and to the left of the current block; and the reference template has the same shape and the same size as the current template.

12. The method according to claim 1, wherein The plurality of samples is a subset of the samples within at least one of the current template of the current block and the reference template of the reference block.

13. The method according to claim 1, wherein The filter shape used for the filtering depends on at least one of the position of the current template, the position of the reference template, the size of the current block, the shape of the current block, the size of the current template, and the shape of the current template.

14. A video encoding method, characterized in that, The method includes: when filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block in a current picture, filtering a plurality of samples within at least one of the current template of the current block and the reference template of the reference block; determining a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; encoding the current block based on the linear model and the reference block; and encoding a syntax element in a bitstream, the syntax element indicating whether filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block.

15. The method according to claim 14, wherein when the current block is encoded according to an Intra Template Matching Prediction (IntraTMP) mode, the reference block is in the current picture; when the current block is encoded according to an inter prediction mode, the reference block is in a reference picture different from the current picture; the current template includes neighboring samples of the current block; and the reference template includes neighboring samples of the reference block.

16. The method according to claim 14, wherein At least one of the current template and the reference template includes the current template and the reference template.

17. The method according to claim 16, wherein The filtering includes: filtering a first sample among the plurality of samples within the current template of the current block with a first filter; and filtering a second sample among the plurality of samples within the reference template of the reference block with a second filter different from the first filter.

18. The method according to claim 14, wherein At least one of the current template and the reference template consists of the current template or the reference template.

19. The method according to claim 14, wherein The syntax element is a flag indicating whether filtering is to be applied to at least one of the current template of the current block and the reference template of the reference block, and the flag is signaled at a block level or a higher level above the block level.

20. A method for processing visual media data, characterized in that, The method includes: processing the bitstream of the visual media data according to format rules, wherein The bitstream includes a syntax element indicating whether to apply filtering to at least one of a current template of a current block and a reference template of a reference block in a current picture, where the current block is predicted based on the reference block of the current block; and The formatting rule specifies that: when the syntax element indicates to apply filtering to at least one of the current template of the current block and the reference template of the reference block, filter a plurality of samples within at least one of the current template of the current block and the reference template of the reference block, and determine a linear model between the current template and the reference template based on the filtered plurality of samples within at least one of the current template and the reference template; and reconstruct the current block based on the linear model and the reference block.