Signaling using template matching costs
By introducing a nonlinear prediction model based on template matching in the video encoding technology, the problem of difficult to effectively utilize template matching prediction in the prior art is solved, and more efficient video encoding and transmission is achieved.
Patent Information
- Application Number
- CN202480004461.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-14
- Filing Date
- 2024-05-28
- Publication Date
- 2025-06-06
AI Technical Summary
When existing video encoding technology processes video data, it is difficult to effectively utilize template matching prediction methods to improve encoding efficiency.
By introducing a template matching based prediction model in the video encoding method, the reference block is processed using a nonlinear function to determine the prediction block of the current block and encode it into the code stream.
The efficiency of video encoding is improved, and through more accurate prediction block determination is reduced, the amount of encoded data is improved, and the quality of video transmission is improved.
Smart Images

Figure CN120113237A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 526,960, filed on July 14, 2023, “On Improvement of Signaling By Using Template Matching Cost,” the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure generally describes aspects related to video encoding. Background Art
[0003] The background description provided herein is intended to present the context of the present disclosure in general. The existing works of the inventors within the scope described in the background section, as well as aspects of the specification that may not belong to the prior art at the time of application, are neither explicitly nor implicitly admitted as the prior art of the present disclosure.
[0004] Image / video compression helps to transmit image / video data between different devices, storages, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which can compress images based on spatial redundancy. For example, intra-frame prediction can use reference data from the current picture being reconstructed for sample prediction. In another example, a video codec can use a technique called inter-frame prediction, which can compress images based on temporal redundancy. For example, inter-frame prediction can use motion compensation to predict samples in the current picture from a previously reconstructed picture. Motion compensation can be indicated by a motion vector (MV). Summary of the invention
[0005] Aspects of the present disclosure include code streams, methods, and devices for video encoding / decoding. In some examples, the device for video encoding / decoding includes a processing circuit.
[0006] According to one aspect of the present disclosure, a method for processing visual media data is provided. In the method, a code stream of the visual media data is processed according to a format rule. In an example, the code stream includes a first syntax element, and the first syntax element indicates whether a prediction method based on a nonlinear function is applied to process a current block based on template matching prediction. The prediction method based on the nonlinear function applies a prediction model, and the prediction model is associated with a template type of a template of a current block and includes a nonlinear function. The nonlinear function is applied to a reference block of the current block. The nonlinear function includes parameters derived based on the template of the current block and the template of the reference block. The format rule stipulates that when the first syntax element indicates that the prediction method based on the nonlinear function is applied to the current block, multiple template matching (Template Matching, TM) costs between the template of the current block and the template of the reference block are determined according to multiple candidate prediction models associated with the prediction method based on the nonlinear function. The format rule stipulates that a prediction model is determined from multiple candidate prediction models based on multiple TM costs. For example, the prediction model corresponds to a selected TM cost among multiple TM costs. The format rule stipulates that the nonlinear function of the prediction model is applied to the sample of the processing reference block. The format rule stipulates that the prediction block of the current block is determined as the reference block processed by the nonlinear function of the prediction model.
[0007] In the example, the format rule stipulates that the function of each candidate prediction model in the plurality of candidate prediction models is applied to the template of the current block and the template of the reference block to obtain the processed template of the current block and the processed template of the reference block. Based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model, each TM cost in the plurality of TM costs is determined. The format rule stipulates that a prediction model corresponding to the minimum TM cost in the plurality of TM costs is determined from the plurality of candidate prediction models.
[0008] In the example, the format rule provides for generating processed samples of the template of the current block and processed samples of the template of the reference block based on a first candidate prediction model among multiple candidate prediction models. The format rule provides for determining a first TM cost among multiple TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block. The format rule provides for filtering samples of the template of the current block and samples of the template of the reference block. The format rule provides for generating processed filtered samples of the template of the current block and processed filtered samples of the template of the reference block based on the first candidate prediction model. The format rule provides for determining a second TM cost among multiple TM costs based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
[0009] According to another aspect of the present disclosure, a video encoding method is provided. In the method, a reference block of a current block is determined based on a template of a current block and a template of a reference block according to intra template matching prediction (intraTMP). A prediction model is determined from a plurality of candidate prediction models based on a plurality of TM costs between the template of the current block and the template of the reference block according to a plurality of candidate prediction models. The prediction model is associated with a template type of the template of the current block and includes a function applied to a reference block of the current block. The function includes parameters derived based on the template of the current block and the template of the reference block. A prediction block of the current block is determined based on the prediction model. The current block is encoded into a bitstream based on the determined prediction block, and a syntax element is encoded into the bitstream. The syntax element indicates whether a function-based prediction method including a prediction model is applied to reconstruct the current block.
[0010] In the example, a function of each candidate prediction model in a plurality of candidate prediction models is applied to a template of a current block and a template of a reference block to obtain a processed template of the current block and a processed template of the reference block. Based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model, each TM cost in a plurality of TM costs is determined. A prediction model corresponding to a minimum TM cost in a plurality of TM costs is determined from a plurality of candidate prediction models.
[0011] In an example, a processed sample of a template of a current block and a processed sample of a template of a reference block are generated based on a first candidate prediction model among a plurality of candidate prediction models. A first TM cost among a plurality of TM costs is determined based on the processed sample of the template of the current block and the processed sample of the template of the reference block. Samples of the template of the current block and samples of the template of the reference block are filtered. Processed filtered samples of the template of the current block and processed filtered samples of the template of the reference block are generated based on the first candidate prediction model. A second TM cost among a plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
[0012] According to another aspect of the present disclosure, a video decoding device is provided. The device includes a processing circuit. The processing circuit is configured to receive a code stream, which includes a first syntax element, and the first syntax element indicates whether to apply a function-based prediction method to reconstruct a current block based on template matching prediction. The function-based prediction method applies a prediction model, the prediction model is associated with the template type of the template of the current block, and the prediction model includes a function applied to a reference block of the current block. The function includes parameters derived based on the template of the current block and the template of the reference block. The processing circuit is configured to: when the first syntax element indicates that the function-based prediction method is applied to reconstruct the current block, based on multiple candidate prediction models associated with the function-based prediction method, determine the prediction model from multiple candidate prediction models based on multiple TM costs between the template of the current block and the template of the reference block. The processing circuit is configured to determine the prediction block of the current block as a reference block processed by the function of the prediction model.
[0013] In an example, the processing circuit is configured to apply a function of each candidate prediction model in a plurality of candidate prediction models to a template of a current block and a template of a reference block to obtain a processed template of the current block and a processed template of the reference block. The processing circuit is configured to determine each TM cost in a plurality of TM costs based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model. The processing circuit is configured to determine a prediction model corresponding to a minimum TM cost in a plurality of TM costs from the plurality of candidate prediction models.
[0014] In the example, the template types include (i) a first template type, the first template type includes a first template area located on the left side of the current block, (ii) a second template type, the second template type includes a second template area located on the upper side of the current block, (iii) a third template type, the third template type includes the first template area located on the left side of the current block and the second template area located on the upper side of the current block, and (iv) a fourth template type, the fourth template type includes the first template area located on the left side of the current block, the second template area located on the upper side of the current block, and an upper left template area located at the upper left corner of the current block and set between the first template area and the second template area.
[0015] In an example, the processing circuit is configured to generate processed samples of a template of a current block and processed samples of a template of a reference block based on a first candidate prediction model among multiple candidate prediction models. The processing circuit is configured to determine a first TM cost among multiple TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block. The processing circuit is configured to filter samples of the template of the current block and samples of the template of the reference block. The processing circuit is configured to generate processed filtered samples of the template of the current block and processed filtered samples of the template of the reference block based on the first candidate prediction model. The processing circuit is configured to determine a second TM cost among multiple TM costs based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
[0016] In an example, the processing circuit is configured to subtract an average value of the samples of the template of the current block from each sample of the samples of the template of the current block to generate an adjusted sample of the template of the current block. The processing circuit is configured to subtract an average value of the samples of the template of the reference block from each sample of the samples of the template of the reference block to generate an adjusted sample of the template of the current block. The processing circuit is configured to determine a third TM cost of the plurality of TM costs based on the adjusted sample of the template of the current block and the adjusted sample of the template of the reference block.
[0017] In an example, the processing circuit is configured to determine a first TM cost among a plurality of TM costs based on samples of a template of a current block and samples of a template of a reference block. The processing circuit is configured to generate processed samples of the template of the current block and processed samples of the template of the reference block based on a first candidate prediction model among a plurality of candidate prediction models. The processing circuit is configured to determine a second TM cost among a plurality of TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block.
[0018] In an example, the processing circuit is configured to determine a TM cost from a plurality of TM costs indicated by a second syntax element in a code stream. The processing circuit is configured to determine a prediction model corresponding to the determined TM cost from a plurality of candidate prediction models. The processing circuit is configured to apply a function of the prediction model to process samples of a reference block. The processing circuit is configured to determine a prediction block of a current block based on samples of the reference block processed by the function of the prediction model.
[0019] In an example, the processing circuit is configured to normalize each of the plurality of TM costs by dividing the corresponding TM cost by a template size corresponding to the corresponding TM cost. The processing circuit is configured to reorder the plurality of normalized TM costs in ascending order based on the normalized TM costs. The processing circuit is configured to generate a candidate list based on the plurality of reordered normalized TM costs.
[0020] In an example, the processing circuit is configured to receive a third syntax element in a code stream, the third syntax element indicating which of a plurality of reordered normalized TM costs is selected. The third syntax element is associated with at least one of a Cb component or a Cr component of a current block. The processing circuit is configured to determine a prediction model corresponding to the TM cost indicated by the third syntax element from a plurality of candidate prediction models.
[0021] Aspects of the present disclosure also provide a video encoding device. The video encoding device includes a processing circuit, which is configured to implement any of the described video encoding methods.
[0022] Aspects of the present disclosure also provide a video decoding method, which includes any method implemented by a video decoding device.
[0023] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions, which, when executed by a computer, causes the computer to perform any of the described video decoding / encoding methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Further features, properties and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0025] Figure 1 is a schematic diagram showing an example of a block diagram of a communication system (100).
[0026] Figure 2 is a diagram showing an example of a block diagram of a decoder.
[0027] Figure 3 is a diagram showing an example of a block diagram of an encoder.
[0028] Figure 4 is a schematic diagram illustrating intra template matching prediction (IntraTMP) according to some aspects of the present disclosure.
[0029] Figure 5 is a schematic diagram illustrating different template types according to some aspects of the present disclosure.
[0030] Figure 6 A flow chart outlining a decoding process according to some aspects of the present disclosure is shown.
[0031] Figure 7 A flow chart outlining an encoding process according to some aspects of the present disclosure is shown.
[0032] Figure 8 is a schematic diagram of a computer system according to one aspect. DETAILED DESCRIPTION
[0033] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, namely a video encoder and a video decoder in a streaming environment. The disclosed subject matter can be equally applicable to other video-enabled applications including, for example, video conferencing, digital television, streaming services, storing compressed video on digital media including CDs, DVDs, memory sticks, etc., and the like.
[0034] The video processing system (100) includes an acquisition subsystem (113), which may include a video source (101), such as a digital camera, for creating, for example, an uncompressed video picture stream (102). In an example, the video picture stream (102) includes samples captured by the digital camera. Compared to the encoded video data (104) (or the encoded video bitstream), the video picture stream (102) is depicted as a thick line to emphasize the high data volume of the video picture stream, and the video picture stream (102) can be processed by an electronic device (120), which includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination of hardware and software to implement or implement various aspects of the disclosed subject matter as described in more detail below. Compared to the video picture stream (102), the encoded video data (104) (or the encoded video bitstream) is depicted as a thin line to emphasize the lower amount of data of the encoded video data (104) (or the encoded video bitstream), which can be stored on the streaming server (105) for future use. One or more streaming client subsystems, such as Figure 1 A client subsystem (106) and a client subsystem (108) in a streaming server (105) may access a copy (107) and a copy (109) of the encoded video data (104). The client subsystem (106) may include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) may decode an input copy (107) of the encoded video data, and the video decoder may create an output video picture stream (111) that may be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), the video data (107), and the video data (109) (e.g., a video bitstream) may be encoded according to certain video encoding / compression standards. Examples of these standards include ITU-T H.265. In an example, the video coding standard under development is informally referred to as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of the VVC standard.
[0035] It should be noted that the electronic device (120) and the electronic device (130) may include other components (not shown). For example, the electronic device (120) may include a video decoder (not shown), and the electronic device (130) may also include a video encoder (not shown).
[0036] Figure 2 An example of a block diagram of a video decoder (210) is shown. The video decoder (210) may be provided in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used to replace Figure 1 A video decoder (110) is shown in an example.
[0037] The receiver (231) can receive one or more encoded video sequences included in a bitstream decoded by the video decoder (210). In one aspect, one encoded video sequence is received at a time, wherein the decoding of each encoded video sequence is independent of the other encoded video sequences. The encoded video sequence can be received from a channel (201), which can be a hardware / software link to a storage device storing the encoded video data. The receiver (231) can receive the encoded video data and other data, such as encoded audio data and / or auxiliary data streams that can be forwarded to their respective use entities (not depicted). The receiver (231) can separate the encoded video sequence from the other data. To prevent network jitter, a buffer memory (215) can be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter referred to as the parser (220)). In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) can be set outside the video decoder (210) (not shown). In other cases, a buffer memory (not shown) is provided outside the video decoder (210) to prevent network jitter, for example. In addition, another buffer memory (215) may be configured inside the video decoder (210) to handle playback timing, for example. When the receiver (231) receives data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, it may not be necessary to configure the buffer memory (215), or the buffer memory may be made smaller. For use on a traffic packet network such as the Internet, a buffer memory (215) may also be required, which may be relatively large and may advantageously have an adaptive size, and may be at least partially implemented in an operating system or similar element (not depicted) external to the video decoder (210).
[0038] The video decoder (210) may include a parser (220) to reconstruct symbols (221) from an entropy coded video sequence. Figure 2As shown, the categories of these symbols include information for managing the operation of the video decoder (210) and potential information for controlling a display device (212) (e.g., a display screen) that is not part of the electronic device (230) but can be coupled to the electronic device (230). The control information for the (one or more) display devices can be a parameter set fragment (not shown) of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI). The parser (220) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can be performed according to a video coding technique or standard, and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser (220) can extract a subgroup parameter set for at least one subgroup of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The parser (220) may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0039] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to create symbols (221).
[0040] Depending on the type of the coded video picture or part of the coded video picture (e.g., inter-frame and intra-frame pictures, inter-frame blocks and intra-frame blocks) and other factors, the reconstruction of the symbol (221) may involve multiple different units. Which units are involved and how they are involved can be controlled by the parser (220) through subgroup control information parsed from the coded video sequence. For the sake of brevity, this subgroup control information flow between the parser (220) and the multiple units is not described below.
[0041] In addition to the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In actual implementations operating under commercial constraints, many of these units interact closely with each other and may be at least partially integrated with each other. However, for the purposes of describing the disclosed subject matter, the conceptual subdivision into the following functional units is appropriate.
[0042] The first unit is a sealer / inverse transform unit (251). The sealer / inverse transform unit (251) receives quantized transform coefficients as symbols (221) from the parser (220) and control information, including which transform method to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit (251) can output blocks including sample values, which can be input into an aggregator (255).
[0043] In some cases, the output samples of the scaler / inverse transform unit (251) may belong to an intra-coded block. An intra-coded block does not use prediction information from a previously reconstructed picture, but may use a block of predictive information from a previously reconstructed portion of a current picture. Such predictive information may be provided by an intra-frame prediction unit (252). In some cases, the intra-frame picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding reconstructed information extracted from the current picture buffer (258). The current picture buffer (258) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. In some cases, the aggregator (255) adds the prediction information generated by the intra-frame prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251) on a per-sample basis.
[0044] In other cases, the output samples of the scaler / inverse transform unit (251) may belong to a block that is inter-coded and possibly motion compensated. In this case, the motion compensated prediction unit (253) may access the reference picture memory (257) to obtain samples for prediction. After the extracted samples are motion compensated according to the block-related symbols (221), these samples may be added to the output of the scaler / inverse transform unit (251) (in this case referred to as residual samples or residual signals) by the aggregator (255), thereby generating output sample information. The acquisition of prediction samples by the motion compensated prediction unit (253) from an address within the reference picture memory (257) may be controlled by a motion vector, which is provided to the motion compensated prediction unit (253) in the form of a symbol (221), which symbol (221) may include, for example, an X component, a Y component and a reference picture component. Motion compensation may also include interpolation of sample values obtained from the reference picture memory (257) when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0045] The output samples of the aggregator (255) may be used by various loop filtering techniques in a loop filter unit (256). The video compression techniques may include in-loop filter techniques controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and available to the loop filter unit (256) as symbols (221) from the parser (220). The video compression may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of an encoded picture or encoded video sequence, and to previously reconstructed and loop filtered sample values.
[0046] The output of the loop filter unit (256) may be a sample stream that may be output to a display device (212) and stored in a reference picture memory (257) for subsequent inter-frame prediction.
[0047] Once fully reconstructed, certain coded pictures may be used as reference pictures for future prediction. For example, once the coded picture corresponding to the current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) may become part of the reference picture memory (257) and a new current picture buffer may be reallocated before starting to reconstruct a subsequent coded picture.
[0048] The video decoder (210) may perform decoding operations according to a predetermined video compression technique recorded in a standard such as ITU-T Rec. H.265. The encoded video sequence may conform to the syntax of the video compression technology or standard used in the sense that the encoded video sequence follows the syntax specified by the video compression technology or standard and the profile recorded in the video compression technology or standard. Specifically, the profile may select certain tools from all available tools available in the video compression technology or standard as the only tools available for use under the profile. For compliance, the complexity of the encoded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, mega samples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further limited by the Hypothetical Reference Decoder (HRD) specification and metadata of the HRD buffer management written in the encoded video sequence.
[0049] In an embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be part of the encoded video sequence. The video decoder (210) may use the additional data to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0050] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmission circuit). The video encoder (303) can be used to replace Figure 1 A video encoder (103) in an example.
[0051] The video encoder (303) can be used to obtain the video source (301) (not Figure 3 In another example, the video source (301) is a part of the electronic device (320) to receive video samples, and the video source can collect video pictures to be encoded by the video encoder (303). In another example, the video source (301) is a part of the electronic device (320).
[0052] The video source (301) may provide a source video sequence in the form of a digital video sample stream to be encoded by the video encoder (303), the digital video sample stream may have any suitable bit depth (e.g. 8-bit, 10-bit, 12-bit, ...), any color space (e.g. BT.601 Y CrCB, RGB, ...) and any suitable sampling structure (e.g. Y CrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) may be a storage device storing pre-prepared videos. In a video conferencing system, the video source (301) may be a camera that captures local picture information as a video sequence. The video data may be provided as a plurality of separate pictures that impart motion when viewed sequentially. The pictures themselves may be organized as a spatial pixel array, where each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The following description focuses on samples.
[0053] According to one aspect, the video encoder (303) can encode and compress pictures of a source video sequence into an encoded video sequence (343) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed is a function of the controller (350). In some aspects, the controller (350) controls other functional units described below and is functionally coupled to the other functional units. For the sake of brevity, couplings are not shown in the figure. Parameters set by the controller (350) can include rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, etc.), picture size, picture group (Group of Picture, GOP) layout, maximum motion vector search range, etc. The controller (350) can be used to have other suitable functions, which may be different from the video encoder (303) optimized for a certain system design.
[0054] In some aspects, the video encoder (303) can operate in an encoding loop. As a simplified description, in an example, the encoding loop can include a source encoder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and (one or more) reference pictures), and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create sample data in a manner similar to how the (remote) decoder creates sample data. The reconstructed sample stream (sample data) is input into a reference picture memory (334). Since the decoding of the symbol stream results in a bit-accurate result that is independent of the decoder location (local or remote), the contents of the reference picture memory (334) are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction part of the encoder are exactly the same as the sample values "seen" by the decoder when using prediction during decoding. This basic principle of reference picture synchronization (and the drift caused when synchronization cannot be maintained, such as due to channel errors) is also used in some related technologies.
[0055] The operation of the "local" decoder (333) can be combined with, for example, Figure 2 The "remote" decoder described in detail for the video decoder (210) is identical. However, reference is also briefly made to Figure 2 , when symbols are available and the entropy encoder (345) and parser (220) are capable of losslessly encoding / decoding the symbols into an encoded video sequence, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and the parser (220), may not be fully implemented in the local decoder (333).
[0056] On the one hand, any decoder technology other than parsing / entropy decoding present in the decoder must also be present in the corresponding encoder in substantially the same functional form. Therefore, the present disclosure focuses on decoder operation. The description of encoder technology can be simplified because encoder technology is mutually inverse to the decoder technology described comprehensively. In some areas, a more detailed description is provided below.
[0057] During operation, in some examples, the source encoder (330) may perform motion compensated predictive coding that predictively encodes an input picture with reference to one or more previously encoded pictures in a video sequence designated as "reference pictures." In this manner, the encoding engine (332) encodes the differences between pixel blocks of the input picture and pixel blocks of (one or more) reference pictures that may be selected as prediction reference (s) for the input picture.
[0058] The local video decoder (333) may decode the encoded video data of the picture that may be designated as the reference picture based on the symbol created by the source encoder (330). The operation of the encoding engine (332) may be a lossy process. When the encoded video data may be decoded at the video decoder ( Figure 3 When the video sequence is decoded at a remote location (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder (333) replicates the decoding process that may be performed by the video decoder on the reference picture and may cause the reconstructed reference picture to be stored in the reference picture memory (334). In this way, the video encoder (303) may locally store a copy of the reconstructed reference picture that has common content (absent transmission errors) with the reconstructed reference picture to be obtained by the remote video decoder.
[0059] The predictor (335) can perform a prediction search on the encoding engine (332). That is, for a new picture to be encoded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that can serve as appropriate prediction references for the new picture. The predictor (335) can operate on a pixel block by pixel block basis to find a suitable prediction reference. In some cases, based on the search results obtained by the predictor (335), it can be determined that the input picture can have prediction references taken from multiple reference pictures stored in the reference picture memory (334).
[0060] The controller (350) can manage the encoding operations of the source encoder (330), including, for example, setting parameters and subgroup parameters for encoding video data.
[0061] The outputs of all the above functional units may be entropy encoded in an entropy encoder (345). The entropy encoder (345) converts the symbols generated by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, arithmetic coding, etc.
[0062] The transmitter (340) can buffer the encoded video sequence created by the entropy encoder (345) in preparation for transmission over a communication channel (360), which can be a hardware / software link to a storage device where the encoded video data will be stored. The transmitter (340) can combine the encoded video data from the video encoder (303) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0063] The controller (350) may manage the operation of the video encoder (303). During encoding, the controller (350) may assign a certain coded picture type to each coded picture, but this may affect the coding techniques that can be applied to the corresponding picture. For example, a picture may generally be assigned to any of the following picture types:
[0064] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other picture in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures.
[0065] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction, which uses at most one motion vector and a reference index to predict sample values for each block.
[0066] Bidirectional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra-frame prediction or inter-frame prediction, which uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-directional predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.
[0067] The source picture may typically be spatially subdivided into blocks of samples (e.g., blocks of 4×4 samples, blocks of 8×8 samples, blocks of 4×8 samples, or blocks of 16×16 samples) and coded block by block. These blocks may be predictively coded with reference to other (already coded) blocks, which are determined according to the coding allocation applied to the block's corresponding picture. For example, blocks of an I picture may be non-predictively coded or may be predictively coded (spatial prediction or intra prediction) with reference to already coded blocks of the same picture. Blocks of pixels of a P picture may be predictively coded by spatial prediction with reference to one previously coded reference picture or by temporal prediction. Blocks of a B picture may be predictively coded by spatial prediction with reference to one or two previously coded reference pictures or by temporal prediction.
[0068] The video encoder (303) may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T Rec. H.265. In operation, the video encoder (303) may perform various compression operations, including predictive encoding operations that exploit temporal and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard used.
[0069] In one aspect, the transmitter (340) may transmit additional data when transmitting the encoded video. The source encoder (330) may include such data as part of the encoded video sequence. The additional data may include time / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0070] The captured video may be taken as a plurality of source pictures (video pictures) in a temporal sequence. Intra-picture prediction (often simplified to intra-prediction) exploits spatial correlations in a given picture, while inter-picture prediction exploits (temporal or other) correlations between pictures. In one example, a particular picture being encoded / decoded is divided into blocks, and the particular picture being encoded / decoded is referred to as the current picture. When a block in the current picture is similar to a reference block in a reference picture that was previously encoded in the video and is still buffered, the block in the current picture may be encoded by a vector called a motion vector. The motion vector points to a reference block in a reference picture, and in the case where multiple reference pictures are used, the motion vector may have a third dimension that identifies the reference picture.
[0071] In some aspects, bidirectional prediction techniques may be used in inter-picture prediction. According to the bidirectional prediction technique, two reference pictures are used, such as a first reference picture and a second reference picture that are both before the current picture in the video in decoding order (but may be in the past and future in display order, respectively). A block in the current picture may be encoded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. The block may be predicted by a combination of the first reference block and the second reference block.
[0072] In addition, merge mode technology can be used in inter-picture prediction to improve coding efficiency.
[0073] According to some aspects of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a video picture sequence are divided into coding tree units (Coding Tree Unit, CTU) for compression, and the CTUs in the pictures have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Typically, a CTU includes three coding tree blocks (Coding Tree Block, CTB), which are a luminance CTB and two chrominance CTBs. Each CTU can be split into one or more coding units (Coding Unit, CU) in a quadtree. For example, a 64×64 pixel CTU can be split into a 64×64 pixel CU, or 4 32×32 pixel CUs, or 16 16×16 pixel CUs. In the example, each CU is analyzed to determine the prediction type of the CU, such as an inter-frame prediction type or an intra-frame prediction type. Depending on temporal and / or spatial predictability, the CU is split into one or more prediction units (Prediction Unit, PU). Typically, each PU includes a luma prediction block (PB) and two chroma PBs. On the one hand, the prediction operation in decoding (encoding / decoding) is performed in units of prediction blocks. Taking the luma prediction block as an example, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, and the like.
[0074] It should be noted that the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) can be implemented using any suitable technology. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) can be implemented using one or more integrated circuits. In one aspect, the video encoder (103) and the video encoder (303) and the video decoder (110) and the video decoder (210) can be implemented using one or more processors that execute software instructions.
[0075] Aspects of the present disclosure include techniques for improving signaling by using template matching costs.
[0076] Video coding has been widely used in many applications, such as broadcasting, video recording, and video streaming. Many early-stage video coding standards, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1, have been published and widely adopted in these video applications. On the one hand, a hybrid video codec may include various coding modules, such as intra-frame prediction, inter-frame prediction, transform coding, quantization, entropy coding, and post-loop filters. In the present disclosure, there is provided an improvement in signaling by using template matching costs to improve coding efficiency not only for intra-frame prediction but also for inter-frame prediction.
[0077] Intra-frame template matching prediction (also called intraTMP) is a special intra-frame prediction mode that copies the best prediction block, for example, from a reconstructed part of the current frame, whose L-shaped template matches the current template (e.g., the template of the current block). For a predefined search range, the encoder can search for the template that is most similar to the current template in the reconstructed part of the current frame, and take the corresponding block as the prediction block, the most similar template is associated with the corresponding block, and the current template is associated with the current block. The encoder then signals that the intraTMP mode is used, and the same prediction operation can be performed on the decoder side. It can be used in Figure 4 A matching block (or corresponding block) (402) is shown in FIG. 4 , and the matching block (or corresponding block) (402) serves as a matching region of the current CU (404).
[0078] like Figure 4 As shown, a prediction signal can be generated by matching an L-shaped causal neighborhood (or L-shaped template) of the current block (404) with another block in a predefined search area. For example, the predefined search area may include R1 (current CTU), R2 (upper left CTU), R3 (upper CTU), and R4 (left CTU).
[0079] In one aspect, the sum of absolute differences (SAD) is used as a cost function in the IntraTMP mode. Within each search region, the decoder can search for a template (406) of a block (402) having a minimum SAD relative to a current template (408) of a current block (404), and use the block with the minimum SAD as the corresponding block of the current block. The corresponding block of the current block can also be used as a prediction block (404) of the current block.
[0080] The sizes of all search regions (e.g., SearchRange_w, SearchRange_h) can be set to be proportional to the block size (e.g., BlkW, BlkH) of the current block. Therefore, a fixed number of SAD comparisons can be obtained in each pixel. For example, the dimensions of the search region (or search range) can be defined in Equation 1 and Equation 2 as follows: SearchRange_w = a * BlkW Formula (1) SearchRange_h = a * BlkH Formula (2) Where "a" is a constant that controls the trade-off between gain and complexity of the search process. In the example, the value of "a" is 5.
[0081] In one aspect, to speed up the template matching process, the search range of all search areas may be subsampled by a factor of 2. The reduction in the search range may result in a reduction in the number of template matching searches by 4. After the best match is found, a further refinement process may be performed. The refinement operation may be performed by performing a second template matching search with a reduced range around the best match. The reduced range may be limited to min(BlkW, BlkH) / 2.
[0082] The intra template matching tool can be enabled for CUs with width and height less than or equal to 64. The maximum CU size for intra template matching is configurable.
[0083] In one aspect, when decoder-side Intra Mode Derivation (DIMD) is not used in the current CU, the intra template matching prediction mode may be signaled at the CU level through a dedicated flag.
[0084] In one aspect, the method in the present disclosure can be used in any of the existing codecs mentioned above, such as H.264, H.265 / HEVC, H.266 / VVC, and AV1.
[0085] In the present disclosure, the template of the coding block (or the current template) may include spatially adjacent reconstructed samples of the coding block. The template may be, but is not limited to Figure 5 One of the template types shown. The size of the template can be, but is not limited to, W×n or n×H, where W and H are the block width and block height of the coding block, and n is a non-zero positive integer.
[0086] like Figure 5 As shown, various templates (or template types) of the current coding block (502) are provided. For example, the current coding block (502) may include a L The first template (or template type) (510), denoted as T a+l The second template (or template type) (512), denoted as T a The third template (or template type) (514) and the third template (or template type) denoted by T l The fourth template (or template type) (516). Figure 5 As shown, the fourth template type (516) may include a template area located on the left side of the current block. The third template type (514) may include a template area located on the upper side of the current block. The second template type (512) may include a first template area located on the left side of the current block and a second template area located on the top of the current block. The first template type (510) may include a first area located on the left side of the current block, a second area located on the upper side of the current block, and an upper left area located at the upper left corner of the current block and between the first area and the second area.
[0087] The template matching process may include several different processes. For example, the template matching process may include a search process to find a template (or reference template) that is most similar to the current template in the reconstructed portion of the current frame or reference frame. The template matching process may also include a matching process to calculate the distortion between the current template of the current block and the reference template of the prediction block in the reconstructed portion of the current frame or reference frame, and the reference template may be pointed to by a given motion vector (Motion Vector, MV) on the reference picture or a block vector (Block Vector, BV) on the reconstructed portion of the current frame.
[0088] In one aspect of the present disclosure, the template matching process may use filtered and / or unfiltered templates. For example, the template matching process may use a filtered current template and a filtered candidate reference template during the template matching process. In an example, the template matching process may use an unfiltered current template and an unfiltered candidate reference template during the template matching process. Similarity may be measured using a cost function that operates on the filtered / unfiltered templates. The cost function may include the sum of absolute differences (SAD), the mean absolute difference (MAD), the mean square error (MSE), and the sum of absolute transformed differences (SATD), etc.
[0089] In one aspect, a first method (or function-based prediction method) for intra-prediction or inter-prediction improvement is provided to reduce or minimize distortion between a current block and a prediction block of the current block from a corresponding reference picture by applying a nonlinear function or a linear function using the original prediction block as input. The nonlinear function or the linear function may be defined by a nonlinear formula or a linear formula, respectively. In an example, the linear function is defined as where n is a non-negative number, and p(x i ,y i ) is the position of the MV on the reference image (x i ,y i ) points to the predicted sample. Therefore, when i = 0, ..., n, p(x i ,y i ) may indicate a set of samples around the sample located (or indicated) by the MV, and may indicate a prediction block for the current block. Parameter α may be derived based on a template of the current block and a template of the prediction block (eg, by using a least squares method). i and β. In the example, the nonlinear function can be defined as Where m is a positive integer, such as 2 or 3.
[0090] The above template-based method can also be used for intra-frame prediction. In the case of intra-frame prediction, BV can be used instead of MV to indicate the position of the reconstructed reference sample in the reference picture, and BV can be derived by using a template matching search process or can be signaled in the bitstream. Parameter α i and β can still be derived through the current block template and the predicted block template.
[0091] In one aspect, a first syntax or first encoded information (e.g., a control flag of a first method) is signaled in a code stream (e.g., at a block level) to indicate whether the first method is applied to a current block. Alternatively, the value of the control flag may be inferred or inherited from another encoded block. In an example, a control flag of the first method associated with the current block is controlled by a control flag of the first method associated with one or more encoded blocks. In addition, a control flag may be derived at a coded block level to adaptively determine whether to apply the first method.
[0092] In one aspect, a second syntax or second encoded information is signaled in the codestream (e.g., at a block level) to indicate which template type is selected. For example, the second syntax indicates the selection of Figure 5 Which template type in .
[0093] In one aspect, some prediction models are applied to improve the prediction block. For example, a first model is derived based on a linear prediction function using a template type and / or a template size, and a second model is derived based on another nonlinear function or another linear function using the same template type and / or the same template size, or another template type and / or another template size. A third syntax or third encoded information can be signaled to select one of the prediction models. For example, a third syntax element indicates which model of the first model and the second model is selected.
[0094] In one aspect, the prediction model refers to a prediction model using a nonlinear function or a linear function.
[0095] In one aspect, a first syntax element or encoded information (such as a flag) is signaled to indicate a decision whether to apply the first method. The decision may be determined based on a plurality of template costs (e.g., two different template matching costs). When the flag (or first syntax element) is true, it indicates that the method with the least cost (or first method) is selected. For example, two different template matching costs are determined based on a first distortion between a current template (or a template of a current block) and a reference template (or a template of a reference block) for a first model and a second distortion between a current template and a reference template for a second model. One of the first model and the second model corresponding to the minimum cost is selected. In addition, a function of the selected model from the first model and the second model may be applied to the reference block to determine a prediction block for the current block. For example, samples of the prediction block of the current block may be defined as samples of the reference block, and the function of the selected model processes the samples of the reference block as Where p(x i ,y i ) are samples of the reference block.
[0096] The first syntax element or encoded information (eg, a flag) may be applicable to one or more color components. In an example, the flag (or first syntax element) is used for all color components (eg, Cb components and Cr components) of the block.
[0097] In the example, each color component has a corresponding flag.
[0098] In an example, the flag (or first syntax element) may only be used for at least one specific one of the color components of the block.
[0099] In an example, one cost is the distortion between the current template and the reference template of the specified model, and the other cost is the distortion between the filtered current template and the filtered reference template using the same model. The specified model can be defined by one or more of a linear prediction function or a nonlinear prediction function, a template type and / or a template size. In an example, the filtered current template and the filtered reference template are obtained by filtering the current reference template and the templated reference based on the filter. The filter can include a smoothing filter, a low pass filter, or any / or other suitable filter.
[0100] In the example, one cost is the distortion between the current template and the reference template, and the other cost is the distortion between the current template using mean removal and the reference template using mean removal. For example, the sample value of the sample of the current template minus the mean value of the sample of the current template. The sample value of the sample of the reference template minus the mean value of the sample value of the reference template.
[0101] In the example, the cost is a normalized cost. The normalized cost can be derived by dividing the distortion by a normalized value (such as template size). For example, first obtain the distortion between the current template and the reference template. The normalized cost is derived by dividing the distortion by the template size of the current template.
[0102] In an example, when the first syntax element (or flag) is used for (or applies to) all color components, distortion of all color components is measured. The first syntax element or flag may indicate that the first method is applied to all color components.
[0103] In an example, when the first syntax element (or flag) is used for all color components, distortion of at least one specific color component is measured.
[0104] In the example, one cost is not applying a linear model or a nonlinear model (e.g., using a pure / raw prediction signal p(x i ,y i ))), and the other cost is the distortion between the current template and the reference template when applying a specific predefined model.
[0105] In an example, one cost is the distortion between the current template and the reference template by applying a first predefined linear model or a first nonlinear model, and the other cost is the distortion between the current template and the reference template by applying a second predefined linear model or a second predefined nonlinear model.
[0106] In one aspect, a second syntax element or encoded information (e.g., a flag) is signaled to indicate a decision to select which prediction model from a plurality of models (e.g., the first and second models described above) in the first method. The decision may be determined based on template matching costs of the plurality of models (e.g., the two different template matching costs of the first model and the second model described above). Each template matching cost may be determined based on the corresponding model applied to the current template and the reference template. When the flag (or second syntax element) is true, it may indicate that a model with the smallest cost is selected from the plurality of models.
[0107] In an example, a first model is derived based on a linear prediction function / non-linear prediction function by using one template type and / or one template size, and a second model is derived based on another non-linear prediction function / another linear prediction function by using the same template type and / or the same template size or another template type and / or another template size. Two template matching costs are calculated based on a first distortion between a current template and a reference template under the first model and a second distortion between the current template and the reference template under the second model.
[0108] In another example, the current template is divided into a first group and a second group, and the first model is applied to the first group and the second model is applied to the second group. Therefore, the reference template can be divided into a first group and a second group, and the first model can be applied to the first group and the second model can be applied to the second group. Two template matching costs can be calculated based on a first distortion between the first group of the current template and the first group of the reference template and a second distortion between the second group of the current template and the second group of the reference template.
[0109] In one aspect, an option in a list of multiple options is selected by using the template matching costs of the multiple options. Each option may correspond to a corresponding template matching cost based on a corresponding model. In an example, a syntax element (or a third syntax element) or coding information is signaled, such as an index i. The index i may represent the i-th candidate in the list of multiple options. The list of options may be constructed in ascending order based on the template matching cost of each option. The list size may be equal to the number of multiple options.
[0110] In an example, the syntax element (or the third syntax element) is used for (or applies to) all color components.
[0111] In an example, each color component has a corresponding syntax element.
[0112] In an example, the syntax element (or the third syntax element) may be used only for one or more specific color components. For example, the syntax element is used for at least one specific color component.
[0113] In an example, the plurality of options include, but are not limited to, a model of the first method. In an example, the first method with the first model and the first method with the second model are applied.
[0114] In an example, the template matching cost is a normalized cost. The normalized cost can be derived by dividing the distortion by a normalized value (eg, template size).
[0115] In an example, when the syntax element (or the third syntax element) is used for all color components, distortion of all color components is measured.
[0116] In an example, when the syntax element (or the third syntax element) is used for all color components, the distortion of at least one specific color component is measured.
[0117] Figure 6 A flow chart outlining a process (600) according to some aspects of the present disclosure is shown. The process (600) may be used in a video decoder. In various aspects, the process (600) is performed by a processing circuit, such as a processing circuit that performs the functions of a video decoder (110), a processing circuit that performs the functions of a video decoder (210), etc. In some aspects, the process (600) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (600). The process starts at (S601) and proceeds to (S610).
[0118] In (S610), a code stream is received. The code stream includes a first syntax element, the first syntax element indicating whether to apply a function-based prediction method based on template matching prediction to reconstruct a current block. The function-based prediction method is applied to a prediction model, the prediction model is associated with a template type of a template of the current block, and the prediction model includes a function applied to a reference block of the current block. The function includes parameters derived based on the template of the current block and the template of the reference block.
[0119] In (S620), when the first syntax element indicates that a function-based prediction method is applied to reconstruct the current block, a prediction model is determined from multiple candidate prediction models based on multiple TM costs between a template of the current block and a template of a reference block according to multiple candidate prediction models associated with the function-based prediction method.
[0120] At (S630), a prediction block of the current block is determined as a reference block processed by a function of a prediction model.
[0121] In an example, a function of each candidate prediction model in a plurality of candidate prediction models is applied to a template of a current block and a template of a reference block to obtain a processed template of the current block and a processed template of the reference block. Each TM cost in a plurality of TM costs is determined based on the processed template of the current block and the processed template of the reference block of the corresponding candidate prediction model. A prediction model corresponding to a minimum TM cost in a plurality of TM costs is determined from the plurality of candidate prediction models.
[0122] In the example, the template types include (i) a first template type, the first template type includes a first template area located on the left side of the current block, (ii) a second template type, the second template type includes a second template area located on the upper side of the current block, (iii) a third template type, the third template type includes the first template area located on the left side of the current block and the second template area located on the upper side of the current block, and (iv) a fourth template type, the fourth template type includes the first template area located on the left side of the current block, the second template area located on the upper side of the current block, and an upper left template area located at the upper left corner of the current block and set between the first template area and the second template area.
[0123] In an example, a processed sample of a template of a current block and a processed sample of a template of a reference block are generated based on a first candidate prediction model among a plurality of candidate prediction models. A first TM cost among a plurality of TM costs is determined based on the processed sample of the template of the current block and the processed sample of the template of the reference block. Samples of the template of the current block and samples of the template of the reference block are filtered. Processed filtered samples of the template of the current block and processed filtered samples of the template of the reference block are generated based on the first candidate prediction model. A second TM cost among a plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
[0124] In an example, an average value of the template samples of the current block is subtracted from each sample of the samples of the template of the current block to generate an adjusted sample of the template of the current block. An average value of the template samples of the reference block is subtracted from each sample of the samples of the template of the reference block to generate an adjusted sample of the template of the current block. A third TM cost of the plurality of TM costs is determined based on the adjusted sample of the template of the current block and the adjusted sample of the template of the reference block.
[0125] In an example, a first TM cost among a plurality of TM costs is determined based on a template sample of a current block and a template sample of a reference block. A processed sample of a current block template and a processed sample of a reference block template are generated based on a first candidate prediction model among a plurality of candidate prediction models. A second TM cost among a plurality of TM costs is determined based on a processed sample of a template of the current block and a processed sample of a template of the reference block.
[0126] In an example, a TM cost among a plurality of TM costs is determined. A TM cost among a plurality of TM costs is indicated by a second syntax element in a bitstream. A prediction model corresponding to the TM cost determined among the plurality of TM costs is determined from a plurality of candidate prediction models. A function of the prediction model is applied to process samples of a reference block. A prediction block of a current block is determined based on samples of the reference block processed by the function of the prediction model.
[0127] In an example, each of the plurality of TM costs is normalized by dividing each of the plurality of TM costs by a template size corresponding to the respective TM cost. The plurality of normalized TM costs are reordered in ascending order based on the normalized TM costs. A candidate list is generated based on the plurality of reordered normalized TM costs.
[0128] In an example, a third syntax element is received in a bitstream, the third syntax element indicating which one of a plurality of reordered normalized TM costs is selected. The third syntax element is associated with at least one of a Cb component or a Cr component of a current block. A prediction model corresponding to the TM cost indicated by the third syntax element is determined from a plurality of candidate prediction models.
[0129] Then, the process proceeds to (S699) and ends.
[0130] The process (600) may be adjusted appropriately. One (or more) steps in the process (600) may be modified and / or omitted. One (or more) additional steps may be added. Any suitable implementation order may be used.
[0131] Figure 7 A flow chart outlining a process (700) according to some aspects of the present disclosure is shown. The process (700) may be used in a video encoder. In various aspects, the process (700) is performed by a processing circuit, such as a processing circuit that performs the functions of a video encoder (103), a processing circuit that performs the functions of a video encoder (303), etc. In some aspects, the process (700) is implemented in software instructions, so when the processing circuit executes the software instructions, the processing circuit performs the process (700). The process starts at (S701) and proceeds to (S710).
[0132] At (S710), according to intraTMP, a reference block of the current block is determined based on a template of the current block and a template of a reference block.
[0133] At (S720), a prediction model is determined from a plurality of candidate prediction models based on a plurality of TM costs between a template of a current block and a template of a reference block according to a plurality of candidate prediction models. The prediction model is associated with a template type of the template of the current block, and the prediction model includes a function applied to a reference block of the current block. The function includes parameters derived based on the template of the current block and the template of the reference block.
[0134] At (S730), a prediction block of the current block is determined based on the prediction model.
[0135] At (S740), the current block is encoded into a code stream based on the prediction block, and a syntax element is encoded into the code stream. The syntax element indicates whether a function-based prediction method including a prediction model is applied to reconstruct the current block.
[0136] In the example, a function of each candidate prediction model in a plurality of candidate prediction models is applied to a template of a current block and a template of a reference block to obtain a processed template of the current block and a processed template of the reference block. Based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model, each TM cost in a plurality of TM costs is determined. A prediction model corresponding to a minimum TM cost in a plurality of TM costs is determined from a plurality of candidate prediction models.
[0137] In an example, a processed sample of a template of a current block and a processed sample of a template of a reference block are generated based on a first candidate prediction model among a plurality of candidate prediction models. A first TM cost among a plurality of TM costs is determined based on the processed sample of the template of the current block and the processed sample of the template of the reference block and the processed sample of the template of the reference block. Samples of the template of the current block and samples of the template of the reference block are filtered. Processed filtered samples of the template of the current block and processed filtered samples of the template of the reference block are generated based on the first candidate prediction model. A second TM cost among a plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
[0138] Then, the process proceeds to (S799) and ends.
[0139] The process (700) may be adjusted appropriately. One (or more) steps in the process (700) may be modified and / or omitted. One (or more) additional steps may be added. Any suitable implementation order may be used.
[0140] In one aspect, a method for processing visual media data includes processing a bitstream of the visual media data according to a format rule. For example, the bitstream may be a bitstream decoded / encoded using any of the decoding and / or encoding methods described herein. The format rule may specify one or more constraints on the bitstream and / or one or more processes performed by a decoder and / or encoder.
[0141] In an example, a code stream includes a first syntax element indicating whether a prediction method based on a nonlinear function is applied to process a current block based on template matching prediction. The prediction method based on a nonlinear function applies a prediction model, the prediction model is associated with a template type of a template of the current block, and the prediction model includes a nonlinear function applied to a reference block of the current block. The nonlinear function includes parameters derived based on the template of the current block and the template of the reference block. The format rule stipulates that when the first syntax element indicates that the prediction method based on a nonlinear function is applied to the current block, multiple template matching (TM) costs between the template of the current block and the template of the reference block are determined according to multiple candidate prediction models associated with the prediction method based on the nonlinear function. The format rule stipulates that a prediction model is determined from multiple candidate prediction models based on multiple TM costs. For example, the prediction model corresponds to a TM cost selected from multiple TM costs. The format rule stipulates that the nonlinear function of the prediction model is applied to the sample of the processing reference block. The format rule stipulates that the prediction block of the current block is determined as a reference block processed by the nonlinear function of the prediction model.
[0142] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. Figure 8 A computer system (800) suitable for implementing certain aspects of the disclosed subject matter is shown.
[0143] Computer software may be encoded using any suitable machine code or computer language, which may be assembled, compiled, linked or similarly constructed to create code comprising instructions that may be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode, etc.
[0144] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, IoT devices, etc.
[0145] Figure 8 The components of the computer system (800) shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing aspects of the present disclosure. Nor should the configuration of components be interpreted as having any dependency or requirement related to any one component or combination of components shown in the exemplary aspects of the computer system (800).
[0146] The computer system (800) may include certain human interface input devices. Such human interface input devices may be responsive to input from one or more human users, for example, through: tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, clapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to a person's conscious input, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video), etc.
[0147] The human-machine interface input device may include one or more of the following (only one of each is shown): keyboard (801), mouse (802), touchpad (803), touch screen (810), data gloves (not shown), joystick (805), microphone (806), scanner (807), camera (808).
[0148] The computer system (800) may also include certain human interface output devices. Such human interface output devices may stimulate one or more senses of a human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through a touch screen (810), a data glove (not shown), or a joystick (805), but may also be a tactile feedback device that is not an input device), audio output devices (e.g., speakers (809), headphones (not shown)), visual output devices (e.g., screens (810) including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities, some of which are capable of outputting two-dimensional visual outputs or outputs of more than three dimensions through devices such as stereoscopic picture output, virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).
[0149] The computer system (800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM / RW (820) with CD / DVD or similar media (821), thumb drives (822), removable hard drives or solid-state drives (823), traditional magnetic media such as tapes and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown), and the like.
[0150] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0151] The computer system (800) may also include an interface (854) to one or more communication networks (855). The network may be, for example, a wireless network, a wired network, an optical network. The network may also be a local network, a wide area network, a metropolitan area network, a vehicle and industrial network, a real-time network, a delay-tolerant network, and the like. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including Global System for Mobile communications (GSM), 3G, 4G, 5G, Long-Term Evolution (LTE), etc., television wired or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television, vehicle and industrial networks including CAN buses, and the like. Some networks typically require an external network interface adapter connected to some common data port or peripheral bus (849) (e.g., a USB port of the computer system (800)); other network interfaces are typically integrated into the kernel of the computer system (800) by connecting to the system bus described below (connected to an Ethernet interface in a PC computer system or to a cellular network interface in a smartphone computer system). The computer system (800) may communicate with other entities using any of these networks. Such communications may be one-way receive only (e.g., broadcast television), one-way send only (e.g., CANBus to certain CANBus devices), or bidirectional, such as connecting to other computer systems using a local area network or wide area network digital network. As described above, certain protocols and protocol stacks may be used on each of these networks and network interfaces.
[0152] The above-mentioned human-machine interface devices, human-accessible storage devices and network interfaces may be connected to the kernel (840) of the computer system (800).
[0153] The kernel (840) may include one or more central processing units (CPUs) (841), graphics processing units (GPUs) (842), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (843), hardware accelerators for certain tasks (844), graphics adapters (850), and the like. These devices, along with read-only memory (ROM) (845), random access memory (846), and internal mass storage (847) such as internal non-user accessible hard drives, SSDs, and the like, may be connected via a system bus (848). In some computer systems, the system bus (848) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, and the like. Peripheral devices may be connected directly to the kernel's system bus (848) or to the kernel's system bus (848) via a peripheral device bus (849). In one example, a screen (810) may be connected to a graphics adapter (850). The architecture of the peripheral bus includes PCI, USB, and the like.
[0154] The CPU (841), GPU (842), FPGA (843) and accelerator (844) can execute certain instructions, which together can constitute the above-mentioned computer code. The computer code can be stored in ROM (845) or RAM (846). Transition data can also be stored in RAM (846), while permanent data can be stored in, for example, internal mass storage (847). Fast storage and retrieval of any memory device can be achieved by using a cache memory, which can be closely associated with one or more CPUs (841), GPUs (842), mass storage (847), ROM (845), RAM (846), etc.
[0155] The computer readable medium may have thereon computer code for performing various computer-implemented operations. The medium and computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and computer code may be of a type well known and available to those skilled in the art of computer software.
[0156] As a non-limiting example, a computer system having an architecture structure (800), in particular a kernel (840), can provide functionality due to the execution of software contained in one or more tangible computer-readable media by one (or more) processors (including CPUs, GPUs, FPGAs, accelerators, etc.). Such computer-readable media can be media associated with a user-accessible mass storage as described above, as well as certain non-temporary memories of the kernel (840), such as a kernel internal mass storage (847) or ROM (845). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the kernel (840). Depending on specific needs, the computer-readable medium may include one or more storage devices or chips. The software can enable the kernel (840), in particular the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform specific processes or specific parts of specific processes described herein, including defining data structures stored in RAM (846) and modifying such data structures according to processes defined by the software. Additionally or alternatively, the computer system provides functionality due to logic hardwired or otherwise embodied in a circuit (e.g., accelerator (844)), which circuit can operate in place of or in conjunction with software to perform a particular process or a particular portion of a particular process described herein. Where appropriate, references to software may include logic and vice versa. Where appropriate, references to computer-readable media may include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0157] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, references to: at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C are intended to include only A, only B, only C, or any combination thereof. References to one of A or B and one of A and B are intended to include A or B or (A and B). Where applicable, such as when the elements are not mutually exclusive, the use of "one of" does not exclude any combination of elements.
[0158] Although the present disclosure has described a number of exemplary embodiments, there are modifications, permutations, and various replacement equivalents that fall within the scope of the present disclosure. It should therefore be understood that those skilled in the art will be able to design a variety of systems and methods that, although not explicitly shown or described in the present disclosure, embody the principles of the present disclosure and therefore fall within the spirit and scope of the present disclosure.
Claims
1. A method for processing visual media data, the method comprising: Processing the code stream of the visual media data according to the format rules, wherein: The code stream includes a first syntax element, the first syntax element indicating whether to apply a prediction method based on a nonlinear function to process a current block based on template matching prediction, the prediction method based on the nonlinear function applies a prediction model, the prediction model is associated with a template type of a template of the current block, the prediction model includes a nonlinear function applied to a reference block of the current block, the nonlinear function includes parameters derived based on the template of the current block and the template of the reference block; and The formatting rules state that: When the first syntax element indicates that the prediction method based on the nonlinear function is applied to the current block, determining a plurality of template matching TM costs between the template of the current block and the template of the reference block according to a plurality of candidate prediction models associated with the prediction method based on the nonlinear function, determining the prediction model from the plurality of candidate prediction models based on the plurality of TM costs, applying the nonlinear function of the prediction model to process samples of the reference block, and A prediction block of the current block is determined as the reference block processed by the nonlinear function of the prediction model.
2. The method according to claim 1, wherein: The formatting rules state that: Applying a function of each candidate prediction model of the plurality of candidate prediction models to the template of the current block and the template of the reference block to obtain a processed template of the current block and a processed template of the reference block; Determine each TM cost of the plurality of TM costs based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model; as well as A prediction model corresponding to a minimum TM cost among the plurality of TM costs is determined from the plurality of candidate prediction models.
3. The method according to claim 1 or 2, wherein: The formatting rules state that: Generate a processed sample of the template of the current block and a processed sample of the template of the reference block based on a first candidate prediction model among the plurality of candidate prediction models; determining a first TM cost of the plurality of TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block; filtering samples of the template of the current block and samples of the template of the reference block; Generate a processed filtered sample of the template of the current block and a processed filtered sample of the template of the reference block based on the first candidate prediction model; as well as A second TM cost of the plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
4. A video encoding method, the method comprising: Predicting intraTMP according to intra-frame template matching, determining the reference block of the current block based on the template of the current block and the template of the reference block; Determining a prediction model from a plurality of candidate prediction models based on a plurality of TM costs between the template of the current block and the template of the reference block, the prediction model being associated with a template type of the template of the current block, the prediction model comprising a function applied to the reference block of the current block, the function comprising parameters derived based on the template of the current block and the template of the reference block; Determine a prediction block for the current block based on the prediction model; as well as The current block is encoded into a code stream based on the prediction block, and a syntax element is encoded into the code stream, the syntax element indicating whether a function-based prediction method including the prediction model is applied to reconstruct the current block.
5. The method according to claim 4, wherein: Determining the prediction model from the plurality of candidate prediction models further comprises: Applying a function of each candidate prediction model of the plurality of candidate prediction models to the template of the current block and the template of the reference block to obtain a processed template of the current block and a processed template of the reference block; Determine each TM cost of the plurality of TM costs based on the processed template of the current block and the processed template of the reference block corresponding to each candidate prediction model; and A prediction model corresponding to a minimum TM cost among the plurality of TM costs is determined from the plurality of candidate prediction models.
6. The method according to claim 4 or 5, wherein: Determining the prediction model comprises: Generate a processed sample of the template of the current block and a processed sample of the template of the reference block based on a first candidate prediction model among the plurality of candidate prediction models; determining a first TM cost of the plurality of TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block; filtering samples of the template of the current block and samples of the template of the reference block; Generate a processed filtered sample of the template of the current block and a processed filtered sample of the template of the reference block based on the first candidate prediction model; and A second TM cost of the plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
7. A video decoding device, the device comprising: A processing circuit, wherein the processing circuit is configured as: Receive a code stream, the code stream comprising a first syntax element, the first syntax element indicating whether to apply a function-based prediction method to reconstruct a current block based on template matching prediction, the function-based prediction method applies a prediction model, the prediction model is associated with a template type of a template of the current block, the prediction model comprises a function applied to a reference block of the current block, the function comprises parameters derived based on the template of the current block and the template of the reference block; When the first syntax element indicates that the function-based prediction method is applied to reconstruct the current block, determining the prediction model from a plurality of candidate prediction models associated with the function-based prediction method based on a plurality of template matching TM costs between a template of the current block and a template of the reference block; as well as A prediction block of the current block is determined as the reference block processed by the function of the prediction model.
8. The device according to claim 7, wherein: The processing circuit is configured to: Applying a function of each candidate prediction model of the plurality of candidate prediction models to the template of the current block and the template of the reference block to obtain a processed template of the current block and a processed template of the reference block; Determine each TM cost of the plurality of TM costs based on the processed template of the current block and the processed template of the reference block corresponding to each of the candidate prediction models; as well as The prediction model corresponding to the minimum TM cost among the plurality of TM costs is determined from the plurality of candidate prediction models.
9. The device according to claim 7 or 8, wherein: The template types include (i) a first template type, which includes a first template area located on the left side of the current block, (ii) a second template type, which includes a second template area located on the upper side of the current block, (iii) a third template type, which includes the first template area located on the left side of the current block and the second template area located on the upper side of the current block, and (iv) a fourth template type, which includes the first template area located on the left side of the current block, the second template area located on the upper side of the current block, and an upper left template area located at the upper left corner of the current block and set between the first template area and the second template area.
10. The apparatus according to any one of claims 7 to 9, wherein: The processing circuit is configured to: Generate a processed sample of the template of the current block and a processed sample of the template of the reference block based on a first candidate prediction model among the plurality of candidate prediction models; determining a first TM cost of the plurality of TM costs based on the processed samples of the template of the current block and the processed samples of the template of the reference block; filtering samples of the template of the current block and samples of the template of the reference block; Generate a processed filtered sample of the template of the current block and a processed filtered sample of the template of the reference block based on the first candidate prediction model; as well as A second TM cost of the plurality of TM costs is determined based on the processed filtered samples of the template of the current block and the processed filtered samples of the template of the reference block.
11. The apparatus according to any one of claims 7 to 9, wherein: The processing circuit is configured to: subtracting an average value of the samples of the template of the current block from each of the samples of the template of the current block to generate adjusted samples of the template of the current block; subtracting an average value of the samples of the template of the reference block from each of the samples of the template of the reference block to generate adjusted samples of the template of the current block; and A third TM cost of the plurality of TM costs is determined based on the adjusted samples of the template of the current block and the adjusted samples of the template of the reference block.
12. The apparatus according to any one of claims 7 to 9, wherein: The processing circuit is configured to: determining a first TM cost of the plurality of TM costs based on samples of the template of the current block and samples of the template of the reference block; Generate a processed sample of the template of the current block and a processed sample of the template of the reference block based on a first candidate prediction model among the multiple candidate prediction models; as well as A second TM cost among the plurality of TM costs is determined based on the processed samples of the template of the current block and the processed samples of the template of the reference block.
13. The apparatus according to any one of claims 7 to 9, wherein: The processing circuit is configured to: determining a TM cost among the plurality of TM costs indicated by a second syntax element in the codestream; determining a prediction model corresponding to the determined TM cost among the plurality of TM costs from the plurality of candidate prediction models; applying the function of the prediction model to process samples of the reference block; and The prediction block of the current block is determined based on samples of the reference block processed by the function of the prediction model.
14. The apparatus according to any one of claims 7 to 13, wherein: The processing circuit is configured to: normalizing each TM cost of the plurality of TM costs by dividing the corresponding TM cost by a template size corresponding to the corresponding TM cost; reordering the plurality of normalized TM costs in ascending order based on the normalized TM costs; as well as A candidate list is generated based on a plurality of re-ranked normalized TM costs.
15. The device according to claim 14, wherein: The processing circuit is configured to: receiving a third syntax element in the code stream, the third syntax element indicating which one of the plurality of reordered normalized TM costs is selected, the third syntax element being associated with at least one of a Cb component or a Cr component of the current block; as well as The prediction model corresponding to the TM cost indicated by the third syntax element is determined from the plurality of candidate prediction models.