A method and an apparatus for video decoding
The improved TIMD method in video coding addresses inefficiencies by using multiple distortion metrics and dynamic template/candidate list formation, enhancing decoding accuracy and efficiency.
Patent Information
- Application Number
- PCT/EP2025/054308
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-26
- Filing Date
- 2025-02-18
- Publication Date
- 2025-10-02
AI Technical Summary
The existing Template-based Intra Mode Derivation (TIMD) technique in video coding is sub-optimal due to its reliance on a single distortion metric, a fixed template size, and a limited intra-prediction mode candidate list, leading to inefficient prediction in video decoding.
An improved TIMD method that determines intra-prediction modes by computing distortions in a template region using multiple distortion metrics and dynamically forming the template region and candidate list based on information extracted from the bitstream and inferred by the decoder.
Enhances the accuracy and efficiency of video decoding by optimizing the intra-prediction process, reducing the need for signaling intra-prediction modes and improving the correlation between the template region and the current block content.
Smart Images

Figure EP2025054308_02102025_PF_FP_ABST
Abstract
Description
[0001] A METHOD AND AN APPARATUS FOR VIDEO DECODING
[0002] Technical Field
[0003] The present solution generally relates to a method and an apparatus and a computer program product for video coding and decoding.
[0004] Background
[0005] This section is intended to provide a background or context to the invention that is recited in the claims. The description herein may include concepts that could be pursued but are not necessarily ones that have been previously conceived or pursued. Therefore, unless otherwise indicated herein, what is described in this section is not prior art to the description and claims in this application and is not admitted to be prior art by inclusion in this section.
[0006] Template-based intra mode derivation (TIMD) is a technique where an intraprediction mode is derived at the decoder side by selecting a mode with minimum distortion from an intra-prediction mode candidate list, where distortions are computed using a template of the block. The candidate list may consist of a list of Most Probable Modes (MPM) extracted from neighboring blocks.
[0007] Summary
[0008] The aim of the present solution is to provide an improved process for the TIMD.
[0009] The scope of protection sought for various embodiments of the invention is set out by the independent claims. The embodiments and features, if any, described in this specification that do not fall under the scope of the independent claims are to be interpreted as examples useful for understanding various embodiments of the invention.
[0010] Various aspects include a method, an apparatus and a computer readable medium comprising a computer program stored therein, which are characterized by what is stated in the independent claims. Various embodiments are disclosed in the dependent claims.
[0011] According to a first aspect, there is provided an apparatus for decoding encoded samples of blocks of video sample data, wherein for a block of video sample data the apparatus comprises means for deriving information to configure a template-based intra-prediction mode derivation process, wherein for implementing the configured template-based intra-prediction mode derivation process the apparatus further comprises means for computing distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra-prediction mode candidate list; means for determining one or more intra mode predictors by the determined one or more intra-prediction modes; and means for computing, based on the intra mode predictors, a prediction for the samples of the video sample data.
[0012] According to a second aspect, there is provided a method for decoding encoded samples of blocks of video sample data, wherein for a block of video sample data the method comprises deriving information to configure a template-based intra-prediction mode derivation process, wherein the configured template-based intra-prediction mode derivation process comprises computing distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra- prediction mode candidate list; determining one or more intra mode predictors by the determined one or more intra-prediction modes; and computing, based on the intra mode predictors, a prediction for the samples of the video sample data.
[0013] According to a third aspect, there is provided an apparatus for decoding encoded samples of blocks of video sample data, the apparatus comprising at least one processor, memory including computer program code, the memory and the computer program code configured to, with the at least one processor, cause the apparatus to derive information to configure a template-based intra- prediction mode derivation process, wherein for implementing the configured template-based intra-prediction mode derivation process the apparatus is further caused to compute distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra- prediction mode candidate list; to determine one or more intra mode predictors by the determined one or more intra-prediction modes; and to compute, based on the intra mode predictors, a prediction for the samples of the video sample data.
[0014] According to a fourth aspect, there is provided computer program product for decoding encoded samples of blocks of video sample data, the computer program product comprising computer program code configured to, when executed on at least one processor, cause an apparatus or a system to derive information to configure a template-based intra-prediction mode derivation process, wherein for implementing the configured template-based intra- prediction mode derivation process the apparatus is further caused to compute distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra-prediction mode candidate list; to determine one or more intra mode predictors by the determined one or more intra-prediction modes; and to compute, based on the intra mode predictors, a prediction for the samples of the video sample data.
[0015] According to an embodiment, the derived information is derived from an encoded bitstream.
[0016] According to an embodiment, the derived information is used to determine a distortion metric among a plurality of distortion metrics.
[0017] According to an embodiment, the distortion metric is one of the following: Sum of Absolute Differences, Sum of Absolute Transformed Differences, Sum of Squared Errors, Mean-removed Sum of Absolute Differences, Mean-removed Sum of Absolute Transformed Differences, Mean-removed Sum of Squared Errors.
[0018] According to an embodiment, the derived information is used to determine a set of reconstructed samples to form the template region.
[0019] According to an embodiment, the derived information is used to determine which intra-prediction mode candidate list of at least two intra-prediction mode candidate lists is used. According to an embodiment, the computer program product is embodied on a non-transitory computer readable medium.
[0020] Description of the Drawings
[0021] In the following, various embodiments will be described in more detail with reference to the appended drawings, in which
[0022] Fig. 1 shows a simplified example of an encoding process;
[0023] Fig. 2 shows a simplified example of a decoding process;
[0024] Figs. 3a, b show simplified examples of template and its reference samples;
[0025] Fig. 4 is a flowchart illustrating a method according to an embodiment ; and
[0026] Fig. 5 shows an apparatus according to an embodiment.
[0027] Description of Example Embodiments
[0028] Certain abbreviations that may be found in the description are herewith defined as follows:
[0029] 2D Two-Dimensional
[0030] 3D Three-Dimensional
[0031] AMVP Advanced Motion Vector Prediction
[0032] AV1 AO Media Video 1
[0033] AVC Advanced Video Coding
[0034] CABAC Context-adaptive binary arithmetic coding
[0035] CDMA Code Division Multiple Access
[0036] CTU Coding Tree unit
[0037] CU Coding Unit
[0038] DC Direct Current
[0039] DCT Discrete Cosine Transform
[0040] DIMD Decoder-side Intra Mode Derivation
[0041] DPB Decoder Picture Buffer
[0042] DSP Digital Signal Processor ECM Enhanced Compression Model
[0043] FDMA Frequency Division Multiple Access
[0044] GSM Global System for Mobile communications
[0045] HEVC High Efficiency Video Coding
[0046] HoG Histogram of Gradients
[0047] IEC International Electrotechnical Commission
[0048] IMS Instant Messaging Service
[0049] IPM Intra Prediction Mode
[0050] ISO Internal Organization for Standardization
[0051] ISOBMFF ISO Base Media File Format
[0052] ITU-T Telecommunications Standardization Sector of International
[0053] Telecommunication Union
[0054] JCT-VC Joint Collaborative Team - Video Coding
[0055] JVT Joint Video Team
[0056] LCU Largest Coding Unit
[0057] MCP Motion-Compensated Prediction
[0058] MHoG Merged Histogram of Gradients
[0059] MIMD Merged Intra Mode derivation
[0060] MMS Multimedia Messaging Service
[0061] MPEG Moving Picture Experts Group
[0062] MPG Most Probable Mode
[0063] MR- Mean-Removed-
[0064] MRL Multiple Reference Line
[0065] MVC Multiview Video Coding
[0066] MV Multiview
[0067] NAL Network Abstraction Layer
[0068] PC Personal Computer
[0069] PU Prediction Unit
[0070] REXT Range Extensions
[0071] SAD Sum of Absolute Differences
[0072] SATD Sum of Absolute Transformed Differences
[0073] SHVC Scalable High efficiency Video Coding
[0074] SMS Short Messaging Service
[0075] SNR Signal-to-Noise Ratio
[0076] SPGM Single-Path Global Modulation
[0077] SSE Sum of Squared Errors
[0078] SVC Scalable Video Coding TDMA Time Divisional Multiple Access
[0079] TCP-IP Transmission Control Protocol - Internet Protocol
[0080] TIMD Template-based Inter Mode Derivation
[0081] Til Transform Unit
[0082] TMRL Template-based Multiple Reference Line
[0083] TV Television
[0084] UMTS Universal Mobile Telecommunications System
[0085] VCEG Video Coding Experts Group
[0086] WC Versatile Video Coding
[0087] The following description and drawings are illustrative to discuss embodiments of the present solution with examples. The specific details are provided for understanding purposes. However, in certain instances, well-known or conventional details are not described in order to avoid obscuring the description. Reference in this specification to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the disclosure.
[0088] The present embodiments relate to an adaptive TIMD.
[0089] Before describing the embodiments further, a brief reference to evolution of video coding standardization is given. The present embodiments are suited within the context of next generation video coding standardization, e.g., H.267 video coding standard and the ECM (Enhanced Compression Model) exploration.
[0090] A video codec comprises an encoder and a decoder. The encoder transforms the input video into a compressed representation suited for storage / transmission. The decoder can un-compress the compressed video representation back into a viewable form. The encoder may discard some information in the original video sequence in order to represent the video in a more compact form (i.e. , at lower bitrate).
[0091] Figure 1 shows a simple example of a principle of an encoding process for 2D pictures, and Figure 2 shows a simple example of a principle of a decoding process for 2D pictures. In Figure 1 , the following have been illustrated: - an image to be encoded (ln);
[0092] - a predicted representation of an image block (P'n);
[0093] - a prediction error signal (Dn);
[0094] - a reconstructed prediction error signal (D'n);
[0095] - a preliminary reconstructed image (l'n);
[0096] - a final reconstructed image (R'n);
[0097] - a transform (T) and inverse transform (T-1);
[0098] - a quantization (Q) and inverse quantization (Q-1);
[0099] - entropy encoding (E);
[0100] - a reference frame memory (RFM);
[0101] - inter prediction (Pinter);
[0102] - intra prediction (Pintra);
[0103] - mode selection (MS), and
[0104] - filtering (F).
[0105] In Figure 2 the following have been illustrated:
[0106] - a predicted representation of an image block (P'n);
[0107] - a reconstructed prediction error signal (D'n);
[0108] - a preliminary reconstructed image (l'n);
[0109] - a final reconstructed image (R'n); an inverse transform (T-1);
[0110] - an inverse quantization (Q-1);
[0111] - an entropy decoding (E-1);
[0112] - a reference frame memory (RFM);
[0113] - a prediction (either inter or intra) (P);
[0114] - and filtering (F).
[0115] An elementary unit for the input to an encoder and the output of a decoder, respectively, in most cases is a picture (also referred to as “an image”). A picture given as an input to an encoder may also be referred to as a source picture, and a picture decoded by a decoded may be referred to as a decoded picture or a reconstructed picture.
[0116] The source and decoded pictures are each comprised of one or more sample arrays, such as one of the following sets of sample arrays:
[0117] - Luma (Y) only (monochrome).
[0118] - Luma and two chroma (YCbCr or YCgCo).
[0119] - Green, Blue and Red (GBR, also known as RGB). Arrays representing other unspecified monochrome or tristimulus color samplings (for example, YZX, also known as XYZ).
[0120] A picture may be defined to be either a frame or a field. A frame comprises a matrix of luma samples and possibly the corresponding chroma samples. A field is a set of alternate sample rows of a frame and may be used as encoder input, when the source signal is interlaced. Chroma sample arrays may be absent (and hence monochrome sampling may be in use) or chroma sample arrays may be subsampled when compared to luma sample arrays.
[0121] The Advanced Video Coding standard (which may be abbreviated AVC or H.264 / AVC) was developed by the Joint Video Team (JVT) of the Video Coding Experts Group (VCEG) of the Telecommunications Standardization Sector of International Telecommunication Union (ITU-T) and the Moving Picture Experts Group (MPEG) of International Organization for Standardization (ISO) I International Electrotechnical Commission (IEC). There have been multiple versions of the H.264 / AVC standard, each integrating several extensions or features to the specification. These extensions include Scalable Video Coding (SVC) and Multiview Video Coding (MVC).
[0122] The High Efficiency Video Coding standard (which may be abbreviated HEVC or H.265 / HEVC) was developed by the Joint Collaborative Team - Video Coding (JCT-VC) of VCEG and MPEG. Extensions to H.265 / HEVC include scalable, multiview, three-dimensional, and fidelity range extensions, which may be referred to as SHVC, MV-HEVC, 3D-HEVC, and REXT, respectively.
[0123] Versatile Video Coding (which may be abbreviated WC, H.266, or H.266A / VC) is a video compression standard developed as the successor to HEVC. WC is specified in ITU-T Recommendation H.266 and equivalently in ISO / IEC 23090-3, which is also referred to as MPEG-I Part 3.
[0124] ECM was developed by JVET (Joint Video Experts Team) of ITU-T VCEG and ISO / IEC MPEG, to provide a future video coding technology the compression capability of which would exceed that of the WC. Some key definitions, bitstream and coding structures, and concepts of H.264 / AVC, HEVC, WC, and / or AV1 and some of their extensions are described in this section as an example of a video encoder, decoder, encoding method, decoding method, and a bitstream structure, wherein the embodiments may be implemented. The aspects of various embodiments are not limited to H.264 / AVC, HEVC, WC, and / or AV1 or their extensions, but rather the description is given for one possible basis on top of which the present embodiments may be partly or fully realized.
[0125] Hybrid video codecs, for example ITU-T H.263, H.264 / AVC, HEVC, and WC, may encode the video information in two phases. At first, pixel values in a certain picture area (or “block”) are predicted for example by motion compensation means or by spatial means.
[0126] In motion compensation based prediction (which may be referred to as inter prediction, temporal prediction or motion-compensated temporal prediction or motion-compensated prediction or MCP) an area in one of the previously coded frames that corresponds closely to the block being coded is found and used for prediction. Inter prediction may reduce temporal redundancy.
[0127] In spatial prediction pixel values around the block to be coded are used. In the first phase, predictive coding may be applied, for example, as so-called sample prediction and / or so-called syntax prediction. In the sample prediction, pixel or sample values in a certain picture area or "block" are predicted. These pixel or sample values can be predicted, for example, using one or more of motion compensation or intra prediction mechanisms.
[0128] Intra prediction, where pixel or sample values can be predicted by spatial mechanisms, involve finding and indicating a spatial region relationship. Intra prediction utilizes the fact that adjacent pixels within the same picture are likely to be correlated. Intra prediction can be performed in spatial or transform domain, i.e., either sample values or transform coefficients can be predicted. Intra prediction is typically exploited in intra coding, where no inter prediction is applied.
[0129] In motion vector prediction, motion vectors e.g., for inter and / or inter-view prediction may be coded differentially with respect to a block-specific predicted motion vector. In many video codecs, the predicted motion vectors are created in a predefined way, for example by calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions, sometimes referred to as advanced motion vector prediction (AMVP), is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signalling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index is typically predicted from adjacent blocks and / or co-located blocks in temporal reference picture. Differential coding of motion vectors is typically disabled across slice boundaries.
[0130] The block partitioning, e.g., from a coding tree unit (CTU) to coding units (CUs) and down to prediction units (PUs), may be predicted.
[0131] In filter parameter prediction, the filtering parameters e.g., for sample adaptive offset may be predicted. Prediction approaches using image information from a previously coded image can also be called as inter prediction methods which may also be referred to as temporal prediction and motion compensation. Prediction approaches using image information within the same image can also be called as intra prediction methods.
[0132] In the second phase of encoding, the prediction error, i.e., the difference between the predicted block of pixels and the original block of pixels, is coded. This may be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of it), quantizing the coefficients and entropy coding the quantized coefficients. By varying the fidelity of the quantization process, encoder can control the balance between the accuracy of the pixel representation (picture quality) and size of the resulting coded video representation (file size of transmission bitrate).
[0133] In some video codecs, such as H.265 / HEVC, video pictures are divided into coding units (CU) covering the area of the picture. A CU consists of one or more prediction units (Pll) defining the prediction process for the samples within the CU and one or more transform units (TU) defining the prediction error coding process for the samples in the said CU. A CU may consist of a square block of samples with a size selectable from a predefined set of possible CU sizes. A CU with the maximum allowed size is typically named as LCU (largest coding unit) or CTU (coding tree unit) and the video picture is divided into non-overlapping CTUs. A CTU can be further split into a combination of smaller CUs, e.g., by recursively splitting the CTU and resultant CUs. Each resulting CU typically has at least one PU and at least one TU associated with it. Each PU and TU can be further split into smaller PUs and TUs in order to increase granularity of the prediction and prediction error coding processes, respectively. Each PU has prediction information associated with it defining what kind of a prediction is to be applied for the pixels within that PU (e.g., motion vector information for inter predicted PUs and intra prediction directionality information for intra predicted PUs). Similarly, each TU is associated with information describing the prediction error decoding process for the samples within the said TU (including e.g., DCT coefficient information). It is typically signaled at CU level whether prediction error coding is applied or not for each CU. In the case there is no prediction error residual associated with the CU, it can be considered there are no TUs for the said CU. The division of the image into CUs, and division of CUs into PUs and TUs is typically signaled in the bitstream allowing the decoder to reproduce the intended structure of these units.
[0134] The decoder reconstructs the output video by applying prediction means similar to the encoder to form a predicted representation of the pixel blocks (using the motion or spatial information created by the encoder and stored in the compressed representation) and prediction error decoding (inverse operation of the prediction error coding recovering the quantized prediction error signal in spatial pixel domain). After applying prediction and prediction error decoding means the decoder sums up the prediction and prediction error signals (pixel values) to form the output video frame. The decoder (and encoder) can also apply additional filtering means to improve the quality of the output video before passing it for display and / or storing it as prediction reference for the forthcoming frames in the video sequence.
[0135] In many video codecs, including H.264 / AVC, HEVC, and WC, motion information is indicated by motion vectors associated with each motion compensated image block. Each of these motion vectors represents the displacement of the image block in the picture to be coded (in the encoder) or decoded (at the decoder) and the prediction source block in one of the previously coded or decoded images (or pictures). In order to represent motion vectors efficiently those are typically coded differentially with respect to block specific predicted motion vectors. In typical video codecs the predicted motion vectors are created in a predefined way, for example calculating the median of the encoded or decoded motion vectors of the adjacent blocks. Another way to create motion vector predictions is to generate a list of candidate predictions from adjacent blocks and / or co-located blocks in temporal reference pictures and signaling the chosen candidate as the motion vector predictor. In addition to predicting the motion vector values, the reference index of previously coded / decoded picture can be predicted. The reference index may be predicted from adjacent blocks and / or or co-located blocks in temporal reference picture. Moreover, high efficiency video codecs may employ an additional motion information coding / decoding mechanism, often called merging / merge mode, where all the motion field information, which includes motion vector and corresponding reference picture index for each available reference picture list, is predicted and used without any modification / correction. Similarly, predicting the motion field information is carried out using the motion field information of adjacent blocks and / or colocated blocks in temporal reference pictures and the used motion field information is signaled among a list of motion field candidate list filled with motion field information of available adjacent / co-located blocks.
[0136] Video codecs may support motion compensated prediction from one source image (uni-prediction) and two sources (bi-prediction). In the case of uniprediction a single motion vector is applied whereas in the case of bi-prediction two motion vectors are signaled and the motion compensated predictions from two sources are averaged to create the final sample prediction. In the case of weighted prediction, the relative weights of the two predictions can be adjusted, or a signaled offset can be added to the prediction signal.
[0137] In addition to applying motion compensation for inter picture prediction, similar approach can be applied to intra picture prediction. In this case the displacement vector indicates where from the same picture a block of samples can be copied to form a prediction of the block to be coded or decoded. This kind of intra block copying methods can improve the coding efficiency substantially in presence of repeating structures within the frame - such as text or other graphics.
[0138] In video codecs the prediction residual after motion compensation or intra prediction may be first transformed with a transform kernel (like DCT) and then coded. The reason for this is that often there still exists some correlation among the residual and transform can in many cases help reduce this correlation and provide more efficient coding.
[0139] Many video encoders utilize Lagrangian cost functions to find optimal coding modes, e.g., the desired Macroblock mode and associated motion vectors. This kind of cost function uses a weighting factor A to tie together the (exact or estimated) image distortion due to lossy coding methods and the (exact or estimated) amount of information that is required to represent the pixel values in an image area:
[0140] C = D + AR (Eq. 1 )
[0141] Where C is the Lagrangian cost to be minimized, D is the image distortion (e.g., Mean Squared Error) with the mode and motion vectors considered, and R the number of bits needed to represent the required data to reconstruct the image block in the decoder (including the amount of data to represent the candidate motion vectors).
[0142] The phrase along the bitstream (e.g., indicating along the bitstream) may be defined to refer to out-of-band transmission, signaling, or storage in a manner that the out-of-band data is associated with the bitstream. The phrase decoding along the bitstream or alike may refer to decoding the referred out- of-band data (which may be obtained from out-of-band transmission, signaling, or storage) that is associated with the bitstream. For example, an indication along the bitstream may refer to metadata in a container file that encapsulates the bitstream.
[0143] There are number of intra-prediction modes available for existing video codecs. These intra-prediction modes comprise different directional intra- prediction modes, as well as prediction modes such as DC or Planar intra- prediction. In Planar intra-prediction, interpolation processes in both the vertical and horizontal directions are performed to obtain vertical and horizontal predictors, respectively. Then, sample-by-sample mean of these two predictors can be computed. Each interpolation process may be carried out by computing weighted averages of the reference samples. As an example, when computing the vertical predictor, each predicted sample at a given horizontal coordinate is obtained as the weighted average between the two reference samples at the same horizontal coordinate extracted from the top row and bottom row of the current block, respectively. Due to the fact the samples located on the bottom row of the current block are not reconstructed when predicting the current block, these are padded using the sample on the bottom-left of the current block. Similarly, the horizontal predictor is computed as the weighted average of the reference samples extracted from the columns on the left and right of the current block.
[0144] ECM provides improved coding tools, such as Template-based intra mode derivation (TIMD) . TIMD relies on an inference process operating on a region, i.e. , template, formed of already reconstructed samples, for example samples in the surrounding of the current block. The intra prediction mode for a block is derived with a template based method both at the encoder and the decoder, instead of being signaled to the decoder. Figure 3a illustrates an example of a template 310 for a current block 300, and the reference samples 320 for the template 310. In TIMD, a distortion is calculated between the prediction and the reconstruction samples of the template. The intra prediction mode or modes with the minimum distortion are selected from an intra-prediction mode candidate list as the TIMD modes, and are used to compute the final intra prediction for the current block.
[0145] The existing version of TIMD technique relies on the MPM intra mode list construction to form the intra-prediction mode candidate list to perform the intra mode template search. In particular, for a given block, a template is considered, formed of reconstructed samples. Depending on the availability of the reference samples, an L-shaped template (as in Figure 3b template 350) is used in the surrounding of the current block (in case of complete availability of neighboring reconstructed samples) or a rectangular template either on the left and / or on top of the current block (as in Figure 3a with templates 310). Also depending on the block size, different template sizes may be used. Then, intra modes present in the MPM list are searched one-by-one. Available already reconstructed samples in the surrounding of the template are used as reference samples, where the samples in the template are used as target. A distortion is then computed between that prediction and the reconstruction template samples using SATD distortion metric. The intra prediction modes with the minimum SATD are selected as TIMD modes and are used for intra prediction of current block.
[0146] TIMD techniques, in its current form, rely on the template formed of already reconstructed samples in the surrounding of the current block. Rather than signalling the correct intra mode to use for the current block, the decoder searches among the different intra modes using the samples within the template as target and predicts the inter-prediction direction. Reconstructed samples in the surrounding of the template are used as reference samples. The decoder performs intra prediction with different modes using the reference samples, and then compares the resulting prediction with the target template, and finally selects the intra mode that minimizes a pre-defined distortion metric.
[0147] TIMD method can successfully reduce the overhead needed to signal a given intra prediction mode, but the resulting prediction is often sub-optimal, due to the fact the template region may not be strongly correlated with the content in the current block. TIMD in its current form is also restricted to use one distortion metric, one intra prediction mode candidate list, and a template size that only depends on the size of the current block.
[0148] Thus, it is an aim of present embodiments to provide an improved version of TIMD.
[0149] A method operating according to the invention applies an intra-prediction for a block, where the intra-prediction process comprises determining one or more intra-prediction modes to compute predictors, o wherein the intra-prediction modes are determined based on computing distortion in a template region using a distortion metric; and wherein at least one of the following applies: • the distortion metric is determined based on information extracted from a bitstream and / or based on information inferred by the decoder; or
[0150] • the template region is formed of a determined set of already reconstructed samples, where the set of already reconstructed samples is determined based on information extracted from the bitstream and / or based on information inferred by the decoder; or
[0151] • the intra-prediction modes are extracted from an intra-prediction mode candidate list based on computing distortions in a template region, where the candidate list is determined based on information extracted from the bitstream and / or based on information inferred by the decoder.
[0152] Use of Distortion metric
[0153] As discussed above, the method according to present embodiments may determine the intra-prediction modes based on computing the distortion in a template region (such as template 350 of Figure 3b) using a distortion metric, where the distortion metric is determined based on information extracted from a bitstream and / or inferred by the decoder.
[0154] As an example, distortions may be computed between already reconstructed samples in the template region 350 and intra-predicted samples. As an example, a number of possible distortion metrics may be considered to compute distortions. As an example, the distortion metric may be determined as the SAD or the SATD or the SSE.
[0155] As an example, when computing a distortion using a given distortion metric, the mean difference between the already reconstructed samples in the template region and the intra-predicted samples may be used. As an example, the mean difference may be subtracted when computing the distortion for a given sample. As an example, the mean difference may be determined as the MR-SAD or the MR-SATD or the MR-SSE. As an example, the usage of removing the mean difference may be determined by reading a flag from the encoded bitstream. According to an embodiment, a number of different distortion metrics can be combined to obtain a distortion. As an example, the sum of two distortions obtained by using two different distortion metrics can be computed. As an example, the sum of two or more distortions determined using different distortion metrics can be used to derive the intra-prediction modes over a template region.
[0156] For example, information indicating which distortion metric to use can be encoded to the bitstream. As another example a flag indicating whether distortions are computed using SAD, SATD, SSE, MR-SAD, MR-SATD, MR- SSE or other metric can be encoded to the bitstream. As another example, an index to determine a given distortion metric from a pre-determined list of available distortion metrics can be encoded to the bitstream. The list of available distortion metrics may be computed based on characteristics, for example a size, of the current block. The decoder is configured to extract the information or the flag or the index from the encoded bitstream.
[0157] As an example, the distortion metric to be used by the decoder may be determined based on information inferred by the decoder. As an example, the decoder may determine a distortion in a template region to determine the distortion metric. As an example, the distortion metric may be determined by measuring a distortion against a given pre-determined threshold. As an example, distortions can be determined for two or more distortion metrics, and a final distortion metric is determined based on these distortions.
[0158] A set of intra-prediction modes with their corresponding distortions may be determined for each distortion metric, and then the final set of intra-prediction modes to be used to predict the current block is determined between these different sets of modes based on the distortions.
[0159] Use of a determined set of already reconstructed samples
[0160] As discussed above, the method according to present embodiments may determine intra-prediction modes based on computing distortions in a template region using a distortion metric. The template region is formed of a determined set of already reconstructed samples, where the decoder is configured to determine the set of already reconstructed samples based on information from the bitstream. This means that the decoder is able to determine how the template should be created.
[0161] For example, for indicating the decoder how to create the template, a flag can have been encoded to the bitstream which indicates which template region to use between two pre-determined template regions. The flag may be extracted by the decoder. For example, the flag may indicate that the decoder should either use a smaller template or a larger template. For example, the flag may indicate that the decoder should either use a smaller template formed of 2 or 4 lines and / or columns of already reconstructed samples depending on the block size, or a larger template formed of 4 or 8 lines and / or columns of already reconstructed samples depending on the block size.
[0162] Alternatively, the flag may indicate whether the number of samples in the template region should be equal to a pre-determined number N, or whether the number of samples should be equal to a different number of samples M = 2N. The flag may be extracted by the decoder.
[0163] According to an embodiment, the set of already reconstructed samples in the template region may be determined based on information inferred by the decoder. The set of already reconstructed samples in the template region may be determined based on the distortion metric used to compute distortions in the template region. As an example, a larger number of samples in the template region may be used if a distortion metric different than the SATD is used.
[0164] Use of Candidate list
[0165] As discussed above, the method according to present embodiments may determine intra-prediction modes from an intra-prediction mode candidate list based on computing distortions in a template region.
[0166] The information on the candidate list may be encoded in the bitstream to be extracted by the decoder.
[0167] For example, a flag can have been encoded to the bitstream which indicates whether a first candidate list should be used, or whether a second candidate list should be used. The second candidate list may be formed based on information extracted from the first candidate list. For example, the second candidate list may be formed based on distortions computed in a template region on modes extracted from the first candidate list. As another example, the second candidate list may be formed in such a way that the modes at minimum distortion among the first candidate list are excluded from being inserted in the second candidate list. As another example, the second candidate list may be determined based on the modes at minimum distortion among the first candidate list.
[0168] According to an embodiment, an index may have been encoded to the bitstream to indicate the method to form the candidate list from number of available methods. Such index may be extracted from the bitstream by the decoder.
[0169] The method according to an embodiment is shown in Figure 4. The method is aimed for decoding encoded samples of blocks of video sample data, wherein for a block of video sample data the method generally comprises deriving 410 information to configure a template-based intra-prediction mode derivation process, wherein the configured template-based intra-prediction mode derivation process comprises computing 420 distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra-prediction mode candidate list; determining 430 one or more intra mode predictors by the determined one or more intra-prediction modes; and computing 440, based on the intra mode predictors, a prediction for the samples of the video sample data. Each of the steps can be implemented by a respective module of a computer system.
[0170] An apparatus for decoding encoded samples of blocks of video sample data is provided, wherein for a block of video sample data the apparatus comprises means for deriving information to configure a template-based intra-prediction mode derivation process, wherein for implementing the configured templatebased intra-prediction mode derivation process the apparatus further comprises means for computing distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra- prediction mode candidate list; means for determining one or more intra mode predictors by the determined one or more intra-prediction modes; and means for computing, based on the intra mode predictors, a prediction for the samples of the video sample data. The means comprises at least one processor, and a memory including a computer program code, wherein the processor may further comprise processor circuitry. The memory and the computer program code are configured to, with the at least one processor, cause the apparatus to perform the method of Figure 4 according to various embodiments.
[0171] Figure 5 illustrates an example of an electronic apparatus 500, being an example of video coding system where the present embodiments can be implemented. In some embodiments, the apparatus may be a mobile terminal or a user equipment of a wireless communication system or a camera device. The apparatus 500 may also be comprised at a local or a remote server or a graphic processing unit of a computer. The apparatus may also be comprised as part of a head-mounted display device.
[0172] The apparatus may be configured to perform various functions, such as for example, gathering information by one or more sensors, encoding and / or decoding information, receiving and / or transmitting information, analyzing information gathered or received by the apparatus. An apparatus configured to encode a video scene may optionally comprise one or more microphones for capturing the scene and / or one or more cameras for capturing information about the physical environment in which the scene is captured. Alternatively, the apparatus configured for encoding may be configured to receive information about an environment in which a scene is captured and / or a simulated environment. An apparatus configured to decode and / or render the video scene may be configured to receive a bitstream comprising encoded video. An apparatus configured to decode and / or render the video scene may comprise one or more speakers / audio transducers and / or displays, and / or may be configured to transmit a decoded scene or signals to a device comprising one or more speakers / audio transducers and / or displays. An apparatus configured to decode and / or render the video scene may comprise a user equipment, a head-mounted display, or another device capable of rendering to a user an AR; VR and / or MR experience.
[0173] The apparatus 500 comprises one or more processors 510 and one or more memories 520 and one or more transceivers interconnected through one or more buses. The one or more memories 520 store computer instructions, for example in respective modules (Modulel , Module2, ModuleN). The one or more memories may store data in the form of image, video and / or audio data, and / or may also store instructions to be executed by the processors or the processor circuitry. The one or more processors may comprise a central processing unit (CPU) and / or a graphical processing unit (GPU). The one or more buses may be address, data or control buses, and may include interconnection mechanism, such as a series of lines on a motherboard or integrated circuit, fiber optics or other optical communication equipment. The apparatus also comprises a codec 530 that is configured to implement various embodiments relating to present solution. According to some embodiments, the apparatus may comprise an encoder or a decoder. The apparatus 500 also comprises a communication interface 540 which is suitable for generating wireless communication signals for example for communication with a cellular communications network, a wireless communications system, or a wireless local area network, and thus enabling data transfer over data transfer network 550.
[0174] The apparatus 500 may comprise a display in the form of a liquid crystal display. In other embodiments of the invention the display may be any suitable display technology suitable to display an image or video. The apparatus 500 may further comprise a keypad. In other embodiments of the invention any suitable data or user interface mechanism may be employed. For example, the user interface may be implemented as a virtual keyboard or data entry system as part of a touch-sensitive display. The apparatus 500 may comprise a microphone or any suitable audio input which may be a digital or analogue signal input. The apparatus 500 may further comprise an audio output device which in embodiments of the invention may be any one of: an earpiece, speaker, or an analogue audio or digital audio output connection. The apparatus 500 may also comprise a battery (or in other embodiments of the invention the device may be powered by any suitable mobile energy device such as solar cell, fuel cell or clockwork generator). The apparatus may further comprise a camera capable of recording or capturing images and / or video. The camera may be a multi-lens camera system having at least two camera sensors. The camera is capable of recording or detecting individual frames which are then passed to the codec 530 or to processor 510. The apparatus may receive the video and / or image data for processing from another device prior to transmission and / or storage. The apparatus 500 may further comprise e.g., the other functional units disclosed in any of the Figures 1 - 2 for implementing any of the present embodiments.
[0175] The apparatus may operate in a system, comprising multiple communication devices, which can communication through one or more networks. The system may comprise any combination of wired or wireless networks including, but not limited to a wireless cellular telephone network (such as a GSM, UMTS, CDMA network etc.), a wireless local area network (WLAN) such as defined by any of the IEEE 802.x standards, a Bluetooth personal area network, an Ethernet local area network, a token ring local area network, a wide area network, and the Internet.
[0176] For example, the system can be a mobile telephone network enabling a connection to the internet. The connection can be, but is not limited to, long range wireless connections, short range wireless connections, and various wired connections including, but not limited to, telephone lines, cable lines, power lines, and similar communication pathways.
[0177] The example communication devices operating in the system may include, but are not limited to, an electronic device or apparatus, a combination of a personal digital assistant (PDA) and a mobile telephone, a PDA, an integrated messaging device (IMD), a desktop computer, a notebook computer, each of which can be a representative of the apparatus according to present embodiments. The apparatus according to present embodiments may be stationary or mobile when carried by an individual who is moving. The apparatus may also be located in a mode of transport including, but not limited to, a car, a truck, a taxi, a bus, a train, a boat, an airplane, a bicycle, a motorcycle, or any similar suitable mode of transport.
[0178] The apparatus may also be a set-top box; i.e. a digital TV receiver, which may / may not have a display or wireless capabilities, a tablet or (laptop) a personal computer (PC), which have hardware or software or combination of the encoder / decoder implementations, in various operating systems, or a chipset, processor, DSP and / or embedded system offering hardware / software based coding. The apparatus according to present embodiments may send and receive calls and messages and communicate with service providers through a wireless connection to a base station. The base station may be connected to a network server that allows communication between the mobile telephone network and the internet.
[0179] The apparatus may communicate using various transmission technologies including, but not limited to, code division multiple access (CDMA), global systems for mobile communications (GSM), universal mobile telecommunications system (UMTS), time divisional multiple access (TDMA), frequency division multiple access (FDMA), transmission control protocolinternet protocol (TCP-IP), short messaging service (SMS), multimedia messaging service (MMS), email, instant messaging service (IMS), Bluetooth, IEEE 802.11 and any similar wireless communication technology. A communications device involved in implementing various embodiments of the present invention may communicate using various media including, but not limited to, radio, infrared, laser, cable connections, and any suitable connection.
[0180] In telecommunications and data networks, a channel may refer either to a physical channel or to a logical channel. A physical channel may refer to a physical transmission medium such as a wire, whereas a logical channel may refer to a logical connection over a multiplexed medium, capable of conveying several logical channels. A channel may be used for conveying an information signal, for example a bitstream, from one or several senders (or transmitters) to one or several receivers.
[0181] An MPEG-2 transport stream (TS), specified in ISO / IEC 13818-1 or equivalently in ITU-T Recommendation H.222.0, is a format for carrying audio, video, and other media as well as program metadata or other metadata, in a multiplexed stream. A packet identifier (PID) is used to identify an elementary stream (a.k.a. packetized elementary stream) within the TS. Hence, a logical channel within an MPEG-2 TS may be considered to correspond to a specific PID value. Available media file format standards include ISO base media file format (ISO / IEC 14496-12, which may be abbreviated ISOBMFF) and file format for NAL unit structured video (ISO / IEC 14496-15), which derives from the ISOBMFF.
[0182] The various embodiments can be implemented with the help of computer program code that resides in a memory and causes the relevant apparatuses to carry out the method. For example, a device may comprise circuitry and electronics for handling, receiving, and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the device to carry out the features of an embodiment. Yet further, a network device like a server may comprise circuitry and electronics for handling, receiving and transmitting data, computer program code in a memory, and a processor that, when running the computer program code, causes the network device to carry out the features of various embodiments.
[0183] If desired, the different functions discussed herein may be performed in a different order and / or concurrently with other. Furthermore, if desired, one or more of the above-described functions and embodiments may be optional or may be combined.
[0184] Although various aspects of the embodiments are set out in the independent claims, other aspects comprise other combinations of features from the described embodiments and / or the dependent claims with the features of the independent claims, and not solely the combinations explicitly set out in the claims.
[0185] It is also noted herein that while the above describes example embodiments, these descriptions should not be viewed in a limiting sense. Rather, there are several variations and modifications, which may be made without departing from the scope of the present disclosure as, defined in the appended claims.
Claims
Claims:1 . An apparatus for decoding encoded samples of blocks of video sample data is provided, wherein for a block of video sample data the apparatus comprises- means for deriving information to configure a template-based intra-prediction mode derivation process, wherein for implementing the configured template-based intra-prediction mode derivation process the apparatus further comprises o means for computing distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra-prediction mode candidate list; o means for determining one or more intra mode predictors by the determined one or more intra-prediction modes; and o means for computing, based on the intra mode predictors, a prediction for the samples of the video sample data.
2. The apparatus according to claim 1 , wherein the derived information is derived from an encoded bitstream.
3. The apparatus according to claim 1 or 2, further comprising means for using the derived information to determine a distortion metric among a plurality of distortion metrics.
4. The apparatus according to claim 3, wherein the distortion metric is one of the following: Sum of Absolute Differences, Sum of Absolute Transformed Differences, Sum of Squared Errors, Mean-removed Sum of Absolute Differences, Mean-removed Sum of Absolute Transformed Differences, Mean-removed Sum of Squared Errors.
5. The apparatus according to claim 1 or 2, further comprising means for using the derived information to determine a set of reconstructed samples to form the template region.
6. The apparatus according to claim 1 or 2, further comprising means for using the derived information to determine which intra-prediction mode candidate list of at least two intra-prediction mode candidate lists is used.
7. A method for decoding encoded samples of blocks of video sample data is provided, wherein for a block of video sample data the method comprises- deriving information to configure a template-based intra- prediction mode derivation process, wherein the configured template-based intra-prediction mode derivation process comprises o computing distortions on a template region based on a distortion metric to determine one or more intra-prediction modes from an intra-prediction mode candidate list; o determining one or more intra mode predictors by the determined one or more intra-prediction modes; and o computing, based on the intra mode predictors, a prediction for the samples of the video sample data.
8. The method according to claim 7, wherein the derived information is derived from an encoded bitstream.
9. The method according to claim 7 or 8, further comprising using the derived information to determine a distortion metric among a plurality of distortion metrics.
10. The method according to claim 9, wherein the distortion metric is one of the following: Sum of Absolute Differences, Sum of Absolute Transformed Differences, Sum of Squared Errors, Mean-removed Sum of Absolute Differences, Mean-removed Sum of Absolute Transformed Differences, Mean-removed Sum of Squared Errors.
11. The method according to claim 7 or 8, comprising using the derived information to determine a set of reconstructed samples to form the template region.
12. The method according to claim 7 or 8, further comprising using the derived information to determine which intra-prediction mode candidate list of at least two intra-prediction mode candidate lists is used.
Citation Information
Patent Citations
Method and apparatus for encoding / decoding an image
US20190379891A1
Method, device, and medium for video processing
WO2022253318A1
Methods and devices for decoder-side intra mode derivation
WO2023141238A1