Intra merge mode
By using decoder-side intra-mode derivation (DIMD) technology, the intra-prediction mode is implicitly derived using the texture analysis information of neighboring blocks, which solves the problem of low efficiency of intra-prediction mode in the existing technology and achieves more efficient film encoding and decoding performance and texture analysis information utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MEDIATEK INC
- Filing Date
- 2024-07-18
- Publication Date
- 2026-05-01
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in intra-frame prediction mode derivation and struggle to effectively utilize texture analysis information to improve encoding and decoding performance and reduce complexity.
The decoder-side intra-frame prediction mode (DIMD) technique is adopted. By analyzing the texture analysis information of neighboring blocks, the intra-frame prediction mode is implicitly derived. Combined with gradient histogram and template matching methods, the final intra-frame prediction mode is generated. Prediction fusion is performed during the encoding process to reduce the bit rate.
It improves the encoding and decoding efficiency of intra-frame prediction mode, reduces the bit rate, and enhances the performance of video encoding and decoding as well as the utilization efficiency of texture analysis information.
Smart Images

Figure CN121970325A_ABST
Abstract
Description
Intra-frame merging mode
[0001] [Cross-Reference] This disclosure is part of a non-provisional application that claims priority to U.S. Provisional Patent Application No. 63 / 514,155, filed July 18, 2023. The contents of that application are incorporated herein by reference. Technical Field
[0002] [Technical Field] This disclosure is generally related to video encoding and decoding. In particular, this disclosure relates to methods for encoding and decoding pixel blocks using intra-frame prediction and merging modes. Background Technology
[0003] [Background Art] Unless otherwise stated herein, the methods described in this section are not prior art to the following claims, and are not considered prior art by virtue of being included in this section.
[0004] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC). HEVC uses a hybrid block-based motion compensation coding architecture similar to Discrete Cosine Transform (DCT). The basic unit of compression, called a Coding Unit (CU), is a 2Nx2N square pixel block. Each CU can be recursively divided into four smaller CUs until a predefined minimum size is reached. Each CU contains one or more Prediction Units (PUs).
[0005] Versatile Video Coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of the International Telecommunication Union (ITU) T Sector SG16 WP3 and the International Organization for Standardization / International Electrotechnical Commission (ISO) JTC1 / SC29 / WG11. The input video signal is predicted from a reconstructed signal derived from encoded (i.e., encoded or decoded) picture regions. The prediction residual signal is processed through block transform. The transform coefficients are quantized and entropy-coded in the bitstream along with other side information. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform and dequantization of the transform coefficients. The reconstructed signal is further processed by loop filtering to remove coding artifacts. The decoded picture is stored in a frame buffer for predicting future pictures in the input video signal.
[0006] In VVC, the encoded image is divided into non-overlapping square block regions, represented by associated Coding Tree Units (CTUs). The leaf nodes of the code tree correspond to Coding Tree Units (CUs). The encoded image can be represented by a series of slices, each slice containing an integer number of CTUs. Individual CTUs within a slice are processed in raster scan order. Bidirectional prediction (B) slices may use intra-frame prediction or inter-frame prediction, with up to two motion vectors (MVs) and a reference index to predict the sample value for each block. Predictive (P) slices use intra-frame prediction or inter-frame prediction, with up to one motion vector and a reference index to predict the sample value for each block. Intra-frame prediction (I) slices use only intra-frame prediction decoding.
[0007] CTU can be partitioned into one or more non-overlapping codec units (CUs) using a quadtree (QT) with a nested multi-type-tree (MTT) structure to accommodate various local motion and texture characteristics. CUs can be further partitioned into smaller CUs using one of five partitioning types: quadtree partitioning, vertical binary tree partitioning, horizontal binary tree partitioning, vertical ternary tree partitioning, and horizontal ternary tree partitioning.
[0008] Each CU contains one or more prediction units (PUs). A prediction unit, along with its associated CU syntax, serves as the basic unit of signal predictor information. The specified prediction process is used to predict the values of associated pixel samples within a PU. Each CU may contain one or more transform units (TUs) to represent prediction residual blocks. A transform unit (TU) consists of a transform block (TB) of one luma sample and / or two corresponding chroma samples, each TB corresponding to a sample from a residual block of a color component. Integer transforms are applied to the transform blocks. The level values of the quantization coefficients, along with other side information, are entropy-encoded in the bitstream. The terms Coding Tree Block (CTB), Coding Block (CB), Prediction Block (PB), and Transform Block (TB) are defined as 2-D sample arrays used to specify the monochrome components associated with the CTU, CU, PU, and TU, respectively. Therefore, a CTU consists of one luma CTB, two chroma CTBs, and / or associated syntax elements. A similar relationship applies to CU, PU, and TU.
[0009] For each inter-frame predicted CU, motion parameters include motion vectors, reference picture indices, and reference picture list usage indices, as well as additional information used for generating inter-frame predicted samples. Motion parameters can be signaled explicitly or implicitly. When a CU is coded in skip mode, the CU is associated with a PU, with no significant residual coefficients, no encoded motion vector difference, and no reference picture index. A merging mode is specified where the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as additional plans introduced in the VVC. The merging mode can be applied to any inter-frame predicted CU. An alternative to the merging mode is explicit transmission of motion parameters, where the motion vector of each CU, the corresponding reference picture index for each reference picture list, the reference picture list usage flag, and other necessary information are explicitly signaled.
[0010] To improve encoding / decoding performance and / or reduce the complexity of systems using texture analysis information, methods and apparatus for encoding / decoding video blocks using texture analysis information are disclosed. Summary of the Invention
[0011] [Summary of the Invention] The following abstract is for illustrative purposes only and is not intended to limit in any way. That is, the following abstract is intended to introduce the concepts, highlights, benefits, and advantages of the novel and non-obvious techniques described herein. Selective and not exhaustive embodiments will be further described in the detailed description. Therefore, the following abstract is not intended to define the essential characteristics of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter.
[0012] Certain embodiments of this disclosure provide a method for modifying or adjusting information related to decoder-side intra-mode derivation (DIMD) to encode or decode pixel blocks. A video encoder / decoder (i.e., an encoder or decoder) receives data to encode or decode pixels of a current block of a current picture of a video. The video encoder / decoder analyzes samples neighboring the current block to generate a first set of texture analysis information (also referred to as, for example, DIMD information). The set of texture analysis information may include histogram values, intra-prediction modes, or both. The video encoder / decoder obtains or derives at least one second set of texture analysis information associated with the current block or at least one reference block. The video encoder / decoder generates a final set of texture analysis information based on the first set of texture analysis information, the second set of texture analysis information, or both. The video encoder / decoder encodes or decodes the current block according to the final set of texture analysis information, or stores the final set of texture analysis information.
[0013] At least one reference block may be determined based on at least one candidate selected from a candidate list. The video codec may examine one or more candidates for the current block. These candidates may include all or any subset of spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates.
[0014] At least one set of second texture analysis information may be inherited from at least one reference block, or from at least one candidate selected from a candidate list. In some embodiments, the inheritance of at least one set of second texture analysis information is permitted only if at least one reference block is intra-predictively coded or contains texture analysis information. In some embodiments, the inheritance of at least one set of second texture analysis information is permitted only if at least one reference block is located in a predefined region.
[0015] In some embodiments, at least one set of second texture analysis information is derived by analyzing samples associated with at least one reference block. In some embodiments, at least one set of second texture analysis information is derived by analyzing predicted samples of the current block or at least one reference block.
[0016] In some embodiments, multiple texture analysis information sets may be combined into a final set of texture analysis information. The video codec may store the final set of texture analysis information, for example, for inheritance or reference in subsequently encoded blocks. The stored final set of texture analysis information is a subset of the final set of texture analysis information. The final set of texture analysis information may include one or more intra-frame prediction modes determined to generate multiple prediction hypotheses to form the final prediction for the current block. The video codec may generate the final prediction based on the intra-frame prediction modes derived from the final set of texture analysis information. Attached Figure Description
[0017] [Figure Description] The accompanying drawings are included to provide a further understanding of this disclosure and are incorporated into and constitute a part of this disclosure. These drawings illustrate embodiments of this disclosure and, together with the description, serve to explain the principles of this disclosure. It is appreciated that the drawings are not necessarily to scale, as some components may be shown out of proportion to their actual size in order to clearly illustrate the concepts of this disclosure.
[0018] Figure 1 shows the intra-prediction modes in different directions.
[0019] Figures 2A-2B conceptually illustrate the extended lengths of the top and left reference samples for wide-angle orientation patterns with different aspect ratios to support non-square blocks.
[0020] Figure 3 illustrates the use of template-based intra modederivation (TIMD) to implicitly derive an intra prediction mode for the current block.
[0021] Figure 4 illustrates the use of decoder-side intra-mode derivation (DIMD) to implicitly derive an intra-prediction mode for the current block.
[0022] Figure 5 illustrates the neighbor reconstruction samples of the current block used for the DIMD chromaticity mode.
[0023] Figure 6 shows the mapping from intra-prediction mode to the transformed set.
[0024] Figure 7 illustrates the construction of gradient histograms using matrix-based intra-frame prediction samples.
[0025] Figure 8 illustrates the locations of spatial merging candidates.
[0026] Figure 9 shows the spatial neighbor blocks used to derive spatial merge candidates.
[0027] Figure 10 illustrates the scaling of motion vectors for time merging candidates.
[0028] Figure 11 shows the candidate positions of the time merging candidates.
[0029] Figure 12 illustrates an example video encoder that may implement DIMD.
[0030] Figure 13 illustrates a portion of the movie encoder that implements intra-frame prediction based on proposed DIMD information.
[0031] Figure 14 conceptually illustrates the process of encoding pixel blocks using the proposed DIMD information.
[0032] Figure 15 illustrates an example video decoder that may implement DIMD.
[0033] Figure 16 illustrates a portion of a video decoder that implements intra-frame prediction based on proposed DIMD information.
[0034] Figure 17 conceptually illustrates the process of decoding pixel blocks using the proposed DIMD information.
[0035] Figure 18 conceptually illustrates some electronic systems that implement the contents of this disclosure. Detailed Implementation
[0036] [Detailed Description] In the following detailed description, numerous specific details are set forth by way of example to provide a thorough understanding of the relevant teaching. Any variations, derivatives, and / or extensions of the teaching described herein are within the scope of this disclosure. In some cases, well-known methods, procedures, components, and / or circuits implemented with respect to one or more examples disclosed herein may be described at a relatively high level without detail to avoid unnecessarily obscuring aspects of the teaching of this disclosure.
[0037] I. Intra-prediction modes a. Directional intra-prediction modes The intra-prediction method uses a reference layer that may be adjacent to the current prediction unit (PU) and an intra-prediction mode to generate the predictor for the current PU. The intra-prediction direction can be selected from a mode set containing multiple prediction directions. For each PU coded via intra-prediction, an index is used and encoded to select one of the intra-prediction modes. The corresponding prediction is generated, and the residual can then be derived and transformed.
[0038] Figure 1 shows intra-prediction modes in different directions. These intra-prediction modes are called directional modes, excluding DC or planar modes. As shown, there are 33 directional modes (V: vertical; H: horizontal), hence H, H+1~H+8, H-1~H-7, V, V+1~V+8, V-1~V-8. A directional mode can typically be represented as H+k or V+k, where k = ±1, ±2..., ±8. Such intra-prediction modes can also be called intra-prediction angles. To capture arbitrary edge directions presented in natural video, the number of directional intra-prediction modes can be expanded from the 33 used in HEVC to 65, so that k ranges from ±1 to ±16. These denser directional intra-prediction modes are suitable for all block sizes and for both luma and chroma intra-prediction. By including DC and planar modes, the number of intra-prediction modes is 35 (or 67).
[0039] Intra-prediction modes used for the luma component can be directly used for the chroma component. This is called chroma DM (direct mode).
[0040] Of the 35 (or 67) intra-prediction modes, some modes (e.g., 3 or 5) are identified as the most probable mode (MPM) for intra-prediction of the current prediction block. The encoder can reduce the bit rate by sending an index to select one of the MPMs instead of an index to select one of the 35 (or 67) intra-prediction modes. For example, the intra-prediction modes used in the left prediction block and the intra-prediction modes used in the upper prediction block are used as MPMs.
[0041] Traditional angular intra-prediction directions are defined from 45 degrees to -135 degrees clockwise. In VVC, several traditional angular intra-prediction modes are adaptively replaced with non-blocky wide-angle intra-prediction modes. The replaced modes are signaled using the original mode index and then remapped to the wide-angle mode index after resolution.
[0042] For some embodiments, the total number of intra-prediction modes remains constant at 67, and the intra-prediction encoding / decoding method remains unchanged. To support these prediction directions, a top reference sample of length 2W+1 and a left reference sample of length 2H+1 are defined. Figures 2A-2B conceptually show the extended lengths of the top and left reference samples for wide-angle orientation modes to support non-blocky blocks with different aspect ratios.
[0043] The number of replacement patterns in the wide-angle orientation mode depends on the block's aspect ratio. The intra-frame prediction replacement patterns for blocks with different aspect ratios are shown in Table 1 below.
[0044] Table 1: Intra-prediction modes replaced by wide-angle mode b. Chroma Intra-Frame Mode Encoding and Decoding For chroma intra-frame mode encoding and decoding, a total of eight chroma intra-frame mode encoding and decoding methods are allowed. These modes include five traditional intra-frame modes and three cross-component linear model modes (CCLM, LM_A, and LM_L). Chroma mode encoding and decoding may directly depend on the intra-prediction mode of the corresponding luma block. Due to the independent block segmentation structure of luma and chroma components enabled in I-frames, one chroma block may correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra-prediction mode of the corresponding luma block directly inherits from or derives from the corresponding luma intra-prediction mode, and the luma block covers the center position of the current chroma block.
[0045] c. Template-based Intra Mode Derivation (TIMD) For mode selection, template matching methods can be applied by calculating the cost between reconstructed samples and predicted samples on a template. One example is Template-based Intra Mode Derivation (TIMD). TIMD is an encoding / decoding method in which the intra-predicted mode of the CU is implicitly derived in the encoder and decoder using neighboring templates, rather than the encoder explicitly predicting the intra-predicted mode of the signal to the decoder.
[0046] Figure 3 illustrates the implicit derivation of intra-prediction modes for the current block 300 using Template-Based Intra-Mode Derivation (TIMD). As shown, neighboring pixels of the current block (Current CU) 300 are used as template 310. For each candidate intra-prediction mode, a prediction sample for the template is generated using reference samples in the L-shaped reference region 320 above and to the left of template 310. The TM cost of the candidate intra-prediction mode is calculated based on the difference (e.g., SATD) between the reconstructed sample of the template and the prediction sample of the template generated by the candidate intra-prediction mode. The candidate intra-prediction mode with the lowest cost (as a TIMD mode, similar to the implicit mode derivation in DIMD mode) is selected and used for intra-prediction of the CU. The candidate intra-prediction modes may include 67 intra-prediction modes (as in VVC) or be extended to 131 intra-prediction modes. MPM can be used to indicate the orientation information of the CU. Therefore, to reduce the intra-modal search space and utilize the characteristics of the CU, the intra-prediction mode is implicitly derived from the MPM list.
[0047] In some embodiments, for each intra-prediction mode in the MPM list, the SATD between the predicted and reconstructed samples of the template is calculated as the TM cost of the intra-prediction mode. The top two intra-prediction modes with the minimum SATD are selected as TIMD modes. These two TIMD modes are weighted together after applying the PDPC procedure, and this weighted intra-prediction is used to encode the current CU. Position-dependent IntraPrediction Combination (PDPC) is included in the derivation of the TIMD modes.
[0048] Compare the cost of two selected intra-prediction modes (mode1 and mode2) with a threshold, for example, by applying a cost factor of 2 as follows: costMode2 < 2 * costMode1. If this condition is true, prediction fusion is applied; otherwise, only mode1 is used. The weights of the modes are calculated from their SATD costs as follows: weight1 = costMode2 / (costMode1 + costMode2) weight2 = 1 - weight1 II. Decoder Side Intra Mode Derivation (DIMD) a. Preliminary DIMD Decoder Side Intra Mode Derivation (DIMD) is a technique in which one or more (e.g., two) intra-prediction modes, such as angles or orientations, are derived from reconstructed neighboring samples (templates) of a block, and these two predictors are combined with a non-angle predictor (such as a planar mode predictor), with weights derived by gradients. DIMD modes are used as alternative prediction modes and / or are always checked in high-complexity RDO modes. To implicitly derive the intra-prediction modes of a block, texture gradient analysis is performed on both the encoder and decoder sides. The process begins with an empty Histogram of Gradients (HoG) with 65 entries, corresponding to 65 intra-frame prediction modes of angle / direction. The amplitudes of these entries are determined during texture gradient analysis.
[0049] The video codec performing DIMD executes the following steps: In the first step, the video codec selects a template of T=3 columns and rows from the left and top of the current block, respectively. This region serves as a reference for gradient-based intra-prediction mode derivation. In the second step, horizontal and vertical Sobel filters are applied to all 3×3 window locations, centered on the pixels of the middle row of the template. At each window location, the Sobel filter computes the intensity in the purely horizontal and vertical directions, respectively. and Then, the texture angle of the window is calculated as follows: This can be converted to one of 65 intra-prediction modes with different angles. Once the intra-prediction mode index for the current window is derived as idx, it is then added... Update the amplitude of the entries in HoG[idx].
[0050] Figure 4 illustrates the use of decoder-side intra-mode derivation (DIMD) to implicitly derive an intra-predictive mode for a current block. The figure shows an example of a Histogram of Gradients (HoG) 410, calculated after applying the above operation to all pixel locations in a template 415, which includes sample lines of neighboring pixels surrounding a current block 400. Once the HoG is calculated, the indices (M1 and M2) of the two highest histogram bars are selected as the two implicitly derived intra-predictive modes (IPMs) for that block. The predictions of these two IPMs are further combined with the planar mode prediction as the prediction of the DIMD mode. The prediction fusion is applied as a weighted average of the three predictors (M1 prediction, M2 prediction, and planar mode prediction). For this purpose, the weight of the planar mode might be set to 21 / 64 (approximately 1 / 3). The remaining weights, 43 / 64 (approximately 2 / 3), are proportionally allocated between the two HoG IPMs based on the magnitude of their HoG bars. The DIMD prediction fusion or combined prediction can be: Pred DIMD = (43*(w1* pred M1 + w2*pred M2 ) + 21*pred planar )>>6w1 = amp M1 / (amp M1 +amp M2 w2 = amp M2 / (amp M1 +amp M2 Furthermore, derived intra-prediction modes, such as two implicitly derived intra-prediction modes, are added to the most probable modes (MPM) list, so the DIMD process is performed before the MPM list is built. The main derived intra-prediction mode of the DIMD block is stored in a block and used for the MPM list construction of neighboring blocks.
[0051] As mentioned earlier, when applying DIMD, two intra-predictive modes are derived from the reconstructed neighboring samples. These two predictors are combined with a planar mode predictor, and the weights are derived from the gradient. The division operation in weight derivation is performed using the same lookup table (LUT) as CCLM, based on an integerization scheme. For example, the division operation in direction calculation... It is calculated using the following LUT-based scheme: x = Floor( Log2( Gx ) )normDiff = ( ( Gx<<4 )>>x )&15x +=( 3 + ( normDiff != 0 ) ? 1 : 0 )Orient = (Gy* ( DivSigTable[ normDiff ] | 8 ) + ( 1<<( x-1 ) ))>>x where DivSigTable
[16] = { 0, 7, 6, 5 ,5, 4, 4, 3, 3, 2, 2, 1, 1, 1, 1, 0}.
[0052] b. DIMD Chroma Mode The DIMD chroma mode uses the DIMD derivation method to derive the chroma intra-prediction mode for the current (chroma) block, based on the reconstructed Y, Cb, and Cr samples from the second nearest neighbor row and column (from the second collocated Y block co-located with the current block), as shown in Figure 5, which illustrates the neighbor reconstructed samples for the current block used in the DIMD chroma mode. Specifically, horizontal and vertical gradients are computed for each collocated collocated Y sample (e.g., neighboring collocated reconstructed Y samples) and neighboring reconstructed Cb and Cr samples of the current chroma block to establish a HoG. The intra-prediction mode with the largest histogram magnitude value is then used to perform chroma intra-prediction for the current chroma block (e.g., the current Cb block and the current Cr block).
[0053] When the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU-level flag is signaled to indicate whether the DIMD chroma mode is applied.
[0054] c. Transform Selection and the DIMD encoder have the ability to apply non-separable or separable mode-dependent transforms to the input, such as predicting residuals or low-frequency coefficients after the master transform, as in DCT Type-II. A non-separable mode-dependent transform can be a Low-Frequency Non-Separable Transform (LFNST), which is performed on the low-frequency coefficients after the master transform. The decoder performs the inverse transform process.
[0055] An example of LFNST is shown below. Three different cores, LFNST4, LFNST8, and LFNST16, are defined to indicate the application to 4xN / Nx4 (N 4) 8xN / Nx8 (N 8) and MxN (M, N) 16) is the LFNST core set. The core size is specified by: (LFSNT4, LFNST8, LFNST16) = (16x16, 32x64, 32x96). Forward LFNST is applied to the upper left low-frequency region, called the Region-Of-Interest (ROI). When LFNST is applied, the principal transform coefficients existing in regions other than the ROI are zeroed out. The ROI of LFNST16 consists of six 4x4 sub-blocks that are consecutive in the scan order. Since the number of input samples is 96, the transform matrix of forward LFNST16 can be Rx96. In this example, R is chosen to be 32, so forward LFNST16 generates 32 coefficients accordingly (two 4x4 sub-blocks).
[0056] In some embodiments, an intra-frame prediction mode can be used to determine which core (or matrix) set of mode-dependent transformations, such as the LFNST set, is used.
[0057] For blocks using matrix-based intra-prediction (MIP), the LFNST set index is shown in Figure 6, which illustrates the mapping from the intra-prediction mode to the LFNST set index. DIMD is used to derive the intra-prediction mode of the current block from the MIP prediction samples before upsampling. Specifically, the horizontal and vertical gradients of each prediction sample are computed to build a gradient histogram (HoG). Figure 7 illustrates the construction of the gradient histogram using MIP prediction samples before upsampling. The intra-prediction mode with the largest histogram amplitude is then used to determine the LFNST transform set and the LFNST transpose flag.
[0058] d. DIMD Merging Modes When using DIMD, up to five DIMD modes, such as intra-prediction modes, can be derived by analyzing the directionality of the content surrounding the current block, such as the template of the current block. A gradient histogram (HoG) is computed on an L-shaped template of width / height three samples formed by reconstructed samples. This is obtained using a Sobel filter, accumulating the magnitude of all gradients in a given direction for all samples along the centerline of the current block template. The direction with the highest accumulated magnitude is selected as the primary and secondary DIMD modes. The predictors obtained using the DIMD modes are then blended to form the final DIMD prediction. Uniform or spatial blending is used, where the DIMD predictors are combined with planar predictors, using weights that depend on the relative magnitudes of the modes in the histogram. DIMD merging modes are described below. When using DIMD merging, DIMD information extracted from neighboring blocks is used to compute the intra-prediction of the current block. Specifically, a new merged gradient histogram (MHoG) is computed for the current block based on the HoG of neighboring blocks. Only neighboring blocks encoded or decoded using DIMD or DIMD merging are considered. When a DIMD or DIMD merge neighboring block is available, its gradient histogram is used to form the MHoG for the current block. If multiple DIMD or DIMD merge neighboring blocks are available, the MHoG is derived by combining the corresponding histograms with magnitude averaging. Finally, the intra-prediction modes and weights are calculated using the MHoG, as with regular DIMD. The five highest-magnitude directional modes and their weights in the MHoG are selected, and the corresponding predictors are mixed as with regular DIMD. DIMD merging is offered as an option only if at least one neighbor of the current block uses DIMD or DIMD merge coding. Under these conditions, the use of DIMD merging is marked using a CABAC-coded CU-level flag. A new CABAC context is added to support the encoding and decoding of the DIMD merging flag. In terms of marking, DIMD merging is treated as a sub-mode of DIMD. That is, the DIMD merging flag is marked only if the DIMD flag = 1 and there is a neighboring CU using DIMD or DIMD merge coding.
[0059] III. Inter-frame Prediction a. Merging Candidate List The following inter-frame prediction encoding / decoding tools are listed: Extended Merging Prediction, Merging Mode with MVD (MMVD), Symmetric MVD (SMVD), Signal Affine Motion Compensation Prediction, Sub-block Based Temporal Motion Vector Prediction (SbTMVP), Adaptive Motion Vector Resolution (AMVR), Motion Field Storage: 1 / 16 thLuminance sample MV storage and 8x8 motion field compression with CU-level weighted dual prediction (BCW) bidirectional optical flow (BDOF) decoder-side motion vector refinement (DMVR) geometric segmentation mode (GPM) combined with inter-frame and intra-frame prediction (CIIP). For (regular) merging mode, the merging candidate list can be constructed by sequentially including the following five types of candidates: spatial MVP from spatially neighboring CUs (spatial merging candidate), temporal MVP from co-located CUs (temporal merging candidate), history-based MVP from FIFO table (HMVP merging candidate), pairwise average MVP (pairwise average candidate), and zero MV.
[0060] Figure 8 illustrates the locations of spatial merge candidates. Up to four merge candidates are selected from the locations depicted in the figure. The derived order is B. 0, A 0, B 1, A1 and B2. Position B2 is considered only when the CU of one or more positions B0, A0, B1, and A1 is unavailable (e.g., because it belongs to another slice or tile or is intra-coded). After adding a candidate for position A1, the addition of other candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding and decoding efficiency.
[0061] In addition to the spatial merge candidates mentioned above, non-neighbor spatial merge candidates are inserted after the TMVP (Temporal MVP, such as the temporal merge candidate) in the regular merge candidate list. Figure 9 shows the spatial neighbor blocks used to derive the spatial merge candidates. The distance between the non-neighbor spatial candidate and the current decoded block is based on the width and height of the current decoded block. Line buffer constraints are not applied.
[0062] For time merge candidates, only one candidate is added to the list. Specifically, when deriving this time merge candidate, the scaled motion vector is derived from the co-located CUs belonging to the co-located reference image. The list of reference images used to derive the co-located CUs and their reference indices are explicitly marked in the slice header. Figure 10 illustrates the motion vector scaling of the time merge candidate. The scaled motion vector is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference image (curr_ref) of the current image (curr_pic) and the current image, and td is defined as the POC difference between the reference image of the co-located image (col_pic) and the co-located image. The reference image index of the time merge candidate is set to zero.
[0063] Figure 11 shows the candidate positions for time merge candidates. As shown, the position of a time merge candidate is selected between candidate C0 and C1. Position C1 is used if the coding unit (CU) at position C0 is unavailable, is intra-coded, or is outside the current coding tree unit (CTU) row. Otherwise, position C0 is used to derive a time merge candidate.
[0064] Historical motion vector prediction (HMVP) merging candidates are added to the merging list after spatial motion vector predictions (such as spatial merging candidates and temporal motion vector predictions (TMVP)). In this approach, motion information from previous coded blocks is stored in a table and used as the motion vector prediction for the current coding unit. A table containing multiple HMVP candidates is maintained during encoding / decoding. When a new CTU row is encountered, the table is reset (cleared). Whenever a non-inter-block coded CU is encountered, the associated motion information is added to the last item of the table as a new HMVP candidate.
[0065] Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing list of merged candidates, using the first two merged candidates. The first merged candidate is defined as p0Cand, and the second merged candidate can be defined as p1Cand. Averaged motion vectors are calculated for each reference list based on the availability of motion vectors for p0Cand and p1Cand. If both motion vectors are available in a list, even if they point to different reference images, they are averaged, with the reference image set to the one for p0Cand; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, the list remains invalid. Furthermore, if the half-pixel interpolation filter indices of p0Cand and p1Cand are different, they are set to 0.
[0066] If the merge list is still not full after adding pairwise average (merged) candidates, zero motion vector predictions will be inserted at the end until the maximum number of merged candidates is encountered.
[0067] IV. Proposed Methods for Improving, Using, or Storing DIMD Information Some embodiments of this disclosure provide methods for enhancing DIMD-related information. For example, in some embodiments, information related to DIMD (e.g., histograms, intra-prediction modes such as angular orientation, etc., also referred to as DIMD information or texture analysis information) can be adjusted, refined, or derived using inheritance or derivation to make it more accurate (e.g., resulting in smaller prediction residuals). In some embodiments, DIMD information is enhanced by settings associated with at least one of the following: (1) proposed DIMD information; (2) use of DIMD information inherited from a previous coding block; (3) re-derived DIMD information from a previous coding block; (4) restrictions associated with inheriting or re-deriving DIMD information. These settings are described in Sections IV.a through IV.d below. In some embodiments, one or more of these settings apply to all prediction modes, e.g., to the codec tool whose syntax is true when the DIMD enable flag is equal to true, or to any subset of intra-prediction modes (sub-modes such as DIMD merging modes), e.g., to the intra-codec tool whose syntax is true when the DIMD enable flag is equal to true.
[0068] a. The proposed DIMD information, in some embodiments, for a block containing DIMD information (e.g., an intra-DIMD codec block), the DIMD information (also referred to as, for example, texture analysis information) includes histogram (bar) values of available DIMD intra-prediction modes (such as DC, planar, and / or directional prediction modes) and / or N intra-prediction modes (with the highest N histogram bars) suggested by the histogram values. The directional prediction modes can be within a predefined directional range, e.g., 0 to 64 (thus a total of 65 directional prediction modes), or 0 to 128 or 130 (thus a total of 129 or 131 directional prediction modes). The histogram values of available intra-prediction modes can be obtained by performing texture analysis on a specified region associated with the block, e.g., performing gradient analysis on samples of the block and / or samples of the block's template to obtain a gradient histogram. An example of obtaining a gradient histogram by performing gradient analysis on reconstructed samples of the block's template is texture gradient analysis for preliminary DIMD. Another example of obtaining a gradient histogram by performing gradient analysis on predicted samples of the block is predictor-DIMD.
[0069] As described in Section II.a above, the initial DIMD information based on the neighbor template is calculated based on the current block. In some embodiments, the initial DIMD information is further adjusted based on the inherited DIMD information obtained in Section IV.b and / or the re-derived DIMD information and / or predictor-DIMD obtained in Section IV.c. Adjustment can be performed through at least one combined operation (e.g., adding inherited and / or re-derived DIMD information with predefined weights to the initial DIMD information). In some embodiments, the weights for the initial DIMD information are higher than the weights for the inherited and / or re-derived DIMD information. For example, the ratio of the weights for the initial DIMD information to the inherited and / or re-derived DIMD information is 3:1. In some embodiments, the weights for the initial DIMD information are lower than the weights for the inherited and / or re-derived DIMD information. For example, the ratio of the weights for the initial DIMD information to the inherited and / or re-derived DIMD information is 1:3.
[0070] In some embodiments, the weights vary depending on whether the inherited and / or re-derived DIMD information is "promising". For example, if the reference block for the inherited and / or re-derived DIMD information is intra-coded, the reference block is identified as "promising", and the weights favor the inherited and / or re-derived DIMD information, for example, the weight of the inherited and / or re-derived DIMD information is higher than the weight of the initial DIMD information.
[0071] In some embodiments, the adjustment uses inherited and / or re-derived DIMD information instead of initial DIMD information. In some embodiments, if multiple inherited and / or re-derived DIMD information are available for the current block, the multiple inherited and / or re-derived DIMD information are combined with predefined weights. For example, if multiple DIMD information comes from multiple reference locations or blocks of merge candidates, such as from the merge candidate list in Section IV.b, then DIMD information from the front end of a candidate (referenced before other candidates), such as a spatially adjacent candidate, may have a higher weight in the merge candidate list than other candidates.
[0072] In some embodiments, texture analysis is performed on the predicted samples of the current block to generate predictor-DIMD information. The predictor-DIMD information may be stored and / or referenced by subsequent / following codec blocks. That is, the current block may be a reference block for subsequent / following codec blocks that may use the predictor-DIMD information. The information generated using the predictor-DIMD is texture analysis information, such as histogram values and / or intra-prediction modes, as previously described. The current block may or may not be a DIMD-coded block to have a predictor-DIMD. The current block may be encoded using MIP, intraTMP, cross-componentchroma modes, or a block that uses the predictor-DIMD to select a transform set, as described in Section II.c above, and / or any block, such as any intra- or inter-frame block. In some embodiments, for each block containing DIMD information (e.g., a DIMD-coded block), the DIMD information is stored for reference or inheritance by subsequent / following codec blocks.
[0073] In some embodiments, the predictor-DIMD can be used to adjust or replace the initial DIMD information by weighting it with the initial DIMD information. The predictor-DIMD is performed on all or any subset of the predictors for the current block (which may be intermediate predictors rather than the final predictor).
[0074] In some embodiments, after adjusting / modifying the DIMD information of the current block using the proposed method, the current block can obtain one or more intra-prediction modes using the adjusted DIMD information (e.g., initial DIMD information, predictor-DIMD information, re-derived DIMD information, and / or inherited DIMD information) as part of the regular DIMD process. These one or more intra-prediction modes will be used to generate multiple prediction hypotheses to form the final prediction for the current block.
[0075] In some embodiments, for each predefined unit containing DIMD information (e.g., a k×k grid in a coded block, such as a DIMD coded block, where k can be 2, 4, 8, 16, or any predefined positive integer), the DIMD information is stored and / or referenced by subsequent / following codec blocks. In some embodiments, to reduce storage, instead of storing all DIMD information for reference, only a subset of the DIMD information is stored. For example, only the top 3 or any predefined positive numbers are stored. For example, the subset is the first 3. The first 3 could be the 3 largest in the histogram bars or the 3 suggested intra-prediction modes.
[0076] b. Using Inherited DIMD Information: The inherited DIMD information is obtained from previously encoded blocks. In some embodiments, multiple merge candidates containing DIMD information are established for the current block, for example, a merge candidate list. As done in a conventional inter-frame merging mode, the merge candidates, for example, the merge candidate list, include candidates for spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates, or any subset of the above candidates. These candidates are similar to those described in Section III above.
[0077] Spatial proximity candidates, which may correspond to, but are not limited to, spatial merge candidates, are neighboring blocks of the current block. These neighboring blocks can be the same as the five spatially neighboring blocks in the regular inter-frame merge mode, or any subset of the current block's neighboring blocks. Non-proximity candidates, which may correspond to, but are not limited to, non-proximity spatial merge candidates, are from the search range surrounding (but not adjacent to) the current block. The search range can be the same as the search range of non-proximity candidates in the regular inter-frame merge mode, or a different search range for the current block.
[0078] Historical candidates, which may correspond to, but are not limited to, HMVP merge candidates, are selected from a history-based buffer array. In this history-based buffer array, DIMD information for each valid previous coded block is stored, where a valid previous coded block refers to any block containing DIMD information (e.g., a DIMD-coded block). Similar to historical candidates in the merge list in a regular inter-frame merge mode, if the buffer array is full, the first stored information may be removed to include information from the latest valid coded block. The buffer array is cleared (becomes empty) at the beginning or end of a predefined unit. A predefined unit can be a CTU, slice, tile, picture, or any predefined region. In some embodiments, merge candidates, such as a merge candidate list, refer only to the history buffer array. That is, only historical candidates are included, and / or candidates from far-from-neighbor regions are not used.
[0079] Temporal candidates, which may correspond to, but are not limited to, temporal merge candidates, are obtained from DIMD information stored in one or more predefined previously encoded images. In some embodiments, temporal candidates are only available for inter-frame slices that have predefined previously encoded images, such as for co-position images in a regular inter-frame merge mode.
[0080] The default candidate is a candidate that includes default DIMD information and / or is derived from candidates that have already been examined, for example, by including them in the list of merged candidates.
[0081] In some embodiments, the merge candidates, such as a merge candidate list, are aligned with the merge candidates of the regular inter-frame merging mode, or any subset of the merge candidates, such as the merge candidate list. In some embodiments, full or partial pruning is used to avoid duplicate DIMD information in the list. Before using a candidate, such as before adding a candidate to the list, all or a subset of the DIMD information of the candidate to be added is checked against the corresponding DIMD information of all or any subset of the candidates, such as those already in the list. All DIMD information refers to all stored DIMD information (e.g., the histogram (bar) values of the available DIMD intra-prediction modes and / or the N intra-prediction modes with the highest histogram bars). A subset of DIMD information can be simply 3 or any predefined positive number of all. For example, the subset is the top 3 of all. The top 3 can be the 3 largest in the histogram bars or the 3 suggested intra-prediction modes.
[0082] In some embodiments, after a list of merge candidates is established, one or more candidates are selected from the list for use in the current block. The selection depends on explicitly marking an index or implicitly selecting the first or more candidates or identified “promising” candidates (e.g., the reference block is intra-frame and / or DIMD encoded, and / or the first candidate in the reordered list).
[0083] c. Using Re-derived DIMD Information In some embodiments, if the current block references the DIMD information of a previously encoded block (e.g., a reference block determined by a candidate, which could be any candidate specified in Section IV.b, or possibly from a list of merge candidates), a DIMD procedure is applied to the reference block to re-derive its DIMD information. In this case, the current block may reference the DIMD information of a previously encoded block. In one embodiment, the DIMD information associated with the reference block is derived by analyzing samples associated with the reference block. For example, samples in the reference block and / or samples from neighboring reference blocks are analyzed. In another embodiment, DIMD information does not need to be stored through re-derived information. Re-derived information can completely replace or conditionally replace the method of storing DIMD information.
[0084] d. Restrictions on Inheritance and / or Derivation In some embodiments, the DIMD information of a block (e.g., a previously encoded block or a reference block) can only be inherited / derived and / or potentially used by the current block to reference it if restrictions on the block are satisfied. In some embodiments, only DIMD information from intra-coded blocks and / or blocks containing DIMD information, such as DIMD-coded blocks, can be used. In some embodiments, only DIMD information from blocks located in a predefined region can be used. For example, the predefined region is determined by the current block's position, block width, block height, and / or block area.
[0085] The methods described in this disclosure can be enabled and / or disabled based on implicit rules (e.g., block width, height, or area) or explicit rules (e.g., syntax regarding blocks, tiles, slices, pictures, SPS, or PPS levels). For example, the proposed methods are applied when the block area is less than or greater than a threshold. The term "block" in this invention can refer to TU / TB, CU / CB, PU / PB, predefined regions, or CTU / CTB. Any combination of methods proposed in this invention can be applied.
[0086] Any of the methods proposed above can be implemented in the encoder and / or decoder. For example, any of the proposed methods can be implemented in the inter-frame / intra-frame / IBC / prediction / transform module of the encoder, and / or in the inter-frame / intra-frame / IBC / prediction / transform module of the decoder. Alternatively, any of the proposed methods can be implemented as circuitry connected to the inter-frame / intra-frame / IBC / prediction / transform module of the encoder and / or decoder to provide the information required by the inter-frame / intra-frame / IBC / prediction / transform module.
[0087] V. Example of a video encoder Figure 12 shows an example of a video encoder 1200 that may implement DIMD. As shown, the video encoder 1200 receives the input video signal from the video source 1205 and encodes the signal into a bitstream 1295. The video encoder 1200 has multiple components or modules for encoding signals from the video source 1205, including at least some components selected from the following modules: Transform module 1210, Quantization module 1211, InverseQuantization module 1214, Inverse Transform module 1215, Intra-PictureEstimation module 1220, IntraPrediction module 1225, MotionCompensation module 1230, MotionEstimation module 1235, In-loopFilter 1245, Reconstructed Picture Buffer 1250, MVBuffer 1265, and MVPrediction module 1275, as well as Entropy Encoder 1290. MotionCompensation module 1230 and MotionEstimation module 1235 are part of Inter-Frame Prediction module 1240.
[0088] In some embodiments, modules 1210-1290 are modules of software instructions executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 1210-1290 are modules of hardware circuitry implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 1210-1290 are depicted as separate modules, some modules may be combined into a single module.
[0089] Video source 1205 provides an uncompressed raw video signal, representing the pixel data for each video frame. A subtractor 1208 calculates the difference between the raw video pixel data of video source 1205 and the predicted pixel data 1213 of motion compensation module 1230 or intra-frame prediction module 1225, as the prediction residual 1209. Transformation module 1210 converts the difference (or residual pixel data or residual signal 1208) into transform coefficients (e.g., by performing a discrete cosine transform, or DCT). Quantization module 1211 quantizes the transform coefficients into quantized coefficients 1212, which are encoded into a bitstream 1295 by entropy encoder 1290.
[0090] Inverse quantization module 1214 inverse quantizes quantized data (or quantization coefficients) 1212 to obtain conversion coefficients. Inverse conversion module 1215 inversely converts the conversion coefficients to generate reconstruction residual 1219. Reconstruction residual 1219 is added to predicted pixel data 1213 to generate reconstructed pixel data 1217. In some embodiments, reconstructed pixel data 1217 is temporarily stored in an online buffer 1227 (or intra-frame prediction buffer) for intra-frame image prediction and spatial MV prediction. Reconstructed pixels are filtered via loop filter 1245 and stored in reconstructed image buffer 1250. In some embodiments, reconstructed image buffer 1250 is a storage device external to video encoder 1200. In some embodiments, reconstructed image buffer 1250 is a storage device internal to video encoder 1200.
[0091] Intra-frame image estimation module 1220 performs intra-frame prediction based on reconstructed pixel data 1217 to generate intra-frame prediction data. The intra-frame prediction data is provided to entropy encoder 1290 for encoding into bitstream 1295. The intra-frame prediction data is also used by intra-frame prediction module 1225 to generate predicted pixel data 1213.
[0092] The motion estimation module 1235 performs inter-frame prediction by generating MVs (Motion Values) that reference pixel data of previously decoded frames stored in the reconstructed image buffer 1250. These MVs are provided to the motion compensation module 1230 to generate predicted pixel data.
[0093] The video encoder 1200 does not encode the complete actual MVs in the bitstream. Instead, it uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 1295.
[0094] The MV prediction module 1275 generates predicted MVs (Motion Compensation MVs) based on the reference MVs generated for encoding previous movie frames. The MV prediction module 1275 retrieves the reference MVs from the previous movie frames in the MV buffer 1265. The movie encoder 1200 stores the MVs generated for the current movie frame in the MV buffer 1265 as reference MVs for generating the predicted MVs.
[0095] The MV prediction module 1275 uses reference MVs to create predicted MVs. Predicted MVs can be calculated by spatial MV prediction or temporal MV prediction. The difference (residual motion data) between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame is encoded into bitstream 1295 by the entropy encoder 1290.
[0096] The entropy encoder 1290 uses entropy encoding techniques such as context-adaptive binary arithmetic encoding (CABAC) or Huffman coding to encode various parameters and data into a bitstream 1295. The entropy encoder 1290 encodes various header elements, flags, quantization conversion coefficients 1212, and residual motion data as syntax elements into the bitstream 1295. The bitstream 1295 is then stored in a storage device or transmitted to the decoder via a communication medium such as a network.
[0097] Loop filter 1245 filters or smooths the reconstructed pixel data 1217 to reduce encoding / decoding distortion, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by loop filter 1245 include deblock filter (DBF), sample adaptive offset (SAO), and / or adaptive loop filter (ALF). In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.
[0098] Figure 13 illustrates a portion of the movie encoder 1200 that implements intra-prediction based on proposed DIMD information. The DIMD module 1310 may receive pixel samples from neighboring reconstructed samples of a Reconstructed Picture Buffer 1250 or an Intra Prediction Buffer 1227, and use these samples to perform DIMD operations (e.g., texture analysis) and provide DIMD information such as histogram values (e.g., HoG) and / or one or more intra-prediction modes. The DIMD module 1310 may perform DIMD on neighboring pixel samples of the current block to generate preliminary DIMD information 1312 (referring to DIMD information generated using preliminary DIMD). DIMD module 1310 may perform DIMD based on pixels related to a reference block (in a previously encoded image or in the current image), such as neighboring pixels of the reference block, to generate Re-Derive DIMD info 1314 (referring to DIMD information generated using re-derived DIMD). DIMD module 1310 may also perform DIMD based on predicted pixels of the block to generate Predictor DIMD info 1313 (referring to DIMD information generated using predictor DIMD). The initial DIMD info 1312 and / or the re-derived DIMD info 1314 and / or the predictor DIMD info 1313 may be stored in Candidate Storage 1335 and / or may be inherited and used by subsequent encoded blocks. In some embodiments, DIMD information derived from a block can only be inherited or re-derived and / or potentially referenced by subsequent blocks if the block meets certain restrictions. For example, only DIMD information from DIMD or intra-coded blocks can be inherited and used, or only DIMD information from blocks located in predefined regions can be used.
[0099] The Candidate Constructor 1330 identifies multiple candidates, such as the current block's Candidate List 1345 (or merged candidate list), which includes spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates, or any subset thereof. DIMD information associated with these candidates can be inherited by the current block. The entropy encoder 1290 may provide a selection index to the Candidate Selector 1348. The selection index may also be signaled in the bitstream 1295. The Candidate Selector 1348 extracts the corresponding DIMD information of the selected candidate from the candidate storage 1335 and provides it as Inherited DIMD info 1316 (referring to DIMD information obtained using inheritance).
[0100] DIMD Adjust 1340 uses preliminary DIMD information 1312, predictor-DIMD information 1313, re-derived DIMD information 1314, and / or inherited DIMD information 1316 to generate a set of final DIMD information (FinalDIMD Info) 1350. Section IV above describes different embodiments of adjusting DIMD information based on preliminary DIMD information, predictor-DIMD information, re-derived DIMD information, and / or inherited DIMD information. For example, in some embodiments, preliminary DIMD information (e.g., histogram values or intra-prediction modes) is adjusted based on inherited or derived DIMD information or predictor-DIMD information. In some embodiments, inherited or derived DIMD information or predictor-DIMD information is used as the final DIMD information instead of the preliminary DIMD information. In some embodiments, multiple inherited or derived DIMD information sets are combined according to predefined weights.
[0101] The intra prediction module 1225 may use the final DIMD information 1350 to perform intra prediction and generate intra prediction samples as pixel data 1213 for prediction. For example, the intra prediction module may use the histogram or intra prediction mode in the final DIMD information 1350 to identify the intra prediction mode of the intra prediction predictor that generated the current block. The final DIMD information 1350 that may be used to encode the current block may be stored in candidate storage 1335 and / or may be inherited by subsequent blocks.
[0102] Figure 14 conceptually illustrates a process 1400 of encoding pixel blocks using adjusted texture analysis information (e.g., DIMD information). In some embodiments, one or more processing units (e.g., processors) of a computing device implementing encoder 1200 execute process 1400 by executing instructions stored in a computer-readable medium. In some embodiments, an electronic device implementing encoder 1200 executes process 1400.
[0103] The encoder receives (in block 1410) data to encode the pixels of the current block in the current image of the movie. The encoder analyzes (in block 1420) samples neighboring the current block to generate a first set of texture analysis (DIMD) information. In some embodiments, the set of texture analysis information may be DIMD information or any other type of texture analysis information, which may include histogram values, intra-prediction modes, or both.
[0104] The encoder acquires (in block 1430) or derives at least one set of texture analysis information associated with the current block or at least one reference block. The at least one reference block is determined based on at least one candidate selected from a candidate list. The encoder may examine one or more candidates for the current block. These candidates may include all or any subset of spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates. In some embodiments, at least one of blocks 1420 and 1430 may be performed to determine the texture analysis information.
[0105] At least one set of second texture analysis information may be inherited from at least one reference block, or from at least one candidate selected from a candidate list. In some embodiments, the inheritance of at least one set of second texture analysis information is permitted only if at least one reference block is intra-predictively coded or contains texture analysis information. In some embodiments, the inheritance of at least one set of second texture analysis information is permitted only if at least one reference block is located in a predefined region.
[0106] In some embodiments, at least one set of second texture analysis information is derived by analyzing samples associated with at least one reference block. In some embodiments, at least one set of second texture analysis information is derived by analyzing predicted samples of the current block or at least one reference block.
[0107] The encoder generates (in block 1440) a final texture analysis set based on the first set of texture analysis information, the second set of texture analysis information, or both. Multiple sets of texture analysis information may be combined to form the final texture analysis information. The encoder may store the final texture analysis information for inheritance or reference by subsequently encoded blocks; the stored texture analysis information is a subset of the original texture analysis information.
[0108] The encoder encodes (e.g., using or storing) the final texture analysis information for the current block (in block 1450) or stores the final texture analysis information for subsequent encoding. The final texture analysis information may include one or more intra-prediction patterns that determine multiple prediction hypotheses to form the final prediction for the current block. The encoder may generate the final prediction based on the intra-prediction patterns in the final texture analysis information. The final prediction may be used to generate the prediction residual for the current block.
[0109] VI. Video Decoder Examples In some embodiments, the encoder may signal (or generate) one or more syntax elements in the bitstream, such that the decoder can parse the specified one or more syntax elements from the bitstream.
[0110] Figure 15 illustrates an example of a video decoder 1500 that may implement DIMD. As shown, the video decoder 1500 is an image or video decoding circuit that receives a bitstream 1595 and decodes the contents of the bitstream into pixel data of video frames for display. The video decoder 1500 has several components or modules for decoding the bitstream 1595, including some components selected from the inverse quantization module 1511, the inverse transform module 1510, the intra prediction module 1525, the motion compensation module 1530, the in-loop filter 1545, the decoded picture buffer 1550, the MV buffer 1565, the MV prediction module 1575, and the parser 1590. The motion compensation module 1530 is part of the motion prediction module 1540.
[0111] In some embodiments, modules 1510-1590 are software instruction modules executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 1510-1590 are hardware circuit modules implemented by one or more integrated circuits of an electronic device. Although modules 1510-1590 are depicted as separate modules, some modules may be combined into a single module.
[0112] Parser 1590 (or entropy decoder) receives bitstream 1595 and performs initial parsing according to the syntax defined by the video-encoding or image-encoding standard. The parsed syntax elements include various header elements, flags, and quantization data (or quantization coefficients) 1512. Parser 1590 uses entropy encoding techniques such as context-adaptive binaryarithmetic encoding (CABAC) or Huffman coding to parse the various syntax elements.
[0113] Inverse quantization module 1511 inversely quantizes the quantized data (or quantized coefficients Quatized Coef.) 1512 to obtain transform coefficients. Inverse transform module 1510 inversely transforms the transform coefficients 1516 to generate a reconstructed residual signal 1519. The reconstructed residual signal 1519 is added to predicted pixel data 1513 from intra-frame prediction module 1525 or motion compensation module 1530 to generate reconstructed pixel data 1517. The reconstructed pixel data is filtered by loop filter 1545 and stored in a deconstructed image buffer 1550. In some embodiments, the deconstructed image buffer 1550 is external storage to the video decoder 1500. In some embodiments, the deconstructed image buffer 1550 is internal storage to the video decoder 1500.
[0114] Intra-prediction module 1525 receives intra-prediction data from bitstream 1595 and generates predicted pixel data 1513 from decoded pixel data 1517 stored in decoded image buffer 1550 based on this data. In some embodiments, decoded pixel data 1517 is also stored in online buffer 1527 (or intra-prediction buffer) for intra-image prediction and spatial MV prediction.
[0115] In some embodiments, the contents of the decoded image buffer 1550 are used for display. The display device 1505 either retrieves the contents of the decoded image buffer 1550 directly for display or retrieves the contents of the decoded image buffer into a display buffer. In some embodiments, the display device receives pixel values from the decoded image buffer 1550 via pixel transfer.
[0116] The motion compensation module 1530 generates predicted pixel data 1513 from the decoded pixel data 1517 stored in the decoded image buffer 1550 based on the motion compensation MV (MC MV). These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1595 to the predicted MV received from the MV prediction module 1575.
[0117] The MV prediction module 1575 generates a predicted MV based on a reference MV used for decoding a previous video frame, such as a motion-compensated MV used for motion compensation. The MV prediction module 1575 retrieves the reference MV of the previous video frame from the MV buffer 1565. The video decoder 1500 stores the motion-compensated MV used for decoding the current video frame in the MV buffer 1565 as a reference MV for generating the predicted MV.
[0118] Loop filter 1545 filters or smooths the decoded pixel data 1517 to reduce encoding / decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering or smoothing operations performed by loop filter 1545 include deblock filter (DBF), sample adaptive offset (SAO), and / or adaptive loop filter (ALF). In some embodiments, luma mapping chromascaling (LMCS) is performed before the loop filters.
[0119] Figure 16 illustrates a portion of the video decoder 1500 that implements intra-frame prediction based on proposed DIMD information. The DIMD module 1610 may receive pixel samples from the Decoded Picture Buffer 1550 or neighboring reconstructed samples from the Intra Prediction Buffer 1527, and use these samples to perform DIMD operations (e.g., texture analysis) and provide DIMD information, such as histogram values (e.g., HoG) and / or one or more intra-frame prediction modes. The DIMD module 1610 may perform DIMD on neighboring pixel samples of the current block to generate preliminary DIMD information 1612 (referring to DIMD information generated using preliminary DIMD). The DIMD module 1610 may perform DIMD based on pixels associated with a reference block (in a previously encoded picture or in the current picture), such as neighboring pixels of the reference block, to generate re-derived DIMD information 1614 (referring to DIMD information generated using re-derived DIMD). DIMD module 1610 may perform DIMD based on the predicted pixels of a block to generate predictor DIMD information 1613 (referring to DIMD information generated using predictor DIMD). Preliminary DIMD information 1612 and / or re-derived DIMD information 1614 and / or predictor DIMD information 1613 may be stored in candidate storage 1635 and / or may be inherited and used by subsequently encoded blocks. In some embodiments, DIMD information derived from a block can only be inherited or re-derived and / or may be referenced and used by subsequent blocks if the block meets certain conditions; for example, only DIMD information from DIMD or intra-coded blocks can be inherited and used, or only DIMD information from blocks located in predefined regions can be used.
[0120] The candidate constructor 1630 identifies multiple candidates, such as the candidate list 1645 (or merged candidate list) of the current block, including spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates, or any subset thereof. DIMD information associated with these candidates can be inherited by the current block. The entropy decoder 1590 may provide a selection index to the candidate selector 1648 based on the content of the bitstream 1595. The candidate selector 1648 extracts the corresponding DIMD information of the selected candidate from the candidate storage 1635 and provides it as inherited DIMD information 1616 (referring to DIMD information obtained using inheritance).
[0121] DIMD Adjust 1640 generates a set of final DIMD information (Final DIMDInfo) 1650 using preliminary DIMD information 1612, predictor-DIMD information 1613, re-derived DIMD information 1614, and / or inherited DIMD information 1616. Section IV above describes adjusting DIMD information in different embodiments based on preliminary DIMD information, predictor-DIMD information, re-derived DIMD information, and / or inherited DIMD information. For example, in some embodiments, preliminary DIMD information (e.g., histogram values or intra-prediction modes) is adjusted based on inherited or derived DIMD information or predictor-DIMD information. In some embodiments, inherited or derived DIMD information or predictor-DIMD information is used as final DIMD information instead of preliminary DIMD information. In some embodiments, multiple inherited or derived DIMD information are combined according to predefined weights.
[0122] The intra prediction module 1525 may use the final DIMD information 1650 to perform intra prediction and generate intra prediction samples as predicted pixel data 1513. For example, the intra prediction module may use the histogram or intra prediction mode in the final DIMD information 1650 to identify the intra prediction mode to generate an intra prediction predictor for the current block. The final DIMD information 1650 that may be used to decode the current block may be stored in candidate storage 1635 and / or may be inherited by subsequent blocks.
[0123] Figure 17 conceptually illustrates process 1700 for decoding pixel blocks using adjusted DIMD information. In some embodiments, one or more processing units (e.g., processors) of a computing device executing decoder 1500 execute process 1700 by executing instructions stored in a readable computing medium. In some embodiments, an electronic device implementing decoder 1500 executes process 1700.
[0124] The decoder receives (in block 1710) data representing the pixels of the current block in the current image and decodes it. The decoder analyzes (in block 1720) samples neighboring the current block to generate a first set of texture analysis information. In some embodiments, the set of texture analysis information may be DIMD information or any other type of texture analysis information, which may include histogram values, intra-frame prediction modes, or both.
[0125] The decoder acquires (in block 1730) or derives at least one set of second texture analysis information associated with the current block or at least one reference block. The at least one reference block is determined based on at least one candidate selected from a candidate list. The decoder may examine one or more candidates for the current block. These candidates may include all or any subset of spatially neighboring candidates, non-neighboring candidates, historical candidates, temporal candidates, and preset candidates. In some embodiments, at least one of blocks 1720 and 1730 is performed to determine the texture analysis information.
[0126] At least one set of second texture analysis information may be inherited from at least one reference block, or from at least one candidate selected from a candidate list. In some embodiments, at least one set of second texture analysis information is allowed to be inherited only if at least one reference block is intra-predictively coded or contains texture analysis information. In some embodiments, at least one set of second texture analysis information is allowed to be inherited only if at least one reference block is located in a predefined region.
[0127] In some embodiments, at least one set of second texture analysis information is derived by analyzing samples associated with at least one reference block. In some embodiments, at least one set of second texture analysis information is derived by analyzing predicted samples of the current block or at least one reference block.
[0128] The decoder generates a final set of texture analysis information (in block 1740) based on the first set of texture analysis information, the second set of texture analysis information, or both. Multiple sets of texture analysis information may be combined into the final set of texture analysis information. The decoder may store the final set of texture analysis information for inheritance or reference in subsequent encoded blocks; the stored final set of texture analysis information may be a subset of the existing texture analysis information.
[0129] The decoder reconstructs the current block (in block 1750) based on the final set texture analysis information or stores the final set texture analysis information for subsequent decoding. The final set texture analysis information may include one or more intra-frame prediction modes, which are determined to generate multiple prediction hypotheses to form the final prediction for the current block. The decoder may generate the final prediction based on the intra-frame prediction modes in the final set texture analysis information. The decoder may then provide the reconstructed current block as part of the reconstructed current image for display.
[0130] VII. Exemplary Electronic Systems Many of the features and applications described above are implemented as software processes, which are specified as a set of instructions recorded on a computer-readable storage medium (also known as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing unit to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disk drives, erasable programmable read / write memory (EPROM), electrically erasable programmable read / write memory (EEPROM), etc. Computer-readable media do not include carrier waves and electronic signals transmitted in wireless or wired connections.
[0131] In this specification, the term "software" means including firmware stored in read-only memory or applications stored in magnetic storage, which can be read into memory for processor processing. Furthermore, in some embodiments, multiple software inventions may be implemented as sub-parts of a larger program while remaining independent software inventions. In some embodiments, multiple software inventions may also be implemented as separate programs. Finally, any combination of separate programs that together implement the software inventions described herein is within the scope of this disclosure. In some embodiments, when software programs are installed and run on one or more electronic systems, they define one or more specific machine implementations that execute and perform the operations of the software programs.
[0132] Figure 18 conceptually illustrates an electronic system 1800 implementing certain embodiments of the present disclosure. The electronic system 1800 may be a calculator (e.g., a desktop calculator, personal calculator, tablet calculator, etc.), a telephone, a PDA, or any other type of electronic device. Such an electronic system includes various calculator-readable media and interfaces for various other types of calculator-readable media. The electronic system 1800 includes a bus 1805, a processing unit(s) 1810, a graphics processing unit (GPU) 1815, system memory 1820, a network 1825, read-only memory (ROM) 1830, permanent storage 1835, input devices 1840, and output devices 1845.
[0133] The bus 1805 set represents all system, peripheral, and chipset buses that communicate with numerous devices within the electronic system 1800. For example, bus 1805 communicates with the processing unit 1810 and GPU 1815, read-only memory 1830, system memory 1820, and permanent storage device 1835.
[0134] From these various memory units, processing unit 1810 retrieves execution instructions and processing data to perform the processes of this disclosure. The processing unit may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by GPU 1815. GPU 1815 may offload various computational or supplemental image processing provided by processing unit 1810.
[0135] Read-only memory (ROM) 1830 stores static data and instructions used by processing unit 1810 and other modules of the electronic system. On the other hand, permanent storage device 1835 is a read-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 1800 is powered off. Some embodiments of this disclosure use mass storage devices (such as disks or optical discs and their corresponding disk drives) as permanent storage device 1835.
[0136] Other embodiments use removable storage devices (such as floppy disks, flash memory devices, etc., and their corresponding disk drives) as permanent storage devices. Like permanent storage device 1835, system memory 1820 is a read-write memory device. However, unlike storage device 1835, system memory 1820 is volatile read-write memory, such as random access memory. System memory 1820 stores some instructions and data used by the processor during runtime. In some embodiments, processes according to this disclosure are stored in system memory 1820, permanent storage device 1835, and / or read-only memory 1830. For example, various memory units include instructions for processing multimedia clips according to some embodiments. From these various memory units, processing unit 1810 retrieves execution instructions and processing data to perform the processes of some embodiments.
[0137] Bus 1805 is also connected to input and output devices 1840 and 1845. Input device 1840 allows users to communicate information and select commands to the electronic system. Input device 1840 includes an alphanumeric keypad and pointing device (also known as a "cursor control device"), a camera (e.g., a webcam), a microphone, or similar device to receive voice commands, etc. Output device 1845 displays images generated by the electronic system or otherwise outputs data. Output device 1845 includes printers and display devices such as cathode ray tube (CRT) or liquid crystal display (LCD), as well as speakers or similar audio output devices. Some embodiments include devices such as touchscreens that function as both input and output devices.
[0138] Finally, as shown in Figure 18, bus 1805 also connects electronic system 1800 to network 1825 via a network adapter (not shown). In this way, the computer can be part of a computer network (e.g., a local area network ("LAN"), a wide area network ("WAN"), an intranet, or a network of networks such as the Internet). Any or all components of electronic system 1800 can be used in conjunction with this disclosure.
[0139] Some embodiments include electronic components, such as microprocessors, storage, and memory, that store calculator program instructions in a machine-readable or calculator-readable medium (also referred to as a calculator-readable storage medium, machine-readable medium, or machine-readable storage medium). Examples of such calculator-readable media include RAM, ROM, read-only optical disc (CD-ROM), recordable optical disc (CD-R), rewritable optical disc (CD-RW), read-only digital multipurpose optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, micro-SD card, etc.), magnetic and / or solid-state drives, read-only and recordable Blu-ray® optical discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. A calculator-readable medium may store a calculator program executed by at least one processing unit, which includes a set of instructions for performing various operations. Examples of calculator programs or calculator code include machine code, such as that generated by a compiler, and files containing higher-level code executed by a calculator, electronic component, or microprocessor using an interpreter.
[0140] While the foregoing discussion primarily refers to microprocessors or multi-core processors that execute software, many of the features and applications described above are executed by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, these integrated circuits execute instructions stored on the circuit itself. Furthermore, some embodiments execute software stored in programmable logic devices (PLDs), ROM, or RAM devices.
[0141] As used in this specification and any claim of this application, the terms “calculator,” “server,” “processor,” and “memory” refer to electronic or other technical devices. These terms do not include people or groups of people. For the purposes of this specification, the term “display” or “show” means “displayed on an electronic device.” As used in this specification and any claim of this application, the terms “calculator-readable medium,” “computer-readable medium,” and “machine-readable medium” are entirely limited to tangible, physical objects that store information in a calculator-readable form. These terms do not include any wireless signals, wired download signals, or any other transient signals.
[0142] While this disclosure has been described with reference to numerous specific details, those skilled in the art will recognize that this disclosure may be embodied in other specific forms without departing from its spirit. Furthermore, many figures (including Figures 14 and 17) conceptually illustrate the processes. The specific operations of these processes may not be performed in the exact order shown and described. Specific operations may not be performed in a continuous series of operations, and different specific operations may be performed in different embodiments. Moreover, the process may be implemented using several sub-processes or as part of a larger macro-process. Therefore, those skilled in the art will understand that this disclosure is not limited to the foregoing illustrative details but should be defined by the appended claims.
[0143] Additional Notes: The topics described herein sometimes demonstrate different components contained within or connected to other components. It should be understood that the architectures depicted are merely examples, and many other architectures can actually be implemented to achieve the same functionality. Conceptually, any arrangement of components to achieve the same functionality is actually “related” in order to achieve the desired function. Therefore, any two components combined here to achieve a particular function can be considered “related” to each other to achieve the desired function, regardless of the architecture or intermediate components. Similarly, any two such related components can also be considered “operably connected,” or “operably coupled,” to each other to achieve the desired function, and any two components that can be so related can also be considered “operably coupled” to each other to achieve the desired function. Specific examples of operational coupling include, but are not limited to, physically matable and / or physically interactive components, wirelessly interactive and / or wirelessly interactive components, and logically interactive and / or logically interactive components.
[0144] Furthermore, regarding any substantially plural and / or singular terms used herein, those skilled in the art may translate from plural to singular and / or from singular to plural depending on the context and / or application. Various singular / plural arrangements may be explicitly set forth herein for clarity.
[0145] Furthermore, those skilled in the art will understand that the terms used herein, particularly in appended claims, such as the body of an appended claim, are generally considered "open" terms. For example, the word "comprising" should be interpreted as "comprising but not limited to," the word "having" should be interpreted as "having at least," and the word "includes" should be interpreted as "including but not limited to," etc. Moreover, those skilled in the art will also understand that if there is an intent to introduce a particular number of claim statements, this intent will be explicitly stated in the claims, and without such statements, this intent does not exist. For example, to aid understanding, the following appended claims may contain the use of the introductory phrases "at least one" and "one or more" to introduce claim statements. However, the use of these phrases should not be construed as meaning that any particular claim introducing a claim statement by the indefinite article "a" or "an" is limited to containing only one such statement, even if the same claim includes the introductory phrase "one or more" or "at least one" and indefinite articles such as "a" or "an," for example, "a" and / or "an" should be interpreted as "at least one" or "one or more"; the same applies to the use of definite articles used to introduce claim statements. Furthermore, even if a specific number of claims is explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as including at least that number. For example, simply stating "two statements" without any other modifiers means at least two statements, or two or more statements. Moreover, when using conventions such as "at least one of A, B, and C," this construction is generally intended according to conventions that will be understood by those skilled in the art. For example, "a system having at least A, B, and C" includes, but is not limited to, systems with only A, only B, only C, A and B together, A and C together, B and C together, and / or systems with A, B, and C together. Similarly, when using conventions such as "at least one of A, B, or C," this construction is generally intended according to conventions that will be understood by those skilled in the art. For example, "a system having at least A, B, or C" includes, but is not limited to, systems with only A, only B, only C, A and B together, A and C together, B and C together, and / or systems with A, B, and C together. Those skilled in the art will further understand that virtually any divergent word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to include the possibility of containing one, any, or both terms. For example, the phrase "A or B" will be understood to include the possibility of "A" or "B" or "A and B". As can be seen from the foregoing, various embodiments of this disclosure have been described herein for illustrative purposes, and various modifications may be made without departing from the scope and spirit of this disclosure. Therefore, the various embodiments disclosed herein are not intended to be restrictive, and the true scope and spirit are indicated by the following claims.
Claims
1. A video encoding / decoding method, comprising: Receive data as pixels of the current block of the current image in the video, encoded or decoded. Determine texture analysis information, including at least one of the following: analyze samples adjacent to the current block to generate a first set of texture analysis information; The process involves obtaining or deriving at least one set of second texture analysis information related to the current block or at least one reference block; generating a final set of texture analysis information based on the first set of texture analysis information, the at least one set of second texture analysis information, or both; and encoding or decoding the current block or storing the final set of texture analysis information according to the final set of texture analysis information.
2. The video encoding / decoding method as described in claim 1, wherein the at least one set of second texture analysis information is inherited from the at least one reference block.
3. The video encoding / decoding method as described in claim 1, wherein the at least one set of second texture analysis information is inherited from at least one candidate selected from a candidate list.
4. The video encoding / decoding method of claim 1, wherein the at least one set of second texture analysis information is allowed to be inherited only when the at least one reference block is intra-frame predictive coded or contains texture analysis information.
5. The video encoding / decoding method as claimed in claim 1, wherein the at least one set of second texture analysis information is allowed to be inherited only when the at least one reference block is located in a predefined region.
6. The video encoding / decoding method of claim 1, wherein the at least one set of second texture analysis information is derived by analyzing samples associated with the at least one reference block.
7. The video encoding / decoding method of claim 1, wherein the at least one set of second texture analysis information is derived by analyzing prediction samples of the current block or the at least one reference block.
8. The video encoding / decoding method of claim 1, wherein the at least one reference block is determined based on at least one candidate selected from a candidate list.
9. The video encoding / decoding method of claim 1 further includes checking one or more candidates of the current block, the one or more candidates including all or any subset of spatially nearby candidates, non-nearby candidates, historical candidates, temporal candidates and preset candidates.
10. The video encoding / decoding method as described in claim 1, wherein multiple texture analysis information sets are combined to form the final set of texture analysis information.
11. The video encoding / decoding method as described in claim 1, further comprising inheriting or referencing the stored final group texture analysis information by subsequently encoded blocks.
12. The video encoding / decoding method as described in claim 1, wherein a set of texture analysis information includes histogram values, intra-frame prediction modes, or both.
13. The video encoding / decoding method as claimed in claim 1, wherein the stored final set of texture analysis information is a subset of the texture analysis information.
14. The video encoding / decoding method of claim 1, wherein encoding or decoding the current block using the final set of texture analysis information includes determining one or more intra-frame prediction modes to generate multiple prediction hypotheses to form a final prediction for the current block.
15. An electronic device comprising: A video encoder / decoder circuit is configured to perform operations including: receiving data to encode or decode pixels of a current block of a current image of a video; determining texture analysis information including at least one of: analyzing samples adjacent to the current block to generate a first set of texture analysis information; and acquiring or deriving at least one set of second texture analysis information associated with the current block or at least a reference block; generating a final set of texture analysis information based on the first set of texture analysis information, the at least one set of second texture analysis information, or both; and encoding or decoding or storing the final texture analysis information based on the final texture analysis information of the current block.
16. A video decoding method includes: Receive data to decode it into pixels of the current block of the current image in the video; Determining texture analysis information includes at least one of the following: analyzing samples adjacent to the current block to generate a first set of texture analysis information; And to acquire or derive at least one set of second texture analysis information related to the current block or at least one reference block; Based on the first set of texture analysis information, the at least one set of second texture analysis information, or both, a final set of texture analysis information is generated. And reconstruct the current block or store the final texture analysis information based on the final texture analysis information of the current block.