Referencing block vectors
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- MEDIATEK INC
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-23
Smart Images

Figure CN2026073053_23072026_PF_FP_ABST
Abstract
Description
REFERENCING BLOCK VECTORSCROSS REFERENCE TO RELATED PATENT APPLICATION (S)
[0001] The present disclosure is part of a non-provisional application that claims the priority benefit of U. S. Provisional Patent Application No. 63 / 745, 841, filed on 16 January 2025. Contents of above-listed applications are herein incorporated by reference.TECHNICAL FIELD
[0002] The present disclosure relates generally to video coding. In particular, the present disclosure relates to methods of coding pixel blocks by using block vectors to perform prediction.BACKGROUND
[0003] Unless otherwise indicated herein, approaches described in this section are not prior art to the claims listed below and are not admitted as prior art by inclusion in this section.
[0004] High-Efficiency Video Coding (HEVC) is an international video coding standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) . HEVC is based on the hybrid block-based motion-compensated DCT-like transform coding architecture. The basic unit for compression, termed coding unit (CU) , is a 2Nx2N square block of pixels, and each CU can be recursively split into four smaller CUs until the predefined minimum size is reached. Each CU contains one or multiple prediction units (PUs) .
[0005] Versatile video coding (VVC) is the latest international video coding standard developed by the Joint Video Expert Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11. The input video signal is predicted from the reconstructed signal, which is derived from the coded picture regions. The prediction residual signal is processed by a block transform. The transform coefficients are quantized and entropy coded together with other side information in the bitstream. The reconstructed signal is generated from the prediction signal and the reconstructed residual signal after inverse transform on the de-quantized transform coefficients. The reconstructed signal is further processed by in-loop filtering for removing coding artifacts. The decoded pictures are stored in the frame buffer for predicting the future pictures in the input video signal.
[0006] In VVC, a coded picture is partitioned into non-overlapped square block regions represented by the associated coding tree units (CTUs) . The leaf nodes of a coding tree correspond to the coding units (CUs) . A coded picture can be represented by a collection of slices, each comprising an integer number of CTUs.
[0007] Each CU contains one or more prediction units (PUs) . The prediction unit, together with the associated CU syntax, works as a basic unit for signaling the predictor information. The specified prediction process is employed to predict the values of the associated pixel samples inside the PU.
[0008] For each inter-predicted CU, motion parameters consisting of motion vectors, reference picture indices and reference picture list usage index, and additional information are used for inter-predicted sample generation. The motion parameter can be signalled in an explicit or implicit manner. When a CU is coded with skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A merge mode is specified whereby the motion parameters for the current CU are obtained from neighbouring CUs, including spatial and temporal candidates, and additional schedules introduced in VVC. The merge mode can be applied to any inter-predicted CU. The alternative to merge mode is the explicit transmission of motion parameters, where motion vector, corresponding reference picture index for each reference picture list and reference picture list usage flag and other needed information are signalled explicitly per each CU.SUMMARY
[0009] The following summary is illustrative only and is not intended to be limiting in any way. That is, the following summary is provided to introduce concepts, highlights, benefits and advantages of the novel and non-obvious techniques described herein. Select and not all implementations are further described below in the detailed description. Thus, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended for use in determining the scope of the claimed subject matter.
[0010] Some embodiments provide methods of coding pixel blocks by referencing block vectors (BVs) . A video coder derives or selects one or more first target block vectors for the current block. The video coder locates one or more reference areas in the current picture by using the first target block vectors. The video coder generates a final predictor of the current block based on coding information derived from samples of the located reference areas and uses the final predictor to encode or decode the current block. The video coder encodes or decodes a subsequent block by referencing the first target block vectors.
[0011] The first target block vectors may be used for intra block copy prediction of the current block. The first target block vectors may be derived by searching the current picture for template areas matching that of the current block. The video coder may select the one or more first target block vectors from an intra prediction mode list of a partition mode of spatial geometric partition mode, the IPM list augmented with BV-based prediction candidates obtained from adjacent and non-adjacent merge candidates coded in IBC or IntraTMP.
[0012] The coding information may include a filter-based model, a parameter model, a cross-component model, or an intra prediction mode. The coding information may include one or more BV-based predictors that are generated based on sample values of the reference areas located by the first target block vectors.
[0013] The final predictor may be a combination of the one or more BV-based predictors that are generated based on the first target block vectors and one or more other predictors. The other predictors may be BV-based predictors or non-BV-based predictors. The one or more other predictors may be generated based on one or more intra direction modes that are determined from a histogram of gradient (HoG) based on a texture gradient analysis (e.g., DIMD process) of samples in the one or more reference areas located by the first target block vectors.
[0014] The subsequent block may be coded by using a merge candidate list comprising one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the second target block vectors including the first target block vectors that are stored for the current block. The subsequent block may be coded based on one or more block vectors selected from one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the neighboring blocks of the subsequent block including the current block and the one or more second target block vectors including the first target block vectors stored for the current block.
[0015] In some embodiments, BV-based predictors may be generated for the subsequent block based on sample values of the reference areas located by the one or more block vectors selected from the second target block vectors, and a final predictor of the subsequent block is a combination of the BV-based predictors and one or more other (BV-based or non-BV-based) predictors. The one or more other predictors may be generated based on one or more intra prediction modes that are determined from a HoG based on a texture gradient analysis of samples in a reference area located by the selected one or more block vectors.BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are included to provide a further understanding of the present disclosure, and are incorporated in and constitute a part of the present disclosure. The drawings illustrate implementations of the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is appreciable that the drawings are not necessarily in scale as some components may be shown to be out of proportion than the size in actual implementation in order to clearly illustrate the concept of the present disclosure.
[0017] FIG. 1 shows an example derivation of auto-relocated block vector prediction.
[0018] FIG. 2 illustrates five positions that are used to derive auto-relocated block vector prediction.
[0019] FIG. 3 illustrates padding candidates for replacing the zero-vector in the intra block copy (IBC) list.
[0020] FIG. 4 shows IBC candidate clustering.
[0021] FIG. 5 illustrates the search regions for IntraTMP for a current block.
[0022] FIG. 6 illustrates using IntraTMP block vector for an IBC block.
[0023] FIG. 7 illustrates a defined type of filter shape having 15 inputs and generating one output for extended intra prediction (EIP) .
[0024] FIG. 8 conceptually illustrates a diagonal order according to which an EIP filter is applied to generate a predictor for the current block.
[0025] FIG. 9 illustrates a reference area used in the block-vector guided EIP.
[0026] FIG. 10 conceptually illustrates a prediction using block vector guided DIMD.
[0027] FIG. 11 illustrates the reference area for calculating CCCM parameters using block vectors.
[0028] FIG. 12 illustrates a spatial GPM candidates list.
[0029] FIG. 13 illustrates a template used for reordering the spatial GPM candidate list.
[0030] FIG. 14 conceptually illustrates a current block that is coded according to a reference block located in the current picture by a block vector.
[0031] FIG. 15 conceptually illustrates a current block that is coded by referencing block vectors from neighboring blocks.
[0032] FIG. 16 illustrates an example video encoder that may use block vectors to encode pixel blocks.
[0033] FIG. 17 illustrates portions of the video encoder that implement prediction by block vector referencing.
[0034] FIG. 18 conceptually illustrates a process for encoding a pixel block by referencing block vectors of previously coded blocks.
[0035] FIG. 19 illustrates an example video decoder that may use block vectors to decode pixel blocks.
[0036] FIG. 20 illustrates portions of the video decoder that implement prediction by block vector referencing.
[0037] FIG. 21 conceptually illustrates a process for decoding a pixel block by referencing block vectors of previously coded blocks.
[0038] FIG. 22 conceptually illustrates an electronic system with which some embodiments of the present disclosure are implemented.DETAILED DESCRIPTION
[0039] In the following detailed description, numerous specific details are set forth by way of examples in order to provide a thorough understanding of the relevant teachings. Any variations, derivatives and / or extensions based on teachings described herein are within the protective scope of the present disclosure. In some instances, well-known methods, procedures, components, and / or circuitry pertaining to one or more example implementations disclosed herein may be described at a relatively high level without detail, in order to avoid unnecessarily obscuring aspects of teachings of the present disclosure. I. Block Vector Based Coding Tools
[0040] A. Intra Block Copy (IBC)
[0041] Intra block copy (IBC) is a block level coding mode, in which block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, a block vector is used to indicate the displacement from the current block to a reference block, which is already reconstructed inside the current picture. An IBC-coded CU is typically considered as the third prediction mode other than intra or inter prediction modes.
[0042] B. IBC Merge / AMVP List Construction
[0043] At CU level, IBC mode is signalled with a flag and it can be signaled as IBC AMVP mode or IBC skip / merge mode. To signal as IBC skip / merge mode, a merge candidate index is used to indicate which of the block vectors in the list from neighboring candidate IBC coded blocks is used to predict the current block. The merge list consists of spatial, HMVP, and pairwise candidates. To signal as IBC AMVP mode, block vector difference is coded in the same way as a motion vector difference. The block vector prediction method uses two candidates as predictors, one from left neighbor and one from above neighbor (if IBC coded) . When either neighbor is not available, a default block vector will be used as a predictor. A flag is signaled to indicate the block vector predictor index.
[0044] In some embodiments, the IBC merge / AMVP list construction is improved according to the following: ● Only if an IBC merge / AMVP candidate is valid, it can be inserted into the IBC merge / AMVP candidate list. ● Above-right, bottom-left, and above-left spatial candidates (belonging to the adjacent spatial candidate category) and one pairwise average candidate can be added into the IBC merge / AMVP candidate list. · Template based adaptive reordering (ARMC-TM) is applied to IBC merge list. ● Candidates from non-adjacent spatial neighboring blocks (a.k.a., non-adjacent candidates) can be added to the candidate lists of IBC merge modes and IBC AMVP. These non-adjacent candidates are inserted between the adjacent spatial candidates and the HBVP candidates for both IBC merge and IBC AMVP. The same reference area of non-adjacent merge in regular inter mode is reused for the IBC. ● Restriction that adjacent spatial candidates cannot be used for IBC merge of a 4x4 CU is removed. · Auto-relocated block vector prediction (AR-BVP) candidates are added to the IBC merge and AMVP candidate list right after the HBVP candidates.
[0045] FIG. 1 shows an example derivation of auto-relocated block vector prediction (AR-BVP. ) As illustrated, a guiding block vector BV0, 1 (i.e., an existing BVP already in the candidate list) associated with the current block B0 points to a reference block B1. If B1 has a BV denoted as BV1, 2 pointing to a reference block B2, then BV0, 2, given by BV0, 2 = BV0, 1 +BV1, 2, is defined as the AR-BVP, guided by BV0, 1. When deriving BVn, n+1 guided by BV0, n, five positions of Bn are checked to find BVn, n+1.
[0046] FIG. 2 illustrates these five positions of Bn that are checked to find BVn, n+1, (with n set to 1) . As illustrated, the five positions include top-left ( “LT” ) , top-right ( “RT” ) , center ( “Ctr” ) , bottom-right ( “RB” ) , and bottom-left ( “LB” ) .
[0047] In some embodiments, the HMVP table size for IBC is increased to 25. After up to 20 IBC merge candidates are derived with full pruning, they are reordered together. After reordering, the first 6 candidates with the lowest template matching costs are selected as the final candidates in the IBC merge list. The zero vectors’ candidates used to pad the IBC Merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. A zero vector is invalid as a block vector in IBC merge mode, and consequently, it is discarded as BVP in the IBC candidate list.
[0048] FIG. 3 illustrates padding candidates for replacing the zero-vector in the IBC list. As illustrated, three candidates are located on the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C) , whose coordinates are determined by the width, and height of the current block and the ΔX and ΔY parameters.
[0049] During the IBC AMVP list construction, a clustering of the BVP candidates may be applied when both BV candidate components are non-zero. FIG. 4 shows IBC candidate clustering. The clustering is based on L2 distance and template matching cost. The clustering as shown in FIG. 4 with L2 distance is applied if there are more than 2 valid BV candidates and up to 6 candidates are clustered, the clustering radius is defined as Radius = log2 ( (cbWidth *cbHeight) ) >> MIN_PU_SIZE)
[0050] The clustering method is applied in the candidate list order, and the candidates assigned to a group are removed from the list for the subsequent clusters. In each group, the BVP with a lowest TM cost is selected as the representative candidate of that group. Finally, the representative candidates of the two first groups are chosen as the candidates for the IBC AMVP list.
[0051] Furthermore, if one of BV candidate components is zero or block is coded in RRIBC, a flag is signalled to indicate this case with a directional flag indicating horizontal or vertical component is non-zero. Instead of usual IBC AMVP list, two new BVP candidates are derived, and the sign of the non-zero BV component is derived at decoder side. The AMVP BVP0 is set to the nearest valid location to the current block (-cbWidth or -cbHeight) , so the non-zero BVD is always negative, pointing to the left for a BV with a zero vertical component or to the above for a BV with a zero horizontal component. Likewise, the AMVP BVP1 is set to the farthest position from the current block in the valid reference region, that is the left boundary or the top boundary of the IBC search region. Consequently, if the BVP1 is selected, the BVD is always positive, pointing to the right for BV with a zero vertical component or to the bottom for BV with a zero-horizontal component.
[0052] The optimal IBC AMVP index is signalled, which allows deriving the sign of the non-zero BVD component at the decoder side. The absolute magnitude of non-zero BVD component is further signalled. In RRIBC, the direction of the flipping mode is derived from the signalled directional flag.
[0053] C. Intra Template Matching (IntraTMP)
[0054] Intra template matching prediction (IntraTMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in a reconstructed part of the current frame and uses the corresponding block as a prediction block. The encoder then signals the usage of this mode, and the same prediction operation is performed at the decoder side.
[0055] The prediction signal is generated by matching the L-shaped, Top-only or Left-Only causal neighbor of the current block with another block in a predefined search area. IntraTMP employs an implicit merge mode, where merge candidates are considered without signaling a merge flag or index. Specifically, the reference positions pointed by the block vectors of all the adjacent and non-adjacent merge candidates (coded in IntraTMP or IBC mode) are used as additional candidates beyond the default search areas, up to 10 merge candidates are derived from the neighboring PUs and by prioritizing the candidates outside the IntraTMP search range. In addition, up to 20 auto-relocated block vector prediction (AR-BVP) candidates are constructed to get more reference positions. Finally, up to 20 AR-BVPs can be constructed and there are up to 30 merge candidates in total. If an AR-BVP candidate is selected for refinement, it would have a refinement range of 3x3.
[0056] The same template matching cost is used to compare the merge positions and the defaults ones. For bi-directional IBC merge candidate, two candidates are retained corresponding to each reference frame. Similarly, for IntraTMP, two candidates are considered corresponding to the best candidate by template search and the coded candidate. Sum of absolute differences (SAD) is used as a cost function for the search.
[0057] A given search order of the 6 regions is utilized for IntraTMP, i.e., R4, R5, R6, R1, R2, and R3. FIG. 5 illustrates the search regions for IntraTMP for a current block 510. Within each region, the decoder constructs a candidate list of up to 19 template matching block vectors that are ranked in ascending order according to their template costs (SAD) . The following modes are supported by IntraTMP: ● Single predictor: A single predictor is selected from the candidate list. · Fusion of multiple predictors: multiple predictors are blended multiple to derive the final prediction block. · Sub-pel precision: When single predictor is used, sub-pel precisions are supported. A new candidate list is constructed by including the selected integer block vector and surrounding 1 / 2-pel and 1-4-pel sub-pel positions. ● Linear filter model: A linear filter can be learned between the reference template and current template and be applied the linear model to reference block. This mode can be used for single predictor when sub-pel precision is not used.
[0058] In some embodiments, one or more block vectors derived from the intra template matching prediction (IntraTMP) is / are used for intra block copy (IBC) . The stored IntraTMP BV of the neighboring blocks along with IBC BV are used as spatial BV candidates in IBC candidate list construction. IntraTMP block vector is stored in the IBC block vector buffer and, the current IBC block can use both IBC BV and IntraTMP BV of neighboring blocks as BV candidate for IBC BV candidate list.
[0059] FIG. 6 illustrates using IntraTMP block vector for an IBC block. In the figure, current block 610 is coded by IBC. An earlier coded neighboring block 620 is coded by IntraTMP, which uses template 625 of the block 620 to match with a template 635 of a reference block 630. A IntraTMP block vector 605 is used to locate the reference block 630. The Intra TMP block vector 605 is also stored for use by the subsequently IBC coded current block 610. IntraTMP block vectors are added to IBC block vector candidate list as spatial candidates. IntraTMP block vectors are stored in quarter-pel resolution for coding of IBC block vectors and HMVP.
[0060] D. Direct Block Vector (DBV) for chroma block
[0061] The direct block vector is used for chroma blocks. A flag is signaled to indicate whether a chroma block is coded using IBC mode. If one of the luma blocks in five specific locations is coded with IBC or intraTMP mode, the block vector of that luma block is scaled and is used as the block vector for the chroma block. Template matching may be used to perform block vector scaling.
[0062] E. Block Vector guided EIP
[0063] Extrapolation Intra Prediction (or EIP) is an intra prediction mode that applies a linear model or filter to generate an intra-prediction of the current block, where the EIP linear model or filter is derived by using reconstructed pixels in an extended template region neighboring the current block. The extended template region extends beyond the edges of the current block by predefined lengths (or a predefined number of pixel positions. )
[0064] FIG. 7 illustrates a defined type of filter shape having 15 inputs and generating one output for EIP. FIG. 8 conceptually illustrates a diagonal order according to which an EIP filter 805 is applied to generate a predictor for the current block. As illustrated, the EIP filter 805 traverses through the samples of the current block 810 to generate the predictor. The EIP filter 805 uses reconstructed samples in an extended template region 820 neighboring the current block 810 and predicted samples within the current block 810 as inputs to the EIP filter 805.
[0065] The block-vector guided EIP (BV-EIP) method uses a block vector to determine the reference area for calculating the EIP filter parameters instead of directly using the adjacent spatial reference area and is coded as a sub-mode of EIP. Then the current coding unit utilizes the calculated EIP filter parameters and neighboring reconstructed samples to create the prediction. FIG. 9 illustrates a reference area 930 used in the block-vector guided EIP (BV-EIP) method. One flag is signaled to indicate the usage of this sub-mode of EIP. The block vector 915 used in this method is derived from the rough searching process of the intraTMP using template 912 of a current block 910. The reference area 930 includes the reference block 920 that is located by the block vector 915 and a template 925 neighboring the reference block 920.
[0066] F1. Decoder-Side Intra Mode Derivation (DIMD)
[0067] Decoder-Side Intra Mode Derivation (DIMD) is a technique in which one or more, for example, two, intra prediction modes such as angles or directions are derived from the reconstructed neighbor samples (template) of a block, and those two predictors are combined with the non-angular predictor such as planar mode predictor with the weights derived from the gradients. To implicitly derive the intra prediction modes of a block, a texture gradient analysis is performed at both encoder and decoder sides. This process starts with an empty Histogram of Gradient (HoG) having 65 entries, corresponding to the 65 angular / directional intra prediction modes. Amplitudes of these entries are determined during the texture gradient analysis.
[0068] In some embodiments, when DIMD is applied, up to five intra modes are derived from the reconstructed neighbor samples, and those five predictors are combined with the non-directional predictor (planar or block vector based predictor) with the weights derived from the histogram of gradients. The decision between for the non-directional modes is taken according to the template cost. Specifically, the block vectors of all adjacent and non-adjacent merge candidates (coded in IntraTMP or IBC) are compared to planar prediction on the reconstructed template. The template cost (SATD) is used to select the best predictor among them.
[0069] F2. Block Vector guided DIMD
[0070] In Block Vector guided DIMD (BVG-DIMD) , directional predictors and non-directional predictor are combined to form a prediction block of the current block. FIG. 10 conceptually illustrates block vector guided DIMD (BVG-DIMD) . As illustrated, reference samples in reference area 1020 pointed at by the block vectors 1015, instead of the neighboring samples of the current block 1010, are used to build an HoG 1030. The HoG 1030 is then used to derive the five intra modes that are used to derive the directional prediction 1040. Up to 5 block vectors 1015 may be obtained by the same searching process as in intraTMP. With each block vector, a reference area in the current picture is determined and its samples are used to derive the intra modes and corresponding amplitudes. To reduce complexity, overlapping reference areas are clustered. All results are counted in constructing the HoG. Two intra modes with the highest amplitudes are selected from the HoG and the prediction block 1060 is the blending of those 2 predictors and a non-directional predictor 1050. The non-directional predictor is the blending of up to 5 BV-based predictors obtained using the block vectors.
[0071] G. Block Vector guided CCCM
[0072] Convolutional cross-component model (CCCM) may be applied to predict chroma samples from reconstructed luma samples. The reconstructed luma samples are down-sampled to match the lower resolution chroma grid when chroma sub-sampling is used. Top, left or top and left reference samples are used as templates for model derivation.
[0073] When the co-located luma prediction is coded with IBC or IntraTMP in Intra slices, the block vector guided CCCM (BVG-CCCM) mode can be used. In this mode, the block vectors of the co-located luma blocks, coded in IBC or intraTMP modes, are used to determine the reference area for calculating the CCCM parameters. FIG. 11 illustrates the reference area for calculating CCCM parameters in BVG-CCCM method. As illustrated, a current chroma block 1110 has a block vector 1115 locating a reference area 1118 for the chroma component. The current chroma block 1110 has a collocated luma block 1120, which has a block vector 1125 locating a reference area 1128 for the luma component. The chroma channel block vector 1115 may be scaled or down-sampled from the luma channel block vector 1125.
[0074] The model parameters can be calculated based on the samples of the reference luma area 1128 and the samples of the chroma reference area 1118. The prediction for the current chroma block 1110 may be performed using the calculated model parameters and co-located luma samples in block 1120. The BVG-CCCM mode uses an 11-tap filter for cross-component prediction as below: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P (C) + c6P (N) + c7P (S) + c8P (W) + c9P (E) + c10B
[0075] The input to the spatial 5-tap component of the filter consists of a center (C) luma sample which is collocated with the chroma sample to be predicted and its above / north (N) , below / south (S) , left / west (W) and right / east (E) neighbors as illustrated CCCM. The nonlinear term P is represented as power of two of the corresponding luma sample and B is the bias term.
[0076] Similar to Direct Block Vector (DBV) mode, five locations in the collocated luma block area are scanned and the associated block vectors are then used for determining the reference area for parameter calculation in BVG-CCCM method.
[0077] H. Spatial Geometric Partitioning Mode (SGPM) with Block Vectors
[0078] Geometric Partitioning Mode (GPM) is a prediction mode in which a CU is split into at least two parts by a geometrically located straight line. The location of the splitting line is mathematically derived from the angle and offset parameters of a specific partition. Each GPM partitioning or GPM split is a partition mode characterized by a distance-angle pairing that defines a segmenting line.
[0079] Spatial GPM (SGPM) is an intra mode that resembles the inter coding tool of GPM, where the two prediction parts are generated from intra predicted process. In this mode, a candidate list is built with each entry comprising a combination of one partition split and two intra prediction modes. FIG. 12 illustrates a spatial GPM candidates list. 26 partition modes and 9 of intra prediction modes are used to form the combinations as the entries of the SGPM candidate list. The length of the candidate list is set equal to 16. The selected candidate index is signalled. The SGPM candidate list may be reordered using a template neighboring the current block, where SAD between the prediction and reconstruction of the template is used for ordering. FIG. 13 illustrates a template used for reordering the spatial GPM candidate list.
[0080] In some embodiments, for each partition mode, a list of intra prediction modes (IPM list) is derived for each partition using an intra-inter GPM list derivation. The IPM list size is set to 3. The list may include intra prediction modes that are derived by template matching, or derived modes with horizontal and vertical orientations. The list is further augmented with block-vector based prediction candidates obtained from the adjacent and non-adjacent merge candidates coded in IntraTMP or IBC mode. The template cost is employed to select the up to 6 block vectors. The final IPM list may contain up to 9 predictors that include 3 regular intra modes and up to 6 BV-based predictors. II. Referencing Block Vectors
[0081] Many block-vector (BV) based coding tools are adopted in video coding standards to improve coding efficiency. Two ways of utilizing block vectors in coding tools are described as following:
[0082] (1) The current block is coded utilizing one or more block vectors of the current block. FIG. 14 conceptually illustrates a current block 1410 that is coded according to a reference block 1430 located in the current picture 1400 by a block vector 1415 of the current block. For example, in some embodiments, the prediction of the current block (e.g., block 1410) may be generated based on coding information derived based on samples in reference blocks (e.g., block 1430) . The coding information may be an EIP model or intra modes. For another example, in some embodiments, the prediction of the current block is generated based on sample values in reference blocks.
[0083] (2) The current block is coded by referencing block vectors from neighboring blocks. In some embodiments, the block vector of the current block is selected / derived from block vectors from neighboring blocks. FIG. 15 conceptually illustrates a current block 1510 that is coded by referencing block vectors 1521 and 1522 from neighboring blocks 1511 and 1512. For example, in some embodiments, the block vector used to locate the reference block for calculating HoG is obtained through an IntraTMP search process, which includes building an IntraTMP candidate list including block vectors from neighboring blocks.
[0084] In some embodiments, when the current block is coded in a mode utilizing the one or more block vectors of the current block (also referred to as “first target block vectors” ) and when a subsequent block is coded by referencing one or more block vectors from neighboring blocks (also referred to as “second target block vectors” ) , and if the current block is referenced by the subsequent block (i.e., the current block is a neighboring block of the subsequent block) , the first target block vectors can be referenced by the subsequent block as the second target block vectors. Thus, in the example of FIG. 15, the block 1511 can be the current block and the block 1510 can be the subsequent block, and that the block 1511 is the neighboring block of the block 1510. The block vector 1521 can be the “first target block vector” that is referenced by the subsequent block 1510 as the “second target block vector” .
[0085] Note that the method of utilizing block vectors of the current block and the method of utilizing / referencing block vectors from neighboring blocks are not used exclusively in a coding tool. For example, in DIMD, the current block may be coded based on a BV-based predictor generated based on a block vector of the current block. And the block vector is selected from block vectors candidates from neighboring blocks. Hence DIMD uses both ways of utilizing block vectors.
[0086] In some embodiments, one or more first target block vectors are derived and used to code a block. In some embodiments, one or more second target block vectors are referenced to code a block. In some embodiments, for IBC merge mode and IntraTMP candidate list, when referencing block vectors from neighboring blocks, in addition to block vectors from IBC or IntraTMP coded blocks, block vectors from blocks coded in modes utilizing block vectors can also be referenced.
[0087] A. Deriving first target block vectors of the current block
[0088] In some embodiments, when coding the current block, one or more block vectors are derived or selected for the current block, and the current block is coded according to reference blocks located by the one or more block vectors of the current block. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks.
[0089] In some embodiments, when coding the current block, one or more block vectors are derived / selected for the current block. Coding information is derived based on samples in reference blocks located by one or more block vectors of the current block. The prediction of the current block is then generated based on the coding information. The coding information may be filter-based models, parameter models, cross-component models, or intra prediction modes. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks. The following are examples of the coding information derived based on samples in reference blocks located by one or more block vectors of the current block:
[0090] For a first example, in some embodiments, the coding information is EIP models derived based on samples in the reference blocks. The method of deriving parameters of EIP models is further described in Section I. E. One or more block vectors are derived / selected for the current block to locate the reference blocks, and the one or more block vectors are the first target block vectors and can be referenced by subsequent blocks.
[0091] For a second example, in some embodiments, the coding information is intra prediction modes derived based on samples in the reference blocks. The method of deriving intra modes can be a DIMD process computing HoG based on the samples in the reference block as described in Section I. F1 or Section I. F2. One or more block vectors are derived / selected for the current block to locate the reference blocks, and the one or more block vectors are the first target block vectors and can be referenced by subsequent blocks.
[0092] For a third example, in some embodiments, the is cross-component models derived based on samples in the reference blocks. The cross-component model can be CCLM, GLM, multi-model cross-component model, CCCM and CCCM variants (such as with different filter shape, non-downsampled inputs, multiple downsampling filters, using location or gradient term as inputs) . If CCCM is to be derived, the method of deriving parameters of CCCM can be as described in Section I.G. One or more block vectors are derived / selected for the current block to locate the reference blocks, and the one or more block vectors are the first target block vectors and can be referenced by subsequent blocks.
[0093] In some embodiments, when coding the current block, one or more block vectors are derived / selected for the current block. BV-based predictors are generated based on sample values of the reference blocks located by the one or more block vectors of the current block, and the final predictor of the current block is a combination of BV-based predictors and other BV-based / non-BV-based predictors. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks. The following are examples of final predictor generated based on BV-based predictors:
[0094] For a first example, in some embodiments, the final predictor can be a combination of intra directional mode predictors and BV-based predictor. The BV-based predictors are generated based on the sample values of reference blocks. One or more block vectors are derived / selected for the current block to locate the reference blocks, and the one or more block vectors are the first target block vectors and can be referenced by subsequent blocks. Take DIMD as example, as described in Section I. F1 or Section I. F2, the final predictor is a combination of intra directional mode predictors and BV-based predictor, if a block vector candidate is selected as the non-directional mode.
[0095] For a second example, in some embodiments, the current block can be partitioned geometrically, and at least the final predictor of one of the partitions is a BV-based predictor. One or more block vectors are derived / selected for the partition with the BV-based predictor, and the one or more block vectors are the first target block vectors and can be referenced by subsequent blocks. Take SGPM as example, as described in Section I. H, the IPM list of each partition mode comprises block-vector candidates. The final predictor of a partition can be a BV-based predictor if a block vector candidate is selected from the IPM list.
[0096] In some embodiment, when coding the current block, one or more block vectors are derived / selected for the current block. Prediction of the current block is generated based on the sample values in reference blocks located by one or more block vectors of the current block. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks. For example, the current block may be a chroma block, and one or more block vectors are derived / selected for the current chroma block. The prediction of the current chroma block is generated based on the sample values of reference blocks located by the one or more block vector of the current chroma block. Method of generating the current prediction are further described in Section I. D. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks.
[0097] In some embodiments, the method of deriving / selecting one or more block vectors can be the same as IBC or IntraTMP, which are further described in Section I. A and Section I. C respectively. The one or more block vectors derived / selected for the current block can be the first target block vectors and can be referenced by subsequent blocks.
[0098] B. Referencing Second Block Vectors from Neighboring Block
[0099] In some embodiments, when coding the current block, one or more block vectors are selected from one or more second target block vectors. The second target block vectors may be from a merge candidate list, and the one or more block vectors are selected from the merge candidate list. The second target block vectors may be obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The current block is coded based on the one or more block vectors selected from the second target block vectors.
[0100] In some embodiments, when coding the current block, one or more block vectors are selected from one or more second target block vectors. Coding information is derived based on samples in reference blocks located by the one or more block vectors. The prediction of the current block is then generated based on the coding information. The coding information may be filter-based models, parameter models, cross-component models, or intra prediction modes. One or more second target block vectors are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The following are examples of coding information derived based on samples in reference blocks located by the one or more block vectors.
[0101] For a first example, the coding information is EIP models derived based on samples in the reference blocks. The method of deriving parameters of EIP models can be as described in Section I. E. The reference blocks are located by the one or more block vectors selected from one or more second target block vectors, which are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The second target block vectors can be obtained from spatial, temporal or history neighbor blocks. The neighboring blocks can be the same as the IBC merge mode or as the IntraTMP search process.
[0102] For a second example, in some embodiments, the coding information is intra prediction modes derived based on samples in the reference block. The method of deriving intra modes can be a DIMD process computing HoG based on the samples in the reference block as described in Section I. F1 or Section I. F2. The reference blocks are located by the one or more block vectors selected from one or more second target block vectors, which are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The second target block vectors can be obtained from spatial, temporal or history neighbor blocks. The neighboring blocks can be the same as the IBC merge mode or as the IntraTMP search process.
[0103] For a third example, in some embodiments, the coding information is cross-component models derived based on samples in the reference blocks. The cross-component model can be CCLM, GLM, multi-model cross-component model, CCCM and CCCM variants (such as with different filter shape, non-downsampled inputs, multiple downsampling filters, using location or gradient term as inputs) . If CCCM is to be derived, the method of deriving parameters of CCCM can be as described in Section I. G. The reference blocks are located by the one or more block vectors selected from one or more second target block vectors, which are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The second target block vectors can be obtained from spatial, temporal or history neighbor blocks. The neighboring blocks can be the same as the IBC merge mode or as the IntraTMP search process.
[0104] In some embodiments, when coding the current block, one or more block vectors are selected from one or more second target block vectors. BV-based predictors are generated based on sample values of the reference blocks located by the one or more block vectors, and the final predictor of the current block is a combination of BV-based predictors and other BV-based / non-BV-based predictors. One or more second target block vectors are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The following are examples of final predictor generated based BV-based predictors:
[0105] For a first example, in some embodiments, the final predictor can be a combination of intra directional mode predictors and BV-based predictor. The BV-based predictors are generated based on the sample values of reference blocks. The reference blocks are located by the one or more block vectors selected from one or more second target block vectors, which are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The second target block vectors can be obtained from spatial, temporal or history neighbor blocks. The neighboring blocks can be the same as the IBC merge mode or as the IntraTMP search process. Take DIMD as example, as described in Section I. F1 or Section I. F2, the final predictor is a combination of intra directional mode predictors and BV-based predictor, if a block vector candidate is selected as the non-directional mode. The second target block vectors are obtained through a IntraTMP search process.
[0106] For a second example, in some embodiments, the current block can be partitioned geometrically, and at least the final predictor of one of the partitions is a BV-based predictor. The BV-based predictors are generated based on the sample values of reference blocks. The reference blocks are located by the one or more block vectors selected from one or more second target block vectors, which are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The second target block vectors can be obtained from spatial, temporal or history neighbor blocks. The neighboring blocks can be the same as the IBC merge mode or as the IntraTMP search process. Take SGPM as example, as described in Section I. H, the IPM list of each partition mode comprises block-vector candidates. The final predictor of a partition can be a BV-based predictor if a block vector candidate is selected from the IPM list. The IPM list is constructed by including block-vector based prediction candidates obtained from the adjacent and non-adjacent merge candidates coded in IntraTMP or IBC mode.
[0107] In some embodiments, when coding the current block, one or more block vectors are selected from one or more second target block vectors. Prediction of the current block is generated based on the sample values in reference blocks located by the one or more block vectors. One or more second target block vectors are obtained from referencing previously coded blocks that are neighboring blocks of the current block and have first target block vectors. When a previously coded neighboring block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. For example, the current block may be a chroma block, and one or more block vectors are selected from one or more second target block vectors. The prediction of the current chroma block is generated based on the sample values of reference blocks located by the one or more block vector. The second target block vectors are obtained by referencing collocated luma blocks of the current chroma block. If the collocated luma block has first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors. The method of obtaining the second target block vector can be as described in Section I. D.
[0108] In some embodiments, the method of obtaining the second target block vectors are further described in Section I. A and Section I. C. When referencing spatial adjacent, spatial non-adjacent, temporal and historical neighboring blocks that have first target block vectors, the first target block vectors are retrieved to be included in the set of the second target block vectors.
[0109] The term “block” in this invention can refer to TU / TB, CU / CB, PU / PB, pre-defined region, or CTU / CTB. Any combination of the proposed methods in this invention can be applied. Any of the foregoing proposed methods can be implemented in encoders and / or decoders. For example, any of the proposed methods can be implemented in an inter / intra / IBC / prediction / transform module of an encoder, and / or an inter / intra / IBC / prediction / transform module of a decoder. Alternatively, any of the proposed methods can be implemented as a circuit coupled to the inter / intra / IBC / prediction / transform module of the encoder and / or the inter / intra / IBC / prediction / transform module of the decoder, so to provide the information needed by the inter / intra / IBC / prediction / transform module. III. Example Video Encoder
[0110] FIG. 16 illustrates an example video encoder 1600 that may use block vectors to encode pixel blocks. As illustrated, the video encoder 1600 receives input video signal from a video source 1605 and encodes the signal into bitstream 1695. The video encoder 1600 has several components or modules for encoding the signal from the video source 1605, at least including some components selected from a transform module 1610, a quantization module 1611, an inverse quantization module 1614, an inverse transform module 1615, an intra estimation module 1624, an intra prediction module 1625, a motion compensation module 1630, a motion estimation module 1635, an in-loop filter 1645, a reconstructed picture buffer 1650, a MV buffer 1665, and a MV prediction module 1675, and an entropy encoder 1690. The motion compensation module 1630 and the motion estimation module 1635 are part of an inter-prediction module 1640. The intra-prediction module 1625 and the intra-estimation module 1624 are part of a current picture prediction module 1620, which uses current picture reconstructed samples as reference samples for prediction of the current block.
[0111] In some embodiments, the modules 1610 –1690 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device or electronic apparatus. In some embodiments, the modules 1610 –1690 are modules of hardware circuits implemented by one or more integrated circuits (ICs) of an electronic apparatus. Though the modules 1610 –1690 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0112] The video source 1605 provides a raw video signal that presents pixel data of each video frame without compression. A subtractor 1608 computes the difference between the raw video pixel data of the video source 1605 and the predicted pixel data 1613 from the motion compensation module 1630 or intra-prediction module 1625 as prediction residual 1609. The transform module 1610 converts the difference (or the residual pixel data or residual signal 1609) into transform coefficients (e.g., by performing Discrete Cosine Transform, or DCT) . The quantization module 1611 quantizes the transform coefficients into quantized data (or quantized coefficients) 1612, which is encoded into the bitstream 1695 by the entropy encoder 1690.
[0113] The inverse quantization module 1614 de-quantizes the quantized data (or quantized coefficients) 1612 to obtain transform coefficients 1618, and the inverse transform module 1615 performs inverse transform on the transform coefficients 1618 to produce reconstructed residual 1619. The reconstructed residual 1619 is added with the predicted pixel data 1613 to produce reconstructed pixel data 1617. In some embodiments, the reconstructed pixel data 1617 is temporarily stored in a line buffer 1627 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction. The reconstructed pixels are filtered by the in-loop filter 1645 and stored in the reconstructed picture buffer 1650. In some embodiments, the reconstructed picture buffer 1650 is a storage external to the video encoder 1600. In some embodiments, the reconstructed picture buffer 1650 is a storage internal to the video encoder 1600.
[0114] The intra estimation module 1624 derives intra-prediction data (e.g., intra prediction modes) based on the reconstructed pixel data 1617 (stored in the line buffer 1627) . The intra-prediction data is provided to the entropy encoder 1690 to be encoded into bitstream 1695. The intra-prediction data is also used by the intra-prediction module 1625 to produce the predicted pixel data 1613.
[0115] The motion estimation module 1635 performs inter-prediction by producing MVs to reference pixel data of previously decoded frames stored in the reconstructed picture buffer 1650. These MVs are provided to the motion compensation module 1630 to produce predicted pixel data.
[0116] Instead of encoding the complete actual MVs in the bitstream, the video encoder 1600 uses MV prediction to generate predicted MVs, and the difference between the MVs used for motion compensation and the predicted MVs is encoded as residual motion data and stored in the bitstream 1695.
[0117] The MV prediction module 1675 generates the predicted MVs based on reference MVs that were generated for encoding previously video frames, i.e., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 1675 retrieves reference MVs from previous video frames from the MV buffer 1665. The video encoder 1600 stores the MVs generated for the current video frame in the MV buffer 1665 as reference MVs for generating predicted MVs.
[0118] The MV prediction module 1675 uses the reference MVs to create the predicted MVs. The predicted MVs can be computed by spatial MV prediction or temporal MV prediction. The difference between the predicted MVs and the motion compensation MVs (MC MVs) of the current frame (residual motion data) are encoded into the bitstream 1695 by the entropy encoder 1690.
[0119] The entropy encoder 1690 encodes various parameters and data into the bitstream 1695 by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding. The entropy encoder 1690 encodes various header elements, flags, along with the quantized transform coefficients 1612, and the residual motion data as syntax elements into the bitstream 1695. The bitstream 1695 is in turn stored in a storage device or transmitted to a decoder over a communications medium such as a network.
[0120] The in-loop filter 1645 performs filtering or smoothing operations on the reconstructed pixel data 1617 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1645 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.
[0121] FIG. 17 illustrates portions of the video encoder 1600 that implement prediction by block vector referencing. Specifically, the figure illustrates a BV-based predictor generator 1710 that uses block vectors of the current block (first target block vectors 1712) and block vectors that are stored for previously coded blocks (second target block vectors 1714) to generate a BV-based predictor 1718 for the current block. The block vector (s) 1716 used for generating the predictor of the current block are then stored in a BV storage 1735 so can be referenced or used by a subsequent block in a merge list. The BV-based predictor 1718 is used to generate a final predictor 1755, which is used as the predicted pixel data 1613.
[0122] The first target block vector (s) 1712 used for generating the predictor 1718 may be provided by an IBC module 1720, which may relay the block vector to the entropy encoder 1690 to be signaled in the bitstream 1695. The first target block vector (s) 1712 may also be provided by an IntraTMP module 1725, which searches the current picture for reference blocks or areas having templates matching that of the current block and provide block vectors that locate those reference areas.
[0123] The second target block vector (s) 1714 may be provided by a merge candidate select module 1730, which selects block vectors from a merge list that includes block vectors that are stored for previously coded blocks or neighboring blocks in the BV storage 1735.
[0124] The BV-based predictor generator 1710 may use the first target block vectors 1712 and / or the second target block vectors 1714 to access the reconstructed picture buffer 1650 or the line buffer 1627 to retrieve samples of the reference areas or reference blocks located by the block vectors. The BV-based prediction generator 1710 may use the retrieved samples (located by the block vectors) as the BV-based predictor 1718 of the current block. The BV-based prediction generator 1710 may also use the retrieved samples to generate a set of coding information 1715, which are then used to generate the BV-based predictor 1718.
[0125] The set of coding information 1715 may include a filter-based model, or a parameter model, or a cross-component model, or a set of intra prediction modes. The filter-based model may be an EIP model described in Section I. A. The parameter model or the cross-component model may be a CCCM model described in Section I. G. The set intra prediction modes may be directional modes identified by using a HoG of texture gradient analysis as described in Section I. F1 or Section I. F2. The retrieved samples of the reference area located by the block vector may also be considered as coding information 1715 for some embodiments. Depending on the type of coding information that is generated and used, the BV-based predictor 1718 may be considered a non-directional predictor (based on block vectors but not any directional intra modes. )
[0126] In some embodiments, the video encoder 1600 generates the final predictor 1755 by using a prediction combiner 1750 to combine the BV-based predictor 1718 with another predictor 1742. The other predictor 1742 may be another BV-based predictor, or a directional intra mode predictor, or a non-BV-based predictor. For example, in some embodiments, a DIMD process 1740 may be used to identify a set of directional intra prediction modes, and the intra prediction module 1625 may generate a directional intra predictor as the other predictor 1742 to be combined with the BV-based predictor 1718. Such a directional intra predictor 1742 is a non-BV-based predictor being combined with the BV-based predictor 1718.
[0127] FIG. 18 conceptually illustrates a process 1800 for encoding a pixel block by referencing block vectors of previously coded blocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the encoder 1600 performs the process 1800 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the encoder 1600 performs the process 1800.
[0128] The encoder receives (at block 1810) data to be encoded as a current block of pixels in a current picture.
[0129] The encoder derives (at block 1820) or selects one or more first target block vectors for the current block. The first target block vectors may be used for IBC prediction of the current block. The first target block vectors may be derived by IntraTMP, i.e., searching the current picture for template areas matching that of the current block. The encoder may select the one or more first target block vectors from an intra prediction mode (IPM) list of a partition mode of spatial geometric partition mode (SGPM) , the IPM list augmented with BV-based prediction candidates obtained from adjacent and non-adjacent merge candidates coded in IBC or IntraTMP.
[0130] The encoder locates (at block 1830) one or more reference areas in the current picture by using the first target block vectors. The encoder generates (at block 1835) a set of coding information based on samples of the located one or more reference areas. The coding information may include a filter-based model, a parameter model, a cross-component model, or an intra prediction mode. The coding information may include one or more BV-based predictors that are generated based on sample values of the reference areas located by the first target block vectors.
[0131] The encoder generates (at block 1840) a final predictor of the current block based on the set of coding information. The final predictor may be a combination of the one or more BV-based predictors that are generated based on the first target block vectors and one or more other predictors. The other predictors may be BV-based predictors or non-BV-based predictors. The one or more other predictors may be generated based on one or more intra direction modes that are determined from a histogram of gradient (HoG) based on a texture gradient analysis (e.g., DIMD process) of samples in a reference area located by the first target block vectors. The encoder encodes (at block 1850) the current block by using the final predictor to produce prediction residuals.
[0132] The encoder encodes (at block 1860) a subsequent block by referencing the first target block vectors. The subsequent block may be encoded by using a merge candidate list comprising one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the second target block vectors including the first target block vectors that are stored for the current block. The subsequent block may be encoded based on one or more block vectors selected from one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the neighboring blocks of the subsequent block including the current block and the one or more second target block vectors including the first target block vectors stored for the current block.
[0133] In some embodiments, BV-based predictors may be generated for the subsequent block based on sample values of the reference areas located by the one or more block vectors selected from the second target block vectors, and a final predictor of the subsequent block is a combination of the BV-based predictors and one or more other (BV-based or non-BV-based) predictors. The one or more other predictors may be generated based on one or more intra prediction modes that are determined from a HoG based on a texture gradient analysis of samples in a reference area located by the selected one or more block vectors. IV. Example Video Decoder
[0134] In some embodiments, an encoder may signal (or generate) one or more syntax element in a bitstream, such that a decoder may parse said one or more syntax element from the bitstream.
[0135] FIG. 19 illustrates an example video decoder 1900 that may use block vectors to decode pixel blocks. As illustrated, the video decoder 1900 is an image-decoding or video-decoding circuit that receives a bitstream 1995 and decodes the content of the bitstream into pixel data of video frames for display. The video decoder 1900 has several components or modules for decoding the bitstream 1995, including some components selected from an inverse quantization module 1914, an inverse transform module 1915, an intra-prediction module 1925, a motion compensation module 1930, an in-loop filter 1945, a decoded picture buffer 1950, a MV buffer 1965, a MV prediction module 1975, and a parser 1990. The motion compensation module 1930 is part of an inter-prediction module 1940. The intra-prediction module 1925 is part of a current picture prediction module 1920, which uses current picture reconstructed samples as reference samples for prediction of the current block.
[0136] In some embodiments, the modules 1914 –1990 are modules of software instructions being executed by one or more processing units (e.g., a processor) of a computing device. In some embodiments, the modules 1914 –1990 are modules of hardware circuits implemented by one or more ICs of an electronic apparatus. Though the modules 1914 –1990 are illustrated as being separate modules, some of the modules can be combined into a single module.
[0137] The parser 1990 (or entropy decoder) receives the bitstream 1995 and performs initial parsing according to the syntax defined by a video-coding or image-coding standard. The parsed syntax element includes various header elements, flags, as well as quantized data (or quantized coefficients) 1912. The parser 1990 parses out the various syntax elements by using entropy-coding techniques such as context-adaptive binary arithmetic coding (CABAC) or Huffman encoding.
[0138] The inverse quantization module 1914 de-quantizes the quantized data (or quantized coefficients) 1912 to obtain transform coefficients, and the inverse transform module 1915 performs inverse transform on the transform coefficients 1918 to produce reconstructed residual signal 1919. The reconstructed residual signal 1919 is added with predicted pixel data 1913 from the intra-prediction module 1925 or the motion compensation module 1930 to produce decoded pixel data 1917. The decoded pixels data are filtered by the in-loop filter 1945 and stored in the decoded picture buffer 1950. In some embodiments, the decoded picture buffer 1950 is a storage external to the video decoder 1900. In some embodiments, the decoded picture buffer 1950 is a storage internal to the video decoder 1900.
[0139] The intra-prediction module 1925 receives intra-prediction data from bitstream 1995 and according to which, produces the predicted pixel data 1913 from the decoded pixel data 1917 stored in the decoded picture buffer 1950. In some embodiments, the decoded pixel data 1917 is also stored in a line buffer 1927 (or intra prediction buffer) for intra-picture prediction and spatial MV prediction.
[0140] In some embodiments, the content of the decoded picture buffer 1950 is used for display. A display device 1905 either retrieves the content of the decoded picture buffer 1950 for display directly, or retrieves the content of the decoded picture buffer to a display buffer. In some embodiments, the display device receives pixel values from the decoded picture buffer 1950 through a pixel transport.
[0141] The motion compensation module 1930 produces predicted pixel data 1913 from the decoded pixel data 1917 stored in the decoded picture buffer 1950 according to motion compensation MVs (MC MVs) . These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1995 with predicted MVs received from the MV prediction module 1975.
[0142] The MV prediction module 1975 generates the predicted MVs based on reference MVs that were generated for decoding previous video frames, e.g., the motion compensation MVs that were used to perform motion compensation. The MV prediction module 1975 retrieves the reference MVs of previous video frames from the MV buffer 1965. The video decoder 1900 stores the motion compensation MVs generated for decoding the current video frame in the MV buffer 1965 as reference MVs for producing predicted MVs.
[0143] The in-loop filter 1945 performs filtering or smoothing operations on the decoded pixel data 1917 to reduce the artifacts of coding, particularly at boundaries of pixel blocks. In some embodiments, the filtering or smoothing operations performed by the in-loop filter 1945 include deblock filter (DBF) , sample adaptive offset (SAO) , and / or adaptive loop filter (ALF) . In some embodiments, luma mapping chroma scaling (LMCS) is performed before the loop filters.
[0144] FIG. 20 illustrates portions of the video decoder 1900 that implement prediction by block vector referencing. Specifically, the figure illustrates a BV-based predictor generator 2010 that uses block vectors of the current block (first target block vectors 2012) and block vectors that are stored for previously coded blocks (second target block vectors 2014) to generate a BV-based predictor 2018 for the current block. The block vector (s) 2016 used for generating the predictor of the current block are then stored in a BV storage 2035 so can be referenced or used by a subsequent block in a merge list. The BV-based predictor 2018 is used to generate a final predictor 2055, which is used as the predicted pixel data 1913.
[0145] The first target block vector (s) 2012 used for generating the predictor 2018 may be provided by an IBC module 2020, which may receive the block vector from the entropy decoder 1990. The first target block vector (s) 2012 may also be provided by an IntraTMP module 2025, which searches the current picture for reference blocks or areas having templates matching that of the current block and provide block vectors that locate those reference areas.
[0146] The second target block vector (s) 2014 may be provided by a merge candidate select module 2030, which selects block vectors from a merge list that includes block vectors that are stored for previously coded blocks or neighboring blocks in the BV storage 2035.
[0147] The BV-based predictor generator 2010 may use the first target block vectors 2012 and / or the second target block vectors 2014 to access the decoded picture buffer 1950 or the line buffer 1927 to retrieve samples of the reference areas or reference blocks located by the block vectors. The BV-based prediction generator 2010 may use the retrieved samples (located by the block vectors) as the BV-based predictor 2018 of the current block. The BV-based prediction generator 2010 may also use the retrieved samples to generate a set of coding information 2015, which are then used to generate the BV-based predictor 2018.
[0148] The set of coding information 2015 may include a filter-based model, or a parameter model, or a cross-component model, or a set of intra prediction modes. The filter-based model may be an EIP model described in Section I. A. The parameter model or the cross-component model may be a CCCM model described in Section I. G. The set intra prediction modes may be directional modes identified by using a HoG of texture gradient analysis as described in Section I. F1 or Section I. F2. The retrieved samples of the reference area located by the block vector may also be considered as coding information 2015 for some embodiments. Depending on the type of coding information that is generated and used, the BV-based predictor 2018 may be considered a non-directional predictor (based on block vectors but not any directional intra modes. )
[0149] In some embodiments, the video decoder 1900 generates the final predictor 2055 by using a prediction combiner 2050 to combine the BV-based predictor 2018 with another predictor 2042. The other predictor 2042 may be another BV-based predictor, or a directional intra mode predictor, or a non-BV-based predictor. For example, in some embodiments, a DIMD process 2040 may be used to identify a set of directional intra prediction modes, and the intra prediction module 1925 may generate a directional intra predictor as the other predictor 2042 to be combined with the BV-based predictor 2018. Such a directional intra predictor 2042 is a non-BV-based predictor being combined with the BV-based predictor 2018.
[0150] FIG. 21 conceptually illustrates a process 2100 for decoding a pixel block by referencing block vectors of previously coded blocks. In some embodiments, one or more processing units (e.g., a processor) of a computing device implementing the decoder 1900 performs the process 2100 by executing instructions stored in a computer readable medium. In some embodiments, an electronic apparatus implementing the decoder 1900 performs the process 2100.
[0151] The decoder receives (at block 2110) data to be decoded as a current block of pixels in a current picture.
[0152] The decoder derives (at block 2120) or selects one or more first target block vectors for the current block. The first target block vectors may be used for IBC prediction of the current block. The first target block vectors may be derived by IntraTMP, i.e., searching the current picture for template areas matching that of the current block. The decoder may select the one or more first target block vectors from an intra prediction mode (IPM) list of a partition mode of spatial geometric partition mode (SGPM) , the IPM list augmented with BV-based prediction candidates obtained from adjacent and non-adjacent merge candidates coded in IBC or IntraTMP.
[0153] The decoder locates (at block 2130) one or more reference areas in the current picture by using the first target block vectors. The decoder generates (at block 2135) a set of coding information based on samples of the located one or more reference areas. The coding information may include a filter-based model, a parameter model, a cross-component model, or an intra prediction mode. The coding information may include one or more BV-based predictors that are generated based on sample values of the reference areas located by the first target block vectors.
[0154] The decoder generates (at block 2140) a final predictor of the current block based on the set of coding information. The final predictor may be a combination of the one or more BV-based predictors that are generated based on the first target block vectors and one or more other predictors. The other predictors may be BV-based predictors or non-BV-based predictors. The one or more other predictors may be generated based on one or more intra direction modes that are determined from a histogram of gradient (HoG) based on a texture gradient analysis (e.g., DIMD process) of samples in a reference area located by the first target block vectors. The decoder reconstructs (at block 2150) the current block by using the final predictor. The decoder may then provide the reconstructed current block for display or output as part of the reconstructed current picture.
[0155] The decoder decodes (at block 2160) a subsequent block by referencing the first target block vectors. The subsequent block may be decoded by using a merge candidate list comprising one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the second target block vectors including the first target block vectors that are stored for the current block. The subsequent block may be decoded based on one or more block vectors selected from one or more second target block vectors that are stored for neighboring blocks of the subsequent block, the neighboring blocks of the subsequent block including the current block and the one or more second target block vectors including the first target block vectors stored for the current block.
[0156] In some embodiments, BV-based predictors may be generated for the subsequent block based on sample values of the reference areas located by the one or more block vectors selected from the second target block vectors, and a final predictor of the subsequent block is a combination of the BV-based predictors and one or more other (BV-based or non-BV-based) predictors. The one or more other predictors may be generated based on one or more intra prediction modes that are determined from a HoG based on a texture gradient analysis of samples in a reference area located by the selected one or more block vectors. V. Example Electronic System
[0157] Many of the above-described features and applications are implemented as software processes that are specified as a set of instructions recorded on a computer readable storage medium (also referred to as computer readable medium) . When these instructions are executed by one or more computational or processing unit (s) (e.g., one or more processors, cores of processors, or other processing units) , they cause the processing unit (s) to perform the actions indicated in the instructions. Examples of computer readable media include, but are not limited to, CD-ROMs, flash drives, random-access memory (RAM) chips, hard drives, erasable programmable read only memories (EPROMs) , electrically erasable programmable read-only memories (EEPROMs) , etc. The computer readable media does not include carrier waves and electronic signals passing wirelessly or over wired connections.
[0158] In this specification, the term “software” is meant to include firmware residing in read-only memory or applications stored in magnetic storage which can be read into memory for processing by a processor. Also, in some embodiments, multiple software inventions can be implemented as sub-parts of a larger program while remaining distinct software inventions. In some embodiments, multiple software inventions can also be implemented as separate programs. Finally, any combination of separate programs that together implement a software invention described here is within the scope of the present disclosure. In some embodiments, the software programs, when installed to operate on one or more electronic systems, define one or more specific machine implementations that execute and perform the operations of the software programs.
[0159] FIG. 22 conceptually illustrates an electronic system 2200 with which some embodiments of the present disclosure are implemented. The electronic system 2200 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc. ) , phone, PDA, or any other sort of electronic device. Such an electronic system includes various types of computer readable media and interfaces for various other types of computer readable media. Electronic system 2200 includes a bus 2205, processing unit (s) 2210, a graphics-processing unit (GPU) 2215, a system memory 2220, a network 2225, a read-only memory 2230, a permanent storage device 2235, input devices 2240, and output devices 2245.
[0160] The bus 2205 collectively represents all system, peripheral, and chipset buses that communicatively connect the numerous internal devices of the electronic system 2200. For instance, the bus 2205 communicatively connects the processing unit (s) 2210 with the GPU 2215, the read-only memory 2230, the system memory 2220, and the permanent storage device 2235.
[0161] From these various memory units, the processing unit (s) 2210 retrieves instructions to execute and data to process in order to execute the processes of the present disclosure. The processing unit (s) may be a single processor or a multi-core processor in different embodiments. Some instructions are passed to and executed by the GPU 2215. The GPU 2215 can offload various computations or complement the image processing provided by the processing unit (s) 2210.
[0162] The read-only-memory (ROM) 2230 stores static data and instructions that are used by the processing unit (s) 2210 and other modules of the electronic system. The permanent storage device 2235, on the other hand, is a read-and-write memory device. This device is a non-volatile memory unit that stores instructions and data even when the electronic system 2200 is off. Some embodiments of the present disclosure use a mass-storage device (such as a magnetic or optical disk and its corresponding disk drive) as the permanent storage device 2235.
[0163] Other embodiments use a removable storage device (such as a floppy disk, flash memory device, etc., and its corresponding disk drive) as the permanent storage device. Like the permanent storage device 2235, the system memory 2220 is a read-and-write memory device. However, unlike storage device 2235, the system memory 2220 is a volatile read-and-write memory, such a random access memory. The system memory 2220 stores some of the instructions and data that the processor uses at runtime. In some embodiments, processes in accordance with the present disclosure are stored in the system memory 2220, the permanent storage device 2235, and / or the read-only memory 2230. For example, the various memory units include instructions for processing multimedia clips in accordance with some embodiments. From these various memory units, the processing unit (s) 2210 retrieves instructions to execute and data to process in order to execute the processes of some embodiments.
[0164] The bus 2205 also connects to the input and output devices 2240 and 2245. The input devices 2240 enable the user to communicate information and select commands to the electronic system. The input devices 2240 include alphanumeric keyboards and pointing devices (also called “cursor control devices” ) , cameras (e.g., webcams) , microphones or similar devices for receiving voice commands, etc. The output devices 2245 display images generated by the electronic system or otherwise output data. The output devices 2245 include printers and display devices, such as cathode ray tubes (CRT) or liquid crystal displays (LCD) , as well as speakers or similar audio output devices. Some embodiments include devices such as a touchscreen that function as both input and output devices.
[0165] Finally, as shown in FIG. 22, bus 2205 also couples electronic system 2200 to a network 2225 through a network adapter (not shown) . In this manner, the computer can be a part of a network of computers (such as a local area network ( “LAN” ) , a wide area network ( “WAN” ) , or an Intranet, or a network of networks, such as the Internet. Any or all components of electronic system 2200 may be used in conjunction with the present disclosure.
[0166] Some embodiments include electronic components, such as microprocessors, storage and memory that store computer program instructions in a machine-readable or computer-readable medium (alternatively referred to as computer-readable storage media, machine-readable media, or machine-readable storage media) . Some examples of such computer-readable media include RAM, ROM, read-only compact discs (CD-ROM) , recordable compact discs (CD-R) , rewritable compact discs (CD-RW) , read-only digital versatile discs (e.g., DVD-ROM, dual-layer DVD-ROM) , a variety of recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc. ) , flash memory (e.g., SD cards, mini-SD cards, micro-SD cards, etc. ) , magnetic and / or solid state hard drives, read-only and recordable discs, ultra-density optical discs, any other optical or magnetic media, and floppy disks. The computer-readable media may store a computer program that is executable by at least one processing unit and includes sets of instructions for performing various operations. Examples of computer programs or computer code include machine code, such as is produced by a compiler, and files including higher-level code that are executed by a computer, an electronic component, or a microprocessor using an interpreter.
[0167] While the above discussion primarily refers to microprocessor or multi-core processors that execute software, many of the above-described features and applications are performed by one or more integrated circuits, such as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) . In some embodiments, such integrated circuits execute instructions that are stored on the circuit itself. In addition, some embodiments execute software stored in programmable logic devices (PLDs) , ROM, or RAM devices.
[0168] As used in this specification and any claims of this application, the terms “computer” , “server” , “processor” , and “memory” all refer to electronic or other technological devices. These terms exclude people or groups of people. For the purposes of the specification, the terms display or displaying means displaying on an electronic device. As used in this specification and any claims of this application, the terms “computer readable medium, ” “computer readable media, ” and “machine readable medium” are entirely restricted to tangible, physical objects that store information in a form that is readable by a computer. These terms exclude any wireless signals, wired download signals, and any other ephemeral signals.
[0169] While the present disclosure has been described with reference to numerous specific details, one of ordinary skill in the art will recognize that the present disclosure can be embodied in other specific forms without departing from the spirit of the present disclosure. In addition, a number of the figures (including FIG. 18 and FIG. 21) conceptually illustrate processes. The specific operations of these processes may not be performed in the exact order shown and described. The specific operations may not be performed in one continuous series of operations, and different specific operations may be performed in different embodiments. Furthermore, the process could be implemented using several sub-processes, or as part of a larger macro process. Thus, one of ordinary skill in the art would understand that the present disclosure is not to be limited by the foregoing illustrative details, but rather is to be defined by the appended claims. Additional Notes
[0170] The herein-described subject matter sometimes illustrates different components contained within, or connected with, different other components. It is to be understood that such depicted architectures are merely examples, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected" , or "operably coupled" , to each other to achieve the desired functionality, and any two components capable of being so associated can also be viewed as being "operably couplable" , to each other to achieve the desired functionality. Specific examples of operably couplable include but are not limited to physically mateable and / or physically interacting components and / or wirelessly interactable and / or wirelessly interacting components and / or logically interacting and / or logically interactable components.
[0171] Further, with respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.
[0172] Moreover, it will be understood by those skilled in the art that, in general, terms used herein, and especially in the appended claims, e.g., bodies of the appended claims, are generally intended as “open” terms, e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles "a" or "an" limits any particular claim containing such introduced claim recitation to implementations containing only one such recitation, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an, " e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more; ” the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number, e.g., the bare recitation of "two recitations, " without other modifiers, means at least two recitations, or two or more recitations. Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention, e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”
[0173] From the foregoing, it will be appreciated that various implementations of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various implementations disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.
Claims
1.A video coding method comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;deriving or selecting one or more first target block vectors for the current block;locating one or more reference areas in the current picture by using the first target block vectors;generating a final predictor of the current block based on coding information derived from samples of the located one or more reference areas;encoding or decoding the current block by using the final predictor; andencoding or decoding a subsequent block by referencing the first target block vectors.2.The video coding method of claim 1, wherein the coding information comprises at least one of a filter-based model, a parameter model, a cross-component model, and an intra prediction mode.3.The video coding method of claim 1, wherein the coding information comprises one or more block-vector-based (BV-based) predictors that are generated based on sample values of the reference areas located by the first target block vectors.4.The video coding method of claim 3, wherein the final predictor is a combination of the one or more BV-based predictors that are generated based on the first target block vectors and one or more other predictors.5.The video coding method of claim 4, wherein the one or more other predictors are generated based on one or more intra direction modes.6.The video coding method of claim 5, wherein the one or more intra direction modes are determined from a histogram of gradient (HoG) based on a texture gradient analysis of samples in the one or more reference areas located by the first target block vectors.7.The video coding method of claim 1, wherein the one or more first target block vectors is selected from an intra prediction mode (IPM) list of a partition mode of spatial geometric partition mode (SGPM) , the IPM list augmented with block vector based prediction candidates obtained from adjacent and non-adjacent merge candidates coded in intra block copy (IBC) or intra template matching (IntraTMP) .8.The video coding method of claim 1, wherein the subsequent block is encoded or decoded by using a merge candidate list comprising one or more second target block vectors that are stored for neighboring blocks of the subsequent block, wherein the second target block vectors include the first target block vectors that are stored for the current block.9.The video coding method of claim 1, wherein the subsequent block is encoded or decoded based on one or more block vectors selected from one or more second target block vectors that are stored for neighboring blocks of the subsequent block, wherein the neighboring blocks of the subsequent block comprise the current block and the one or more second target block vectors comprise the first target block vectors stored for the current block.10.The video coding method of claim 9, wherein block vector based (BV-based) predictors are generated for the subsequent block based on sample values of the reference areas located by the one or more block vectors selected from the second target block vectors, and a final predictor of the subsequent block is a combination of the BV-based predictors and one or more other predictors.11.The video coding method of claim 1, wherein the first target block vectors are used for intra block copy (IBC) prediction of the current block.12.The video coding method of claim 1, wherein the first target block vectors are derived by searching the current picture for template areas matching that of the current block.13.An electronic apparatus comprising:a video coder circuit configured to perform operations comprising:receiving data to be encoded or decoded as a current block of pixels of a current picture of a video;deriving or selecting one or more first target block vectors for the current block;locating one or more reference areas in the current picture by using the first target block vectors;generating a final predictor of the current block based on coding information derived from samples of the located one or more reference areas;encoding or decoding the current block by using the final predictor; andencoding or decoding a subsequent block by referencing the first target block vectors.14.A video decoding method comprising:receiving data to be decoded as a current block of pixels of a current picture of a video;deriving or selecting one or more first target block vectors for the current block;locating one or more reference areas in the current picture by using the first target block vectors;generating a final predictor of the current block based on coding information derived from samples of the located one or more reference areas;reconstructing the current block by using the final predictor; anddecoding a subsequent block by referencing the first target block vectors.