Encoding / Decoding of Video Image Data
By integrating intra-template matching prediction and intra-block copy modes through shared block vectors, the method enhances video compression efficiency by reducing bitrate costs in IBC prediction, addressing the suboptimal integration of these modes in existing systems.
Patent Information
- Application Number
- JP2024577100
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-04-24
- Publication Date
- 2025-07-23
AI Technical Summary
Existing video compression systems like HEVC and VVC do not effectively integrate intra-template matching prediction (ITMP) and intra-block copy (IBC) modes, leading to suboptimal compression performance when both modes are enabled for the same video image.
The method involves determining a first block vector through intra-template matching prediction (ITMP) and storing it in a block-based motion information buffer, then using this vector as a candidate for intra-block copy (IBC) prediction, allowing interaction between the two modes to enhance compression efficiency.
This approach improves the overall coding efficiency by reducing bitrate costs for signaling block vector information in the IBC prediction mode, resulting in more efficient video image compression.
Smart Images

Figure 2025523586000001_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims priority to European Patent Application No. "22306005.4" filed on July 5, 2022, the entire content of which is incorporated herein by reference.
Technical Field
[0002] This application generally relates to the encoding and decoding of video images. Specifically and non-exclusively, the technical field of this application relates to the prediction of video image blocks based on intra-template matching.
Background Art
[0003] This section is intended to introduce the reader to various aspects of the art, which are related to aspects of at least one exemplary embodiment of the present application described below and / or claimed for protection. This description is believed to be useful in providing the reader with background information to better understand the various aspects of this application. Therefore, these techniques do not admit prior art and should be read from the above perspective.
[0004] In state-of-the-art video compression systems such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level image partitions are provided to divide the video image into image regions called coding tree units (CTUs). In the case of HEVC, the size of the coding tree unit (CTU) is usually between 16×16 pixels and 64×64 pixels, and in the case of VVC, the size of the coding tree unit (CTU) may be 32×32, 64×64 or 128×128 pixels.
[0005] The CTU segmentation of a video image forms a grid of CTUs of a fixed size, i.e., a CTU grid, where the upper and left boundaries of this grid spatially overlap with the upper and left boundaries of the video image. The CTU grid represents the spatial partition of the video image.
[0006] In VVC and HEVC, the CTU sizes (CTU width and CTU height) of all CTUs in the CTU grid are equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) may be 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is encoded in the bitstream, for example, at the sequence level in the sequence parameter set (SPS).
[0007] The spatial position of a CTU within the CTU grid is determined based on the CTU address ctuAddr, which defines the spatial position from the origin of the upper left corner of the CTU. As shown in FIG. 1, the CTU address can define the spatial position from the upper left corner of the upper-level spatial structure S that contains the CTU.
[0008] An encoding tree is associated with each CTU to determine the tree partition of the CTU.
[0009] As shown in FIG. 1, in HEVC, the coding tree is a quadtree partition of a CTU, and each leaf is called a coding unit (CU). The spatial position of a CU in a video image is defined by a CU index cuIdx, and the CU index cuIdx indicates the spatial position from the upper left corner of the CTU. A CU is spatially divided into one or more prediction units (PUs). The spatial position of a PU in a video image VP is defined by a PU index puIdx, the PU index puIdx defines the spatial position from the upper left corner of the CTU, and the spatial position of the elements of the divided PU is defined by a PU partition index puPartIdx, and the PU partition index puPartIdx defines the spatial position from the upper left corner of the PU. Some intra or inter prediction data is assigned to each PU.
[0010] The intra or inter coding mode is assigned at the CU level. This means that the same intra / inter coding mode is assigned to each PU of a CU, although the prediction parameters may vary from PU to PU.
[0011] Based on a quadtree called a transform tree, a CU can be spatially divided into one or more transform units (TUs). A transform unit is a leaf of the transform tree. The spatial position of a TU in a video image is defined by a TU index tuIdx, and the TU index tuIdx defines the spatial position from the upper left corner of the CU. Some transform parameters are assigned to each TU. The transform type is assigned at the TU level, and a 2D independent transform is performed at the TU level during the encoding or decoding of an image block.
[0012] The PU partition types in HEVC are as shown in Figure 2. They include the square partitions (2N×2N and N×N), which are the only partitions used in both intra-predicted CUs and inter-predicted CUs, the symmetric non-square partitions (2N×N, N×2N, used only in inter-predicted CUs), and the asymmetric partitions (used only in inter-predicted CUs). For example, PU type 2N×nU represents an asymmetric horizontal partition of the PU, and the small partition is located at the top of the PU. According to another example, PU type 2N×nL represents an asymmetric horizontal partition of the PU, and the small partition is located at the top of the PU.
[0013] As shown in Figure 3, in VVC, the coding tree starts from the root node (i.e., CTU). Then, the quadtree (or quad tree) splitting divides the root node into four nodes corresponding to four sub-blocks (solid lines) of the same size. Subsequently, the leaf of the quad tree (or quadtree) can be further divided by a so-called multi-type tree, which is related to a binary split or a ternary split of one of the four split modes shown in Figure 4. These split types are the vertical and horizontal binary split modes denoted as SBTV and SBTH, and the vertical and horizontal ternary split modes SPTTV and STTH.
[0014] In the case of a common coding tree where the luma component and the chroma component are shared, the leaf of the coding tree of the CTU is the CU.
[0015] Contrary to HEVC, in VVC, in most cases, the CUs, PUs, and TUs have the same size, which means that except for some specific coding modes, the coding unit is generally not divided into PUs or TUs.
[0016] Figures 5 and 6 provide an overview of the video encoding / decoding method used in current video standard compression systems (e.g., HEVC or VVC).
[0017] FIG. 5 shows an exemplary block diagram of the steps of a method 100 for encoding a video image VP based on the prior art.
[0018] In step 110, the video image VP is divided into sample blocks, and the division information data is signaled to the bitstream. Each block contains samples of one component of the video image VP. Thus, these blocks contain samples of each component that defines the video image VP.
[0019] For example, in HEVC, an image is divided into coding tree units (CTUs). Each CTU can be further subdivided by quadtree partitioning, where each leaf of the quadtree represents a coding unit (CU). And the division information data may include data describing the CTUs and the quadtree subdivision of each CTU.
[0020] Thereafter, each sample block (abbreviated as block) may be a CU (when the CU contains a single PU) or a PU of the CU.
[0021] Each block is encoded along an encoding loop (also called "inside the loop") using an intra or inter prediction mode.
[0022] Intra prediction (step 120) uses intra prediction data. Intra prediction predicts the current block using an intra prediction block based on samples that are already encoded, decoded, and reconstructed and are located around the current block (usually at the top and left of the current block). Intra prediction is performed in the spatial domain.
[0023] In the inter prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for a candidate reference block as a good predictor of the current block in one or more reference video images that predictively encode the current video image. For example, a good predictor of the current block is a predictor similar to the current block. The output of the motion estimation step 130 is inter prediction data including motion information (usually one or more motion vectors and one or more reference video image indexes) associated with the current block, and other information for obtaining the same prediction block on the encoding / decoding side. Subsequently, motion compensation (step 135) obtains a prediction block using the (one or more) motion vectors and (one or more) reference video image indexes determined in the motion estimation step 130. Basically, the block belonging to the selected reference video image and pointed to by the motion vector is available as the prediction block of the current block. Also, since the motion vector is represented as a fraction of an integer pixel position (known as sub-pixel MV accuracy representation), motion compensation usually includes spatial interpolation of some reconstructed samples of the reference video image to calculate the prediction block.
[0024] The prediction information data is signaled in the bitstream. The prediction information may include a prediction mode, intra / inter prediction data, and any other information for obtaining the same prediction CU on the decoding side.
[0025] Method 100 selects a prediction mode (intra or inter prediction mode) by optimizing the rate-distortion trade-off by considering, for example, the encoding of the prediction residual block calculated by reducing candidate prediction blocks from the current block and the signaling of the prediction information data necessary for determining the candidate prediction block on the decoding side.
[0026] Normally, the best prediction mode is given as follows as the prediction mode of the best encoding mode p* for the current block.
Number
[0027] Here, P is the set of all candidate coding modes of the current block, p is a candidate coding mode in the set, and RD cost (p) is the rate-distortion cost of candidate coding mode p, and is usually expressed as follows. RD cost(p) =D(p)+λ.R(p)
[0028] D(p) is the distortion between the current block and the reconstructed block obtained by encoding / decoding the current block using candidate coding mode p, R(p) is the rate cost associated with the encoding of the current block by coding mode p, λ is the Lagrange parameter representing the rate constraint for encoding the current block, and is generally calculated based on the quantization parameter for encoding the current block.
[0029] Normally, the current block is encoded based on the prediction residual block PR. More precisely, for example, the prediction residual block PR is calculated by subtracting the best prediction block from the current block. Then, for example, a transform of the DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type or any other suitable transform is used to transform the prediction residual block PR (step 140), and the obtained transform coefficient block is quantized (step 150).
[0030] In a variant, method 100 can skip the transform step 140 and directly apply quantization (step 150) to the prediction residual block PR by skipping the so-called transform coding mode.
[0031] The quantized transform coefficient block (or the quantized prediction residual block) is entropy encoded into a bitstream (step 160).
[0032] Subsequently, as part of the encoding loop, inverse quantization (step 170) and inverse transformation (180) (or not inverse transformed) are performed on the quantized transform coefficient block (or quantized residual block) to generate a decoded prediction residual block. Thereafter, the decoded prediction residual block and the prediction block are combined, usually summed, which provides a reconstructed block.
[0033] Furthermore, other information data may be entropy encoded in step 160 so as to encode the current block of the video image VP.
[0034] Artifacts can be reduced by applying an in-loop filter (step 190) to the reconstructed image (including the reconstructed blocks). After all image blocks have been reconstructed, the loop filter can be applied. For example, these may include a deblocking filter, sample adaptive offset (SAO), or an adaptive loop filter.
[0035] The reconstructed block or the filtered reconstructed block is formed as a reference image, which can be stored in a decoded picture buffer (DPB), whereby it can be used as a reference image for the next current block of the video image VP or the next encoded video image to be encoded.
[0036] FIG. 6 shows an exemplary block diagram of the steps of a method 200 for decoding a video image VP based on the prior art.
[0037] In step 210, split information data, prediction information data, and a quantized transform coefficient block (or quantized residual block) are obtained by entropy decoding the bitstream of the encoded video image data. For example, this bitstream is generated based on method 100.
[0038] The current block of the video image VP can also be decoded from the bitstream by entropy-decoding other information data.
[0039] In step 220, based on the segmentation information, the reconstructed image is segmented into current blocks. Each current block is entropy-decoded from the bitstream along a decoding loop (also referred to as "inside the loop"). Each decoded current block is a quantized transform coefficient block or a quantized prediction residual block.
[0040] In step 230, the current block is inverse-quantized and, optionally, inverse-transformed (step 240) to obtain a decoded prediction residual block.
[0041] On the other hand, the prediction information data is used to predict the current block. A predicted block is obtained by its intra prediction (step 250) or motion-compensated temporal prediction (step 260). The prediction process executed on the decoding side is the same as the prediction process executed on the encoding side.
[0042] Subsequently, the decoded prediction residual block and the predicted block are merged, usually summed, which provides a reconstructed block.
[0043] In step 270, the in-loop filter is applicable to the reconstructed image (including the reconstructed blocks), and the reconstructed blocks or the filtered reconstructed blocks are formed as a reference image, which can be stored in the above-described decoded picture buffer (DPB) (Figure 5).
[0044] In step 130 / 135 of Figure 5 or step 260 of Figure 6, an inter prediction block is defined from the inter prediction data associated with the current block (CU or PU of the CU) of the video image. This inter prediction data includes motion information that can be displayed (encoded) based on the so-called AMVP mode (adaptive motion vector prediction) or the so-called merge mode.
[0045] In HEVC, in the AMVP mode, the motion information for defining an inter prediction block is represented by up to two reference video picture indices, and the up to two reference video picture indices are each associated with up to two reference video picture lists (usually represented as L0 and L1). The reference video picture reference list L0 contains at least one reference video picture, and the reference video picture reference list L1 contains at least one reference video picture. Each reference video picture index is a temporal prediction for the current block. The motion information further includes up to two motion vectors, and each motion vector is associated with one of the reference video picture indices of one of the two reference video picture lists. Each motion vector is predictively encoded and signaled in the bitstream, that is, one motion vector difference MVd is derived from the motion vector, one AMVP (adaptive motion vector predictor) candidate is selected from the AMVP candidate list (constructed on both the encoding and decoding sides), and MVd is signaled in the bitstream. The AMVP candidate index of the AMVP candidate selected from the AMVP candidate list is similarly signaled in the bitstream.
[0046] FIG. 7 shows an illustrative example for constructing an AMVP candidate list, which is used to define an inter prediction block of the current block of the current video picture.
[0047] The AMVP candidate list may include two spatial MVP (motion vector predictor) candidates derived from the current video picture. The first spatial MVP candidate is derived from the motion information associated with the inter prediction block of the current block and, if it exists, is located at the adjacent positions A0, A1 to the left of the current block. The second MVP candidate is derived from the motion information associated with the inter prediction block of the current block and, if it exists, is located at the adjacent positions B0, B1, and B2 at the top of the current block. Then, a redundancy check is performed among the derived spatial MVP candidates, i.e., the overlapping derived MVP candidates are discarded. The AMVP candidate list may further include a temporal MVP candidate, which is derived from the motion information associated with the co-located block (if it exists) at the spatial position H in the reference video picture or the spatial position C in other cases. The temporal MVP candidate is scaled based on the temporal distance between the current video picture and the reference video picture. Finally, if the AMVP candidate list contains less than two MVP candidates, the AMVP candidate list is filled with zero motion vectors.
[0048] In HEVC, in the merge mode, the motion information for defining an inter prediction block is represented by one merge index that merges the MVP candidate list. Each merge index points to the motion predictor information indicating which MVP is used to derive the motion information. The motion information is represented by one unidirectional or bidirectional temporal prediction type, a maximum of two reference video picture indices, and a maximum of two motion vectors, where each motion vector is associated with one of the two reference video picture lists (L0 or L1).
[0049] Information other than the merge index is not signaled. This means that the motion vector of the current block is equal to the motion vector of the MVP candidate indicated by the merge index. Therefore, contrary to the merge mode, the MVd and the reference picture index are not signaled in the bitstream. Only the index of the merge candidate selected from the merge candidate list is signaled in the bitstream.
[0050] Therefore, contrary to the AMVP mode, in the merge mode, the MVd and the reference picture are not signaled. Only the merge index is signaled in the bitstream.
[0051] The merge MVP candidate list may include five spatial MVP candidates, which are derived from the current video image shown in FIG. 7. The first spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the left adjacent position A1. The second spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the upper adjacent position B1. The third spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the upper-right adjacent position B0. The fourth spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the lower-left adjacent position A0. The fifth spatial MVP candidate is derived from the above motion associated with the inter-prediction block and, if it exists, is at the left adjacent position B2. Then, a redundancy check is performed among the derived spatial MVPs, i.e., the duplicate derived MVP candidates are discarded. The merge candidate list may further include a temporal MVP candidate called TMVP candidate, which is derived from the motion information associated with the collocated block (if it exists) located at position H or the central spatial position "C" of the reference video image. And a redundancy check is performed among the derived spatial MVPs, i.e., the duplicate derived MVP candidates are discarded. Finally, when using the bidirectional temporal prediction type, if the merge candidate list contains less than five MVP candidates, merge candidates are added to the merge candidate list. The merge candidate is associated with one reference video image list and is derived from the motion information of one MVP candidate existing in the merge candidate list, and the motion information corresponds to another MVP candidate existing in the merge candidate list and associated with another reference video image list. Finally, if the merge candidate list is still not (filled with five merge candidates), the merge candidate list is filled with zero motion vectors.
[0052] In ECM (Explanation of the Algorithm of Extended Compression Model 4 (ECM 4), M. Coban, F. Le Leannec, K. Naser, J. Strom, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 23rd meeting, teleconference, July 7 - 16, 2021, Document JVET-Y202-v2, https: / / jvet-experts.org / doc_end_user / documents / 25_Teleconference / wg11 / JVET-Y2025-v2.zip), the template matching (TM) method may be used to determine some of the MVP candidates in the AMVP and merge modes. The TM method is a method for subdividing motion vectors on the decoder side. As shown in Figure 8, it subdivides the motion vector of a block by matching a so-called L-shaped template located above and to the left of the current block with a template in the search area of the reference image. Search for a more appropriate MV around the initial MV of the current block, for example, search within a [-8, +8]-pixel search range. Obtain the subdivided MV by minimizing the so-called template matching cost TMcost between the template around the current block and the candidate template in the reference video image. Select the MVP candidate with the lowest template matching cost and further subdivide it.
[0053] When encoding motion information according to VVC, more abundant motion information can be represented than HEVC encoding.
[0054] In VVC, motion representation can be provided based on the AMVP mode or the merge mode.
[0055] In VVC, in the AMVP mode, the motion information for defining an inter prediction block is represented in the same way as in the HEVC AMVP mode. If present, the AMVP candidate list may include spatial and temporal MVP candidates, similar to HEVC. If present, the AMVP candidate list may further include four additional HMVP (history-based motion vector prediction) candidates. Finally, if the AMVP candidate list contains less than two MVP candidates, the AMVP candidate list is filled with zero motion vectors.
[0056] The HMVP candidates are derived from the previously encoded MVPs of adjacent or non-adjacent blocks associated with the current block. Therefore, an HMVP candidate table is maintained on both the encoder and decoder sides and updated in real time as a first-in first-out (FIFO) buffer for MVPs. There are up to five HMVP candidates in the table. After encoding one block, the relevant motion information is added to the end of the table to update the table as a new HMVP candidate. The table is managed by applying the FIFO rule, and in addition to the basic FIFO mechanism, redundant candidates in the HMVP table are deleted first instead of the first candidate. The table is reset for each CTU row to enable parallel processing.
[0057] In the AMVP mode, the AMVR (adaptive motion vector resolution) algorithm is adopted. The AMVR tool can signal the MVd at a luminance sample resolution of 1 / 4 pixel, 1 / 2 pixel, integer pixel, or 4 pixels. This can save bits when encoding the MVd information. In AMVR, the resolution of the motion vector is selected at the block level.
[0058] Finally, the internal motion vector representation is realized with a luminance sample accuracy of 1 / 16 instead of 1 / 4 luminance sample accuracy in HEVC.
[0059] In VVC, in the merge mode, the motion information for defining an inter prediction block is represented in the same way as in the HEVC merge mode.
[0060] The merge candidate list is different from the merge candidate list used in HEVC.
[0061] In VVC, both the MMVD (Merge Mode with Motion Vector Difference) mode and the CIIP (Composite Intra / Inter Prediction) mode can establish a merge candidate list.
[0062] The MMVD mode can represent the motion information associated with an inter prediction block by encoding a limited motion vector difference (MVd) in the selected merge candidate. MMVD encoding is limited to four vector directions, eight amplitude values, and from 1 / 4 luminance samples to 32 luminance samples, as shown in FIG. 9. The MMVD mode can perform an intermediate trade-off between rate cost and MV (motion vector) accuracy to provide an intermediate accuracy level and signal the motion information. The MMVD offset can also be signaled in the bitstream. Then, the motion vector is derived by adding the MV predictor to the MMVD offset.
[0063] CIIP combines an inter prediction signal and an intra prediction signal to predict the current block of a video image. The inter prediction signal in the CIIP mode is derived using a prediction process similar to that applied to the merge mode. The intra prediction signal is derived according to a normal intra prediction process having a Planar mode. The planar prediction mode predicts a block through a spatial interpolation process of adjacent reconstructed samples of the prediction blocks at the top and left of the block. Then, the intra and inter prediction signals are blended (composited) using a weighted average, where the calculation of the weighting value is based on the coding modes of the adjacent blocks at the top and left as follows. PCIIP = (Wmerge × Pmerge + Wintra × Pintra + 2) >> 2 Weight W merge and W intra The sum of is 4, and these weights are constant throughout the current block.
[0064] In short, in VVC, the merge candidate list may contain spatial MVP candidates similar to HEVC, and only two first candidates are swapped. During the construction of the merge candidate list, candidate B1 is considered before candidate A1. The merge candidate list may further include TMVP, for example, the HEVC, HMVP candidates in the VVC AMVP mode. Some HMVP candidates are inserted into the merge candidate list so that the merge candidate list reaches the maximum allowable number (-1) of MVP candidates. The merge candidate list may further include an average candidate that results in at most one pair calculated as follows. Consider two first merge candidates present in the merge candidate list and average their motion vectors. This average value is calculated separately for each reference video picture list L0 and L1. Therefore, if both MVPs are bi-directional, the motion vectors associated with the two lists L0 and L1 are averaged. If there is only one motion vector in the reference video picture list, it is taken as it is and a paired candidate is formed. Finally, if the merge candidate list has not reached the maximum allowable number of MVP candidates, the merge candidate list is filled with zero motion vectors.
[0065] In HEVC and VVC, screen content is encoded using the Intra Block Copy (IBC) prediction mode. The IBC prediction mode is known to significantly improve the encoding efficiency of screen content materials. Since the IBC prediction mode is implemented as a block-level encoding mode, block matching (BM) is performed in the encoder to find the optimal block vector (or motion vector) for each block to be predicted. The block vector indicates the displacement from the block to be predicted in the video image to the reference block (i.e., the predicted block in the reconstructed area of the video image that has already been reconstructed (decoded) within the video image). The block vector of the IBC prediction block of the chroma samples can also be rounded to integer precision. When merged with AMVR, the IBC prediction mode can switch between 1-pixel and 4-pixel motion vector precisions. The IBC prediction mode is regarded as a third prediction mode other than the intra or inter prediction mode.
[0066] At the CU level, the IBC prediction mode is signaled as prediction information having a flag indicating that the IBC-AMVP mode or the IBC-skip / merge mode is being used.
[0067] In the IBC-skip / merge mode, the block vector for defining the IBC prediction block is indicated by a single merge index. This merge index indicates an element in the list of block vector prediction candidates (merge candidate list). The merge candidate list may include spatial candidates, HMVP candidates, and paired block vector candidates.
[0068] The HMVP (History-based Motion Vector Prediction) candidates are obtained in the same way as in the case of conventional interlaced coding. HMVP is related to the buffer of block vector (motion vector) candidates, and the buffer is provided as long as the blocks of the video image are predicted using the IBC prediction mode. And the HMVP buffer includes the buffer of block vector candidates, and provides block vector candidates to predict the current block of the video image in the IBC prediction mode.
[0069] The paired block vector candidates can be generated by averaging two IBC block vector candidates (i.e., the block vectors derived from the IBC prediction mode). This means averaging two first block vector candidates in the constructed merge candidate list to form a so-called paired block vector candidate. This paired block vector candidate is added to the merge candidate list after the HMVP candidates.
[0070] For HMVP, the block vectors are inserted into the history buffer for future reference.
[0071] In the IBC-AMVP mode, two block vectors are determined, one from the adjacent area to the left of the block to be predicted and the other from the adjacent area above the block to be predicted. If either adjacent area is not available, the default block vector is considered. The flag is signaled as prediction information indicating which block vector to use for predicting the block vector of the current block. In fact, since a maximum of two block vector candidates are considered in the IBC-AMVP mode, the flag is sufficient to identify the block vector candidate used for the coding of the block vector information. The block vector difference is coded in a similar way to the motion vector difference of the inter prediction block.
[0072] The IBC prediction mode cannot be used in combination with the inter prediction tools of VVC such as CIIP and MMVD.
[0073] In VVC, the IBC prediction mode shares the same process as the normal MV merge mode that includes a pair of merge candidates and a motion predictor based on history, but TMVP and the zero vector are not used because they are invalid for the IBC prediction mode.
[0074] The intra-template matching prediction (ITMP) mode is a special intra-prediction mode for the current block of a video image. It copies the optimal prediction block from the reconstructed block of the video image, and its L-shaped template matches the L-shaped template of the current block to be predicted in the video image. Basically, the encoder searches for the L-shaped template of the reconstructed block in the video image that most resembles the L-shaped template of the current block within a predefined search range and uses the corresponding block as the optimal prediction block. Then, the encoder signals the use of the ITMP mode as prediction information, and the same prediction operation is performed on the decoder side.
[0075] In ECM, as shown in FIG. 10, the search range includes four spatial regions R1, R2, R3, and R4. Region R1 is the current CTU containing the current block b to be predicted in the video image, region R2 is the upper left CTU, R3 is the upper CTU, and R4 is the left CTU. The cost function is used to evaluate the matching between the L-shaped template of the current block b and the L-shaped template of the prediction block candidate. The optimal prediction block corresponds to the minimum value of the cost function. For example, the cost function calculated between two L-shaped templates is the sum of absolute differences (SAD) between the samples of the L-shaped templates.
[0076] The width SearchRange_w and height SearchRange_h of regions R1 - R3 are set to be proportional to the width BLkW and height BLkH of the current block so as to have a certain number of SAD comparisons per pixel. It is as follows.
Equation
[0077] For blocks with a width and height of 64 or less of the maximum value, enable the ITMP mode. This maximum size is configurable.
[0078] Use the dedicated flag of the current block to signal the ITMP mode as prediction information at the CU level.
[0079] The problem to be solved by the present invention is to improve the compression performance of ECM and VVC, especially when enabling the prediction modes of IBC and ITMP simultaneously for the same video image.
[0080] In the prior art, the ITMP or IBC prediction mode can be used to predict the blocks of a slice of a video image. However, these are used separately and there is no interaction between the two prediction modes.
[0081] Considering the above-described situation, at least one exemplary embodiment of the present application is designed.
Summary of the Invention
[0082] The following section provides a basic understanding of some aspects of the present application by showing a simplified summary of at least one exemplary embodiment. This summary is not a detailed summary of the exemplary embodiment. It does not identify the essential or important elements of the exemplary embodiment. The following summary only shows some aspects in a simplified form as a prelude to the more detailed description provided in other parts of the document for at least one exemplary embodiment.
[0083] According to a first aspect of the present application, a method for predicting a block of a video image is provided. The first prediction block is obtained by an intra-template matching prediction mode. The intra-template matching prediction mode determines the first prediction block by minimizing a cost function calculated between the sample values of the L-shaped template of the block of the video image and the sample values of the L-shaped template of the reconstructed block of the video image. Here, the method further includes a step of determining a first block vector as the displacement between the first prediction block and the block of the video image. The first block vector identifies the first prediction block as a prediction block candidate of the block of the video image.
[0084] In an exemplary embodiment, the first block vector is stored in a block-based motion information buffer.
[0085] In an exemplary embodiment, the first block vector is stored based on sub-blocks.
[0086] In an exemplary embodiment, the method further includes a step of adding the first block vector to a list of block vector prediction candidates associated with the video image.
[0087] In an exemplary embodiment, the block of the video image is predicted by a second prediction block. The second prediction block is obtained by an intra-block copy prediction mode. The intra-block copy prediction mode determines the second block vector as the displacement between the block of the video image and the prediction block of the reconstructed region of the video image by block matching. Here, the method further includes a step of predicting the second block vector by the block vectors in the list of block vector prediction candidates.
[0088] In an exemplary embodiment, the method further includes a step of storing the first block vector in a history-based motion vector prediction table.
[0089] In an exemplary embodiment, the motion vector prediction table based on the history further stores at least one third block vector associated with at least one prediction block of at least one reconstructed block of the video image, and each of the at least one third block vector is obtained by an intra-block copy prediction mode, and the intra-block copy prediction mode determines the third block vector as a displacement between the prediction block of the reconstructed area of the video image and the reference block of the video image by block matching.
[0090] In an exemplary embodiment, a block of the video image is predicted by a second prediction block, and the second prediction block is obtained by an intra-block copy prediction mode, and the intra-block copy prediction mode determines the second block vector as a displacement between the block of the video image and the prediction block of the reconstructed area of the video image by block matching, where the method further includes predicting the second block vector by the block vector of the motion vector prediction table based on the history.
[0091] According to a second aspect of the present application, there is provided a method for encoding a video image block based on a prediction block derived from the method of the first aspect.
[0092] According to a third aspect of the present application, there is provided a method for decoding a video image block based on a prediction block derived from the method of the first aspect.
[0093] According to a fourth aspect of the present application, there is provided a bitstream formatted to include the encoded video image data obtained from the method of the first aspect.
[0094] According to a fifth aspect of the present application, there is provided an apparatus including means for executing one of the methods of the first, second, and / or third aspects.
[0095] According to a sixth aspect of the present application, when a program is executed by one or more processors, there is provided a computer program product including instructions for causing the one or more processors to execute the method according to the first, second, and / or third aspects.
[0096] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying instructions of program code for executing the method according to the first, second, and / or third aspects.
[0097] The specific nature of at least one embodiment in the exemplary embodiments and other objects, advantages, features, and uses of the at least one embodiment in the exemplary embodiments will become apparent from the description given with reference to the examples in combination with the following drawings.
Brief Description of the Drawings
[0098] Here, by way of example, reference is made to the drawings of the exemplary embodiments of the application.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Best Mode for Carrying Out the Invention
[0099] Hereinafter, at least one exemplary embodiment will be described more fully with reference to the drawings, where an example of at least one exemplary embodiment of the exemplary embodiment is described. However, the exemplary embodiments can be implemented in many alternative forms and should not be construed as limited to the examples described herein. Therefore, it should be understood that there is no intention to limit the exemplary embodiments to the specific forms disclosed. Rather, this application is intended to cover all modifications, equivalent substitutions, and alternatives within the spirit and scope of this application.
[0100] At least one of these aspects generally relates to the encoding and decoding of video images, another aspect generally relates to the transmission of a bitstream to be provided or encoded, and one of the other aspects relates to the reception / access of a decoded bitstream.
[0101] Although at least one exemplary embodiment of these exemplary embodiments is described as encoding / decoding video images, since each video image is sequentially encoded / decoded as described below, it is also extended to encoding / decoding of video images (image sequences).
[0102] Also, at least one exemplary embodiment is not limited to the current version of VVC. The at least one exemplary embodiment can be applied to existing or future-developed VVC and extensions of recommendations. Unless otherwise specified or technically excluded, each aspect described in this application can be used individually or in combination.
[0103] A pixel corresponds to the smallest display unit on the screen, and it can be composed of one or more light sources (one for a monochrome screen, three or more for a color screen).
[0104] A video image, also referred to as a frame or an image frame, includes at least one component (also called an image component or channel) determined by a specific image / video format, and the specific image / video format specifies all information associated with pixel values and all information for displaying and / or decoding video image data related to the video image by being used by a display unit and / or any other device.
[0105] A video image usually includes at least one component represented in the form of an array of samples.
[0106] A monochrome video image may include a single component, and a color video image may include three components.
[0107] For example, if the image / video format is in the well-known (Y,Cb,Cr) format, a color video image may include a luma (or luminance) component and two chroma components, or if the image / video format is in the well-known (R,G,B) format, a color video image may include three color components (one for red, one for green, and one for blue).
[0108] Each component of a video image may contain a number of samples relative to the number of pixels of the screen on which the video image is displayed. In a variant, the number of samples contained in a component may be a multiple (or fraction) of the number of samples contained in other components of the same video image.
[0109] For example, if a video format includes a luma component and two chroma components (e.g., a (Y,Cb,Cr) format), then depending on the color format considered, the chroma components may include half the number of samples in width and / or height compared to the luma component.
[0110] A sample is the smallest visual information unit of a component that makes up a video image. A sample value may be, for example, a luminance or chroma value or a color value in (R,G,B) format.
[0111] A pixel value is the value of a screen pixel. For monochrome video images, a pixel value can be represented by one sample, and for color video images, a pixel value can be represented by multiple co-located samples. The co-located samples associated with a pixel refer to the samples that correspond to the pixel's location on the screen.
[0112] A video image is typically viewed as a set of pixel values, with each pixel represented by at least one sample.
[0113] A block of a video image is a set of samples of one component of the video image. When the image / video format is the well-known (Y, Cb, Cr) format, a block of at least one luminance sample or a block of at least one chrominance sample is considered, or when the image / video format is the well-known (R, G, B) format, a block of at least one color sample is considered.
[0114] At least one exemplary embodiment is not limited to a specific image / video format.
[0115] Generally, as described above with respect to the prior art, when predicting the current block (coding unit or prediction unit) of a video image in the ITMP mode, a search is performed based on an L-shaped template of the optimal prediction block of the current block in the coded, decoded, and reconstructed regions for coding.
[0116] The present invention also determines a first block vector between the current block and the optimal prediction block.
[0117] Compared with the prior art, the present invention converts the difference in the spatial position between the optimal prediction block and the current block into a block vector connecting the two blocks, that is, the block vector represents the displacement between the current block and the optimal prediction block.
[0118] The present invention provides an interaction between ITMP and the IBC prediction mode (IBC skip / merge mode or IBC-AMVP mode) because the block vector of the adjacent block of the current block to be predicted, which is predicted using the IBC prediction or the ITMP mode, can predict the block vector determined for the current block using the IBC prediction mode.
[0119] Thereafter, the block vector of the adjacent block may be reused and spatially propagated to a subsequent block to be predicted in the video image. In this way, the IBC prediction mode can potentially benefit from the block vector information transmitted from the ITMP mode.
[0120] This results in more efficient signaling of the block vector information in the case of the IBC prediction mode, that is, reducing the bitrate cost for signaling the block vector information in the IBC prediction block.
[0121] As a result, the overall coding efficiency of the video image is improved.
[0122] FIG. 11 is a block diagram of a method 300 for predicting a video image block according to an exemplary embodiment.
[0123] Method 300 provides a prediction block for a video image block. This prediction block can be used for the selection of the prediction mode of method 100 or the prediction mode of the prediction process of method 200.
[0124] Signal the prediction information related to method 300, that is, when using method 300 in method 100, write the prediction information into the bitstream, and when using method 300 in method 200, analyze the prediction information from the bitstream.
[0125] In step 301, as in the introduction part, the first prediction block PB1 is obtained by the ITMP mode. Basically, the ITMP mode determines the first prediction block PB1 by minimizing a cost function calculated between the sample values of the L-shaped template of the block of the video image (to be predicted) and the sample values of the L-shaped template of the reconstructed block of the video image.
[0126] In step 302, as shown in FIG. 12, the first block vector BV1 is determined as the displacement between the first prediction block PB1 and the block B of the video image, and the first block vector BV1 identifies the first prediction block as a prediction block candidate of the block of the video image.
[0127] As introduced, the block B in FIG. 12 is predicted by the optimal prediction block PB1 obtained by the ITMP mode. Basically, the first prediction block PB1 is searched by minimizing the cost function calculated between the sample values of the L-shaped template of block B and the sample values of the L-shaped template of the reconstructed block of the video image. The block vector BV1 is obtained as the displacement between block B and the optimal prediction block PB1 in the video image: BV1=(px - x, py - y) Here, (px, py) is the spatial position of the first prediction block PB1, and (x, y) is the spatial position of block B.
[0128] In an exemplary embodiment, the first block vector BV1 can be stored in a block-based motion information buffer.
[0129] In a variant, the first block vector BV1 is stored based on an N×M sub-block.
[0130] For example, N = M = 4.
[0131] For example, as shown in FIG. 13, the first block BV1 is stored in the sub-block associated with the adjacent block PBC of block PB2. In this example, the adjacent block having sub-blocks A0, A1, B0, B1, and B2 is predicted by the prediction block determined by the ITMP or IBC prediction mode, and the block vectors associated with these IBC- or ITMP-based prediction blocks can be stored in association with these sub-blocks.
[0132] In an exemplary embodiment, in step 303, the first block vector BV1 is added to a list L of block vector prediction candidates associated with the video image.
[0133] Therefore, the list L of block vector prediction candidates can include block vectors associated with prediction block candidates obtained by the IBC skip / merge mode or the IBC-AMBP mode or the ITMP mode.
[0134] In a variant of the exemplary embodiment, in step 304, a block of the video image is predicted by the second prediction block PB2. As described in the introduction, the second prediction block PB2 is obtained by the IBC prediction mode. Basically, the IBC prediction mode determines the second block vector BV2 as the displacement between the block of the video image and the reference block in the reconstructed region of the video image by block matching. In step 305, the second block vector BV2 is predicted by the block vector BV3 in the list L of block vector prediction candidates.
[0135] The difference between the second block vector BV2 and the block vector BV3 is signaled as prediction information.
[0136] According to the exemplary embodiment and the variant, the second block vector BV2 can be predicted by the block vector BV3 associated with the IBC or ITMP prediction block.
[0137] FIG. 14 is a block diagram schematically showing a method 400 for constructing a list L of block vector prediction candidates for predicting a second block vector BV2 according to an exemplary embodiment.
[0138] Briefly, the method 400 includes a loop of several spatial positions around the second prediction block PB2, and in each iteration, it is evaluated whether a certain block vector candidate BV3 is available as a potential candidate for predicting the second block vector BV2.
[0139] In step 401, the spatial position of the prediction block candidate PB3 around the second prediction block PB2 is considered.
[0140] In step 402, method 400 checks whether the prediction block candidate PB3 is associated with the block vector determined by the IBC or ITMP prediction mode.
[0141] If NO, the spatial position of the new prediction block candidate PB3 around the second prediction block PB2 is considered.
[0142] If YES, after step 402, steps 403 - 405 follow.
[0143] In step 403, the block vector candidate BV3 associated with the considered prediction block PB3 is obtained from the block - based motion information buffer.
[0144] In step 404, method 400 checks whether the block vector candidate BV3 is valid as the IBC block vector of the second block vector BV2, that is, whether the spatial propagation of the block vector BV3 is permitted to predict the second block vector BV2.
[0145] If the block pointed to by the block vector BV3 is inside the reconstructed region of the video image and conforms to a predefined range with allowable values of the block vector components, the block vector BV3 is considered valid as the IBC block vector of the second block vector BV2.
[0146] If NO, after step 404, step 406 follows.
[0147] If YES, in step 405, the block vector BV3 is added to the list L of block vector prediction candidates.
[0148] In step 406, method 400 checks whether all prediction block candidates PB3 around the second prediction block PB2 have been considered.
[0149] If YES, method 400 ends.
[0150] If NO, after step 406, step 407 follows, and the spatial position of the new prediction block candidate PB3 around the second prediction block PB2 is considered. After step 406, step 402 follows.
[0151] In an exemplary embodiment, in step 306, the first block vector BV1 is stored in the history-based motion vector prediction table HMVP.
[0152] In a variant, the table HMVP further stores at least one third block vector BV3 associated with at least one prediction block of at least one reconstructed block of the video image, each of the at least one third block vector BV3 being obtained by an IBC prediction mode.
[0153] Compared with the prior art, the HMVP table becomes richer, and a better block vector predictor candidate can be obtained to predict subsequent IBC prediction blocks by utilizing the HMVP block vector propagation mechanism of ECM.
[0154] In a variant, when predicting a block of the video image by the second prediction block PB2, in step 307, the second block vector PB2 is predicted by the block vector BV4 of the table HMVP.
[0155] In a variant of the method 400 in FIG. 14, a certain block vector prediction candidate stored in the HMVP table can be added to the list L of block vector prediction candidates constructed by the method 400.
[0156] FIG. 15 is an exemplary schematic block diagram of a system 600 that implements various aspects and exemplary embodiments.
[0157] System 600 may be incorporated as one or more devices that include various components described below. In various exemplary embodiments, system 600 may be configured to implement one or more aspects described in this application.
[0158] Examples of devices that make up all or part of system 600 include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles, and associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "caves" (systems that include multiple displays), servers, video encoders, video decoders, post-processors that process the output from a video decoder, pre-processors that provide input to a video encoder, web servers, video servers (e.g., broadcast servers, video broadcast servers, or network servers), still or video cameras, encoding or decoding chips, or any other arbitrary communication device. The elements of system 600 can be implemented singly or in combination within a single integrated circuit (IC), multiple ICs, and / or individual components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of system 600 can be distributed across multiple ICs and / or individual components. In various exemplary embodiments, system 600 can be communicatively coupled to other similar systems or electronic devices, for example, via a communication bus or dedicated input and / or output ports.
[0159] System 600 may include at least one processor 610, and the at least one processor 610 is configured to execute instructions loaded to implement each aspect described in this application. The processor 610 may include an embedded memory, an input / output interface, and various other circuits known in the art. System 600 may include at least one memory 620 (e.g., a volatile memory device and / or a non-volatile memory device). System 600 may include a storage device 640 including non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. By way of non-limiting example, the storage device 640 may include an internal storage device, a connected storage device, and / or a network-accessible storage device.
[0160] System 600 may include an encoder / decoder module 630 configured to provide encoded / decoded video image data, for example, by processing data. The encoder / decoder module 630 may include its own processor and memory. The encoder / decoder module 630 can represent one or more modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device can include one or both of an encoding and a decoding module. Also, the encoder / decoder module 630 can be implemented as a separate element of the system 600 or can be incorporated within the processor 610 as a combination of hardware and software known to those skilled in the art.
[0161] The program code loaded into the processor 610 or the encoder / decoder 630 to execute each aspect described in this application is stored in the storage device 640, loaded into the memory 620, and executed by the processor 610. According to various exemplary embodiments, during the execution of the processes described in this application, one or more of the processor 610, the memory 620, the storage device 640, and the encoder / decoder module 630 can store one or more of various items. Such stored items include, but are not limited to, video image data, information data used for encoding / decoding video image data, bitstreams, matrices, variables, equations, mathematical formulas, operations, and intermediate or final results of operation logic processing.
[0162] In some exemplary embodiments, the memory inside the processor 610 and / or the encoder / decoder module 630 may be used to store instructions and provide a working memory for the processes executed during encoding or decoding.
[0163] However, in other exemplary embodiments, a memory external to the processing device (for example, the processing device may be the processor 610 or the encoder / decoder module 630) is used for one or more of these functions. The external memory may be the memory 620 and / or the storage device 640, and for example, may be dynamic volatile memory and / or non-volatile flash memory. In some exemplary embodiments, the external non-volatile flash memory may be used to store the operating system of a television. In at least one exemplary embodiment, a high-speed external dynamic volatile memory such as RAM can be utilized as a working memory for video encoding and decoding operations such as, for example, the second part of MPEG-2 (also called ITU-T Recommendation H.262 and ISO / IEC 13818-2, and also called MPEG-2 video), AVC, HEVC, EVC, VVC, AV1, etc.
[0164] As shown in block 690, inputs to the elements of system 600 can be provided via various input devices. Such input devices may include, but are not limited to, the following (i) - (v). (i) An RF portion capable of receiving RF signals wirelessly transmitted, for example, by a broadcast device, etc., (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) a bus such as CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802 - 3) bus when the present invention is implemented in the automotive field.
[0165] In various exemplary embodiments, the input devices of block 690 may each have associated input processing elements as known in the art. For example, the RF portion may be associated with elements necessary for (i) selecting a desired frequency (also called selecting a signal or restricting a signal to a frequency band), (ii) down - converting the selected signal, (iii) re - band - limiting to a narrower frequency band to select a signal frequency band, called a channel in a particular exemplary embodiment, for example, (iv) demodulating the down - converted and band - limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF portion of various exemplary embodiments may include one or more elements for performing these functions, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down - converter, a demodulator, an error corrector, and a demultiplexer. The RF portion down - converts the received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. It may include a tuner for performing various functions.
[0166] In one example of a set-top box, the RF portion and associated input processing elements can receive RF signals transmitted via a wired (e.g., cable) medium. The RF portion can then perform filtering, down-conversion, and re-filtering to perform frequency selection in the desired frequency band.
[0167] In various exemplary embodiments, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements performing similar or different functions are added.
[0168] Adding elements includes inserting elements between conventional elements, for example, inserting amplifiers or analog or digital converters. In various exemplary embodiments, RF may include an antenna.
[0169] Also, the USB and / or HDMI terminals may include corresponding interface processors to connect the system 600 to other electronic devices via SB and / or HDMI connections. U. Note that, when necessary, each aspect of input processing (e.g., Reed-Solomon error correction) can be implemented within an independent input processing IC or within the processor 610. Similarly, when necessary, each aspect of USB or HDMI interface processing can be implemented within an independent interface IC or within the processor 610. Demodulation, error correction, and demultiplexed streams can be provided to various processing elements, for example, including the processor 610 and the encoder / decoder 630, which operate in combination with memory and storage elements to process the data stream and display it on an output device when necessary.
[0170] The various elements of the system 600 can be provided within an integrated housing. Within the integrated housing, appropriate connection arrangements 690 can be used, for example, internal buses (including I2C buses), wiring, and printed circuit boards known in the art to interconnect the various types of elements and transmit data therebetween.
[0171] The system 600 may include a communication interface 650 to enable communication with other devices via a communication channel 651. The communication interface 650 may include, but is not limited to, a transceiver for transmitting and receiving data over the communication channel 651. The communication interface 650 includes, but is not limited to, a modem or a network card, and the communication channel 651 can be implemented, for example, within a wired and / or wireless medium.
[0172] In various exemplary embodiments, data can be streamed to the system 600 using a Wi-Fi network such as IEEE802.11. The Wi-Fi signals of these exemplary embodiments can be received via a communication channel 651 and a communication interface 650 that are compliant with Wi-Fi communication. The communication channel 651 of these exemplary embodiments can typically be connected to an access point or a router, and the access point or router enables streaming applications and other over-the-top communications by providing access to an external network including the Internet.
[0173] Other exemplary embodiments can provide data streamed to the system 600 using a set-top box, and the set-top box distributes the data via an HDMI connection of the input block 690.
[0174] Other exemplary embodiments can provide the system 600 with data streamed using an RF connection of the input block 690.
[0175] The streamed data can be used as a method for signaling information used by the system 600. The signaling information may include information on the bitstream B and / or the number of pixels of the video image and / or any encoding / decoding setting parameters.
[0176] Note that signaling can be implemented in various ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. may be used to signal information to the corresponding decoder.
[0177] System 600 can provide output signals to various output devices including display 661, speaker 671, and other peripheral devices 681. In various examples of the exemplary embodiments, other peripheral devices 681 may include one or more of an independent DVR, a disc player, a stereo system, an illumination system, and other devices that provide functions based on the output of system 600.
[0178] In various exemplary embodiments, the control signal uses, such as signaling of AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable control between devices regardless of the presence or absence of user intervention, to communicate between system 600 and display 661, speaker 671, or other peripheral devices 681.
[0179] The output devices can be communicatively coupled to system 600 via dedicated connections through corresponding interfaces 660, 670, and 680.
[0180] Optionally, the output devices can be connected to system 600 using communication channel 651 via communication interface 650. Display 661 and speaker 671 may be coupled in a single unit with other components of system 600 in an electronic device (e.g., a television).
[0181] In various exemplary embodiments, display interface 660 may include a display driver such as a timing controller (T Con) chip.
[0182] For example, if the RF portion of input terminal 690 is part of a separate set-top box, display 661 and speaker 671 may be selectively configured separately from one or more of the other components. In various exemplary embodiments where display 661 and speaker 671 may be external components, an output signal can be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output side).
[0183] In FIGS. 1-15, various methods are described, and each method includes one or more steps or operations for implementing the described method. The order and / or use of specific steps and / or operations can be changed or combined, provided that a specific order of steps or operations is not required for the exact operation of the method.
[0184] Several examples of block diagrams and / or flowcharts of operations are described. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified (one or more) logical functions. Note that in other embodiments, the (one or more) functions shown in the blocks may not be executed in the indicated order. For example, based on the related functions, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may be executed in the reverse order.
[0185] The embodiments and aspects described herein can be implemented, for example, as a method or process, an apparatus, a computer program, a data flow, a bit stream, or a signal. Even when described in the context of a single form of embodiment (e.g., only described as a method), the embodiments of the described features may be implemented in other forms (e.g., an apparatus or a computer program).
[0186] The method can be implemented, for example, within a processor, which generally refers to a processing device including a computer, a microprocessor, an integrated circuit, or a programmable logic device, etc. The processor further includes a communication device.
[0187] Also, the method can be realized by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored in a computer-readable storage medium. The computer-readable storage medium can be implemented in and have computer-readable program code executable by a computer in one or more computer-readable media in the form of a computer-readable program product 54C1. Considering its inherent ability to store information therein and retrieve information therefrom, the computer-readable storage medium used herein can be regarded as a non-transitory storage medium. The computer-readable storage medium may include, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. The following shows more specific examples of computer-readable storage media to which this exemplary embodiment is applicable, but it should be noted that these are merely examples and not an exhaustive list, as can be easily understood by those skilled in the art. The above computer-readable storage medium can be, for example, a portable computer floppy disk, a hard disk, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0188] The instructions can form an application program specifically implemented on a readable medium.
[0189] For example, instructions can exist in hardware, firmware, software, or a combination. For example, instructions can be found in an operating system, a single application, or a combination of both. Thus, a processor may be characterized as a device configured to execute a process, for example, a device including a processor-readable medium (such as a storage device) having instructions for executing the process. Also, in addition to or instead of instructions, a processor-readable medium can store data values generated by embodiments.
[0190] The apparatus can be implemented, for example, in suitable hardware, software, and firmware. Examples of such an apparatus include a personal computer, a laptop, a smartphone, a tablet, a digital multimedia set-top box, a digital television receiver, a personal video recording system, connected home appliances, a head-mounted display device (HMD, see-through glasses), a projector (beamer), a "cave" (a system including multiple displays), a server, a video encoder, a video decoder, a post-processor that processes the output from the video decoder, a pre-processor that provides an input to the video encoder, a web server, a set-top box, and any other device for processing video images, or other communication devices. Note that the apparatus can be mobile and mounted on a moving vehicle.
[0191] The computer software can be implemented by a processor 610, or by hardware, or by a combination of hardware and software. As a non-limiting example, exemplary embodiments can also be implemented with one or more integrated circuits. The memory 620 can be of any type suitable for the technical environment and can be implemented by any appropriate data storage technology (non-limiting examples include, for example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). The processor 610 can be of any type applicable to the technical environment and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0192] As will be apparent to those skilled in the art, embodiments can generate signals that are formatted, for example, to carry information that is stored or transmitted. The information can include instructions for executing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the exemplary embodiments described. This signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as baseband. The formatting can include encoding the data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. As is well known, signals can be transmitted via a variety of different wired or wireless links. The signal can be stored on a processor-readable medium.
[0193] The terms used in this specification are for the purpose of describing specific exemplary embodiments and are not intended to be limiting. Unless otherwise indicated from the context, the singular forms "a", "an", and "the" used in this specification are intended to include the plural forms as well. Additionally, the terms "include" and / or "including" used in this specification can specify the presence of the stated features, integers, steps, operations, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof. Further, when one element is referred to as "responding to", "connected to", or "associated with" another element, it may respond directly to the other element, or be associated with the other element, or intermediate elements may exist. In contrast, when one element is referred to as "responding directly to", "directly connected to", or "directly associated with" another element, no intermediate elements exist.
[0194] Note that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the symbols / terms " / ", "and / or", and "at least one of" is intended to include the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A, B, and C). As is well known to those skilled in the art, it can be extended by the number of items listed.
[0195] In this application, various numerical values can be used. Specific values may be used for illustrative purposes and each described aspect is not limited to these specific values.
[0196] Note that terms such as first, second, etc. may be used in this specification to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, unless departing from the teachings of this application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No order is implied between the first element and the second element.
[0197] References to "exemplary embodiment" or "exemplary embodiments" or "one embodiment" or "embodiments" and other variations are often used to convey that a particular feature, structure, characteristic, etc. (described in combination with an exemplary embodiment / embodiment) is included in at least one exemplary embodiment / embodiment. Thus, the appearance of the terms "in an exemplary embodiment" or "in an exemplary embodiment" or "in one embodiment" or "in an embodiment" and any variations thereof that appear throughout this application do not necessarily refer to the same exemplary embodiment.
[0198] Similarly, references to "according to an exemplary embodiment / illustration / embodiment" or "in an exemplary embodiment / illustration / embodiment" and other variations are often used to convey that a particular feature, structure, or characteristic (described in combination with an exemplary embodiment / illustration / embodiment) may be included in at least one exemplary embodiment / illustration / embodiment. Thus, the expressions "according to an exemplary embodiment / illustration / embodiment" or "in an exemplary embodiment / illustration / embodiment" that appear throughout this application do not necessarily refer to the same exemplary embodiment / illustration / embodiment, and individual or alternative exemplary embodiments / illustrations / embodiments are not necessarily mutually exclusive of other exemplary embodiments / illustrations / embodiments.
[0199] The reference numbers recited in the claims are for purposes of illustration only and do not limit the scope of the claims. Although not explicitly described, any combination or sub - combination of these exemplary embodiments / illustrations and variations can be used.
[0200] It should be understood that when the drawings are represented as flowcharts, a block diagram of the corresponding apparatus is also provided. Similarly, when the drawings are shown as block diagrams, a flowchart of the corresponding method / process is also provided.
[0201] It should be understood that in some of the drawings, arrows are included on communication paths to indicate the main direction of communication, but the communication may occur in the direction opposite to the drawn arrows.
[0202] Various embodiments relate to decoding. As used herein, "decoding" includes all or part of a process of an ultimate output that is adapted for further processing in a displayed or reconstructed video region, which is performed on a received video image (which may include a received bitstream encoding one or more video images). In various exemplary embodiments, this process includes one or more of the processes typically performed by a decoder. In various exemplary embodiments, for example, such a process can optionally include the processes performed by the decoders of the various embodiments described herein.
[0203] As a further example, in one exemplary embodiment, "decoding" refers only to inverse quantization. In one exemplary embodiment, "decoding" can refer to entropy decoding. In another exemplary embodiment, "decoding" may refer only to differential decoding. In another exemplary embodiment, "decoding" may refer to a combination of inverse quantization, entropy decoding, and differential decoding. Based on the specifically described context, it is considered that those skilled in the art will fully understand whether the term "decoding process" refers to a subset of operations or a broader decoding process.
[0204] Various embodiments relate to encoding. Similar to the above description of "decoding", "encoding" as used herein can include all or part of a process performed on an input video image to generate an output bitstream. In various exemplary embodiments, this process includes one or more of the processes typically performed by an encoder. In various exemplary embodiments, this process further includes or optionally includes the processes performed by the encoders of the various embodiments described herein.
[0205] As a further example, in an exemplary embodiment, "encoding" may refer only to quantization; in another exemplary embodiment, "encoding" may refer only to entropy encoding; in yet another exemplary embodiment, "encoding" may refer only to differential encoding; and in other exemplary embodiments, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Based on the specifically described context, it is clear whether the term "encoding process" refers specifically to a subset of operations or to a broader encoding process, and it is considered to be well understood by those skilled in the art.
[0206] Also, in this application, the "acquisition" of various information is mentioned. The acquisition of information may include, for example, any one or more of the estimation of information, the calculation of information, the prediction of information, the retrieval of information from memory, the processing of information, the transfer of information, the copying of information, the deletion of information, the calculation of information, the determination of information, the prediction of information, or the estimation of information.
[0207] Also, in this application, the "receiving" of various information is mentioned. The receiving of information may include, for example, any one or more of accessing the information or receiving the information from a communication network.
[0208] And, as used herein, the term "signal" specifically refers to indicating something to the corresponding decoder. For example, in some exemplary embodiments, the encoder signals specific information such as encoded parameters or encoded video image data. In this way, in exemplary embodiments, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicitly signal) specific parameters to the decoder, whereby the decoder can use the same specific parameters. In contrast, when the decoder has specific parameters and other parameters, it signals (implicitly signals) without transmission to inform and obtain the specific parameters by the decoder. By avoiding the transmission of actual functions, bit savings are achieved in various exemplary embodiments. It should be understood that multiple signaling methods can be achieved. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. The foregoing relates to the verb form of the term "signal", and the term "signal" may also be used as a noun in this specification.
[0209] Although many embodiments have been described, it should be understood that various modifications are possible. For example, elements of different implementations can be combined, supplemented, changed, or deleted to generate other embodiments. Further, as would be understood by those skilled in the art, the disclosed structures and processes can be replaced with other structures and processes, and the resulting embodiments perform at least substantially the same functions and achieve at least substantially the same results in at least substantially the same way as the disclosed implementations. Accordingly, these embodiments and other embodiments are contemplated by this application.
Claims
1. A method for predicting a block of a video image, comprising: A first prediction block (PB1) is obtained by an intra-template matching prediction mode (301), and the intra-template matching prediction mode determines the first prediction block (PB1) by minimizing a cost function calculated between sample values of an L-shaped template of the block of the video image and sample values of an L-shaped template of a reconstructed block of the video image. The method further includes a step of determining (302) a first block vector (BV1) as a displacement between the first prediction block (PB1) and the block of the video image, and the first block vector (BV1) identifies the first prediction block as a prediction block candidate of the block of the video image. A method for predicting a block of a video image.
2. Storing the first block vector (BV1) in a block-based motion information buffer. The method for predicting a block of a video image according to Claim 1.
3. Storing the first block vector (BV1) based on sub-blocks. The method for predicting a block of a video image according to Claim 2.
4. The method further includes a step of adding (303) the first block vector (BV1) to a list of block vector prediction candidates associated with the video image. The method for predicting a block of a video image according to any one of Claims 1 to 3.
5. Predicting the block of the video image by a second prediction block (PB2), the second prediction block (PB2) is obtained by an intra-block copy prediction mode (304), the intra-block copy prediction mode determines a second block vector (BV2) as a displacement between the block of the video image and a prediction block of a reconstructed area of the video image by block matching, and the method further includes a step of predicting (305) the second block vector (BV2) by a block vector (BV3) in the list of block vector prediction candidates. The method for predicting a block of a video image according to Claim 4.
6. The method further includes a step of storing (306) the first block vector (BV1) in a history-based motion vector prediction table. A method for predicting blocks of a video image according to any one of claims 1 to 3.
7. The motion vector prediction table based on the history further stores at least one third block vector (BV3) associated with at least one prediction block of at least one reconstructed block of the video image, and each of the at least one third block vector (BV3) is obtained by an intra block copy prediction mode, and the intra block copy prediction mode determines the third block vector (BV3) as the displacement between a prediction block in a reconstructed area of the video image and a reference block of the video image by block matching. A method for predicting blocks of a video image according to claim 6.
8. Predicting the block of the video image by a second prediction block (PB2), the second prediction block (PB2) being obtained by an intra block copy prediction mode, the intra block copy prediction mode determining a second block vector (BV2) as the displacement between the block of the video image and a prediction block in a reconstructed area of the video image by block matching, and the method further includes the step of predicting (307) the second block vector (PB2) by a block vector (BV4) in the motion vector prediction table based on the history. A method for predicting blocks of a video image according to claim 7.
9. A method for encoding blocks of a video image based on prediction blocks obtained by the method according to any one of claims 1 to 8.
10. A method for decoding blocks of a video image based on prediction blocks obtained by the method according to any one of claims 1 to 8.
11. A bitstream formatted to include encoded video image data obtained by the method according to any one of claims 1 to 10.
12. An apparatus comprising means for performing the method according to any one of claims 1 to 10.
13. A computer program product including instructions, when the program is executed by one or more processors, causing the one or more processors to perform the method according to any one of claims 1 to 10. A computer program product.
14. A non-transitory storage medium including instructions of program code for executing the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method for encoding blocks of an image and method for reconstructing blocks of an image
JP2013511874A
Method and apparatus for video coding
US20190246113A1
Intra block copy for video coding
US20190246143A1
Method and apparatus for block vector signaling and derivation in intra picture block compensation
US20200014934A1
Method and apparatus for video coding
US20200021798A1