Encoding / Decoding of Video Image Data
By adapting the IBC prediction mode with sub-pixel accuracy and bidirectional prediction for video content captured by a camera, the encoding efficiency for natural video content is enhanced, addressing the performance degradation in existing video compression standards.
Patent Information
- Application Number
- JP2024577352
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-05
- Filing Date
- 2023-04-24
- Publication Date
- 2025-07-10
- Estimated Expiration
- 2043-04-24
AI Technical Summary
The IBC prediction mode in existing video compression standards like VVC is more suitable for screen content and graphic content, leading to degraded performance when applied to video content captured by a camera.
Adapt the IBC prediction mode to predict video content captured by a camera by signaling specific syntax elements and deriving block vectors at sub-pixel accuracy, allowing bidirectional prediction and template matching, enhancing encoding efficiency for natural video content while maintaining performance for screen and graphic content.
Improves compression efficiency for video content captured by a camera by optimizing the IBC prediction mode, providing better encoding performance compared to standard VVC and ECM designs.
Smart Images

Figure 2025521852000001_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to the encoding and decoding of video images. Specifically and non-exclusively, the technical field of this application relates to the prediction of blocks of video images based on intra block copy.
Background Art
[0002] This section is intended to introduce the reader to various aspects of the art, which are related to aspects of at least one exemplary embodiment of the present application described hereinafter and / or claimed for protection. This description is believed to be useful in providing the reader with background information to better understand the various aspects of the present application. Accordingly, these techniques are not admitted as prior art and should be read from the above perspective.
[0003] In state-of-the-art video compression systems such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level image partitions are provided to divide a video image into image regions called coding tree units (CTUs). In the case of HEVC, the size of a coding tree unit (CTU) is usually between 16×16 pixels and 64×64 pixels. In the case of VVC, the size of a coding tree unit (CTU) may be 32×32, 64×64 or 128×128 pixels.
[0004] The CTU partitioning of a video image forms a grid of fixed-size CTUs, i.e., a CTU grid, whose upper and left boundaries spatially overlap with the upper and left boundaries of the video image. The CTU grid represents the spatial partitioning of the video image.
[0005] In VVC and HEVC, the CTU sizes (CTU width and CTU height) of all CTUs in the CTU grid are equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) may be 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is encoded in the bitstream, for example, encoded at the sequence level in the sequence parameter set (SPS).
[0006] The spatial position of a CTU within the CTU grid is determined based on the CTU address ctuAddr, and the CTU address ctuAddr defines the spatial position from the origin of the upper left corner of the CTU. As shown in FIG. 1, the CTU address can define the spatial position from the upper left corner of the upper-level spatial structure S containing the CTU.
[0007] An encoding tree is associated with each CTU to determine the tree partitioning of the CTU.
[0008] As shown in FIG. 1, in HEVC, the encoding tree is a quadtree partitioning of the CTU, and each leaf is called a coding unit (CU). The spatial position of a CU in the video image is defined by the CU index cuIdx, and the CU index cuIdx indicates the spatial position from the upper left corner of the CTU. A CU is spatially divided into one or more prediction units (PUs). The spatial position of a PU in the video image VP is defined by the PU index puIdx, the PU index puIdx defines the spatial position from the upper left corner of the CTU, and the spatial position of the elements of the divided PU is defined by the PU partition index puPartIdx, and the PU partition index puPartIdx defines the spatial position from the upper left corner of the PU. Some intra or inter prediction data is assigned to each PU.
[0009] The intra or inter coding mode is assigned at the CU level. This means that the same intra / inter coding mode is assigned to each PU of the CU, even though the prediction parameters may vary by PU.
[0010] Based on a quadtree called the transform tree, a CU can be spatially divided into one or more transform units (TUs). A transform unit is a leaf of the transform tree. The spatial position of a TU in a video image is defined by a TU index tuIdx, which defines the spatial position from the top-left corner of the CU. Some transform parameters are assigned to each TU. The transform type is assigned at the TU level, and a 2D independent transform is performed at the TU level during the encoding or decoding of an image block.
[0011] The PU partition types existing in HEVC are as shown in Figure 2. They include the square partitions (2N×2N and N×N), which are the only partitions used in both intra-predicted CUs and inter-predicted CUs, the symmetric non-square partitions (2N×N, N×2N, used only in inter-predicted CUs), and the asymmetric partitions (used only in inter-predicted CUs). For example, the PU type 2N×nU represents an asymmetric horizontal partition of the PU, and the small partition is located at the top of the PU. According to another example, the PU type 2N×nL represents an asymmetric horizontal partition of the PU, and the small partition is located at the top of the PU.
[0012] As shown in FIG. 3, in VVC, the coding tree starts from the root node (i.e., CTU). Then, the quadtree (or quad tree) splitting divides the root node into four nodes corresponding to four sub-blocks (solid lines) of the same size. Subsequently, the leaf of the quad tree (or quadtree) can be further divided by a so-called multi-type tree, which is related to a binary split or a ternary split of one of the four split modes shown in FIG. 4. These split types are the vertical and horizontal binary split modes denoted as SBTV and SBTH, and the vertical and horizontal ternary split modes SPTTV and STTH.
[0013] In the case of a common coding tree where the luma component and the chroma component are shared, the leaf of the coding tree of the CTU is the CU.
[0014] Contrary to HEVC, in VVC, in most cases, the CU, PU, and TU have the same size, which means that except for some specific coding modes, the coding unit is generally not divided into PUs or TUs.
[0015] FIGS. 5 and 6 provide an overview of the video encoding / decoding method used in current video standard compression systems (e.g., HEVC or VVC).
[0016] FIG. 5 shows an exemplary block diagram of the steps of a method 100 for encoding a video image VP based on the prior art.
[0017] In step 110, the video image VP is divided into sample blocks, and the division information data is signaled to the bitstream. Each block contains samples of one component of the video image VP. Therefore, these blocks contain samples of each component that defines the video image VP.
[0018] For example, in HEVC, an image is divided into coding tree units (CTUs). Each CTU can be further divided by quadtree partitioning, where each leaf of the quadtree represents a coding unit (CU). The partitioning information data may include data describing the CTUs and the quadtree subdivision of each CTU.
[0019] Thereafter, each sample block (hereinafter abbreviated as block) may be a CU (when the CU contains a single PU) or a PU of the CU.
[0020] Using an intra or inter prediction mode, each block is encoded along an encoding loop (also referred to as "within the loop").
[0021] Intra prediction (step 120) uses intra prediction data. Intra prediction predicts the current block by utilizing an intra prediction block based on samples that are already encoded, decoded, and reconstructed and are located around the current block (usually at the top and left of the current block). Intra prediction is performed in the spatial domain.
[0022] In the inter prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for a candidate reference block as a good predictor of the current block in one or more reference video images that predictively encode the current video image. For example, a good predictor of the current block is a predictor similar to the current block. The output of the motion estimation step 130 is inter prediction data that includes motion information (usually one or more motion vectors and one or more reference video image indices) associated with the current block, and other information for obtaining the same predicted block on the encoding / decoding side. Subsequently, motion compensation (step 135) utilizes the (one or more) motion vectors and (one or more) reference video image indices determined in the motion estimation step 130 to obtain a predicted block. Basically, the block belonging to the selected reference video image and pointed to by the motion vector is available as the predicted block of the current block. Also, since the motion vector is represented as a fraction of an integer pixel position (known as sub-pixel MV accuracy representation), motion compensation typically includes spatial interpolation of several reconstructed samples of the reference video image to calculate the predicted block.
[0023] The prediction information data is signaled in the bitstream. The prediction information may include a prediction mode (intra or inter or skip), intra / inter prediction data, and any other information for obtaining the same predicted CU on the decoding side.
[0024] Method 100 selects a prediction mode (intra or inter prediction mode) by optimizing the rate-distortion trade-off by considering, for example, the encoding of the prediction residual block calculated by reducing candidate prediction blocks from the current block, and the signaling of the prediction information data necessary for determining the candidate prediction block on the decoding side.
[0025] Generally, the best prediction mode is given as follows as the prediction mode of the best encoding mode p* for the current block.
[0026]
Number
[0027] Here, P is the set of all candidate encoding modes of the current block, p is a candidate encoding mode in the set, and RD cost (p) is the rate-distortion cost of the candidate encoding mode p, and is usually shown as follows. RD cost (p) = D(p) + λ.R(p)
[0028] D(p) is the distortion between the current block and the reconstructed block obtained by encoding / decoding the current block using the candidate encoding mode p, R(p) is the rate cost associated with the encoding of the current block by the encoding mode p, λ is the Lagrange parameter representing the rate constraint for encoding the current block, and is generally calculated based on the quantization parameter for encoding the current block.
[0029] Generally, the current block is encoded based on the prediction residual block PR. More precisely, for example, the prediction residual block PR is calculated by subtracting the best prediction block from the current block. Then, for example, a DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type transform or any other suitable transform is used to transform the prediction residual block PR (step 140), and the obtained transform coefficient block is quantized (step 150).
[0030] In a variant, method 100 can skip the transform step 140, skip the encoding mode by so-called transform, and directly apply quantization (step 150) to the prediction residual block PR.
[0031] Entropy-encode the quantized transform coefficient block (or quantized prediction residual block) into a bitstream (step 160).
[0032] Subsequently, as part of the encoding loop, inverse quantization (step 170) and inverse transform (180) (or not inverse-transformed) are performed on the quantized transform coefficient block (or quantized residual block) to generate a decoded prediction residual block. Then, the decoded prediction residual block and the prediction block are combined, usually summed, which provides a reconstructed block.
[0033] Furthermore, other information data may be entropy-encoded in step 160 so as to encode the current block of the video image VP.
[0034] Artifacts can be reduced by applying an in-loop filter (step 190) to the reconstructed image (including the reconstructed blocks). After all image blocks are reconstructed, the loop filter can be applied. For example, these may include a deblocking filter, sample adaptive offset (SAO), or an adaptive loop filter.
[0035] The reconstructed block or the filtered reconstructed block is formed as a reference image, which can be stored in a decoded picture buffer (DPB), whereby it can be used as a reference image for the next current block of the video image VP or for the next encoded video image to be encoded.
[0036] FIG. 6 shows an exemplary block diagram of steps of a method 200 for decoding a video image VP based on the prior art.
[0037] In step 210, by performing entropy decoding on the bitstream of the encoded video image data, split information data, prediction information data, and quantized transform coefficient blocks (or quantized residual blocks) are obtained. For example, this bitstream is generated based on method 100.
[0038] By performing entropy decoding on other information data, the current block of the video image VP can also be decoded from the bitstream.
[0039] In step 220, based on the split information, the reconstructed image is split into current blocks. Each current block is entropy decoded from the bitstream along the decoding loop (also referred to as "inside the loop"). Each decoded current block is a quantized transform coefficient block or a quantized prediction residual block.
[0040] In step 230, the current block is inverse quantized, and in some cases, inverse transformation (step 240) is performed to obtain the decoded prediction residual block.
[0041] On the other hand, the prediction information data is used to predict the current block. The prediction block is obtained by its intra prediction (step 250) or motion compensated temporal prediction (step 260). The prediction process executed on the decoding side is the same as the prediction process executed on the encoding side.
[0042] Subsequently, the decoded prediction residual block and the prediction block are merged, usually summed, which provides the reconstructed block.
[0043] In step 270, the in-loop filter is applicable to the reconstructed image (including the reconstructed blocks), and the reconstructed block or the filtered reconstructed block is formed as a reference image, which can be stored in the decoded image buffer (DPB) described above (Figure 5).
[0044] In step 130 / 135 of FIG. 5 or step 260 of FIG. 6, an inter prediction block is defined from the inter prediction data associated with the current block (CU or PU of the CU) of the video image. This inter prediction data includes motion information that can be displayed (encoded) based on the so-called AMVP mode (adaptive motion vector prediction) or the so-called merge mode.
[0045] In HEVC, in the AMVP mode, the motion information for defining the inter prediction block is represented by up to two reference video image indexes, and the up to two reference video image indexes are each associated with up to two reference video image lists (usually represented as L0 and L1). The reference video image reference list L0 includes at least one reference video image, and the reference video image reference list L1 includes at least one reference video image. Each reference video image index is a temporal prediction for the current block. The motion information further includes up to two motion vectors, and each motion vector is associated with one of the reference video image indexes of the two reference video image lists. Each motion vector is predictively encoded and signaled in the bitstream, that is, one motion vector difference MVd is derived from the motion vector, one AMVP (adaptive motion vector predictor) candidate is selected from the AMVP candidate list (constructed on both the encoding and decoding sides), and MVd is signaled in the bitstream. The AMVP candidate index of the AMVP candidate selected from the AMVP candidate list is similarly signaled in the bitstream.
[0046] FIG. 7 shows an illustrative example for constructing the AMVP candidate list, which is used to define the inter prediction block of the current block of the current video image.
[0047] The AMVP candidate list may include two spatial MVP (motion vector predictor) candidates derived from the current video picture. The first spatial MVP candidate is derived from the motion information associated with the inter prediction block of the current block and, if it exists, is located at the adjacent positions A0, A1 to the left of the current block. The second MVP candidate is derived from the motion information associated with the inter prediction block of the current block and, if it exists, is located at the adjacent positions B0, B1 and B2 above the current block. Then, a redundancy check is performed among the derived spatial MVP candidates, i.e., the overlapping derived MVP candidates are discarded. The AMVP candidate list may further include a temporal MVP candidate, which is derived from the motion information associated with the co-located block (if it exists) at the spatial position H in the reference video picture or the spatial position C in other cases. The temporal MVP candidate is scaled based on the temporal distance between the current video picture and the reference video picture. Finally, if the AMVP candidate list contains less than two MVP candidates, the AMVP candidate list is filled with zero motion vectors.
[0048] In HEVC, in merge mode, the motion information for defining an inter prediction block is represented by one merge index that merges the MVP candidate list. Each merge index points to the motion predictor information indicating which MVP is used to derive the motion information. The motion information is represented by one unidirectional or bidirectional temporal prediction type, up to two reference video picture indices, and up to two motion vectors, where each motion vector is associated with one of the two reference video picture lists (L0 or L1).
[0049] Information other than the merge index is not signaled. This means that the motion vector of the current block is equal to the motion vector of the MVP candidate indicated by the merge index. Therefore, contrary to the merge mode, the MVd and the reference picture index are not signaled in the bitstream. Only the index of the merge candidate selected from the merge candidate list is signaled in the bitstream.
[0050] Therefore, contrary to the AMVP mode, in the merge mode, the MVd and the reference picture are not signaled. Only the merge index is signaled in the bitstream.
[0051] The merge MVP candidate list may include five spatial MVP candidates, which are derived from the current video image shown in FIG. 7. The first spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the left adjacent position A1. The second spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the upper adjacent position B1. The third spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the upper-right adjacent position B0. The fourth spatial MVP candidate is derived from the motion information associated with the inter-prediction block and, if it exists, is located at the lower-left adjacent position A0. The fifth spatial MVP candidate is derived from the above motion associated with the inter-prediction block and, if it exists, is at the left adjacent position B2. Then, a redundancy check is performed among the derived spatial MVPs, i.e., duplicate derived MVP candidates are discarded. The merge candidate list may further include a temporal MVP candidate called TMVP candidate, which is derived from the motion information associated with a collocated block (if it exists) located at position H in the reference video image or the central spatial position "C". Then, a redundancy check is performed among the derived spatial MVPs, i.e., duplicate derived MVP candidates are discarded. Finally, when using the bidirectional temporal prediction type, if the merge candidate list contains less than five MVP candidates, merge candidates are added to the merge candidate list. The merge candidate is derived from the motion information of one MVP candidate associated with one reference video image list and existing in the merge candidate list, and the motion information corresponds to another MVP candidate associated with another reference video image list and existing in the merge candidate list. Finally, if the merge candidate list is still not (filled with five merge candidates), the merge candidate list is filled with zero motion vectors.
[0052] In ECM (Explanation of the Algorithm of Extended Compression Model 4 (ECM 4), M. Coban, F. Le Leannec, K. Naser, J. Strom, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, 23rd meeting, teleconference, July 7 - 16, 2021, Document JVET-Y202-v2, https: / / jvet-experts.org / doc_end_user / documents / 25_Teleconference / wg11 / JVET-Y2025-v2.zip), the template matching (TM) method may be used to determine some of the MVP candidates in the AMVP and merge modes. The TM method is a method for subdividing motion vectors on the decoder side. As shown in Figure 8, it subdivides the motion vector of a block by matching the template regions above and to the left of the current block with the templates within the search region of the reference image. A more appropriate MV is searched around the initial MV of the current block, for example, within the [-8, +8]-pixel search range. The subdivided MV is obtained by minimizing the so-called template matching cost TMcost between the template around the current block and the candidate templates in the reference video image. The MVP candidate with the lowest template matching cost is selected and further subdivided.
[0053] When encoding motion information according to VVC, more abundant motion information can be represented than in HEVC encoding.
[0054] In VVC, motion representation can be provided based on the AMVP mode or the merge mode.
[0055] In VVC, in the AMVP mode, the motion information for defining an inter prediction block is represented in the same way as in the HEVC AMVP mode. If present, the AMVP candidate list may include spatial and temporal MVP candidates as in HEVC. If present, the AMVP candidate list may further include four additional HMVP (history-based motion vector prediction) candidates. Finally, if the AMVP candidate list contains fewer than two MVP candidates, the AMVP candidate list is filled with zero motion vectors.
[0056] HMVP candidates are derived from previously encoded MVPs of adjacent or non-adjacent blocks associated with the current block. Therefore, an HMVP candidate table is maintained on both the encoder and decoder sides and updated in real time as a first-in-first-out (FIFO) buffer for MVPs. There are up to five HMVP candidates in the table. After encoding one block, the associated motion information is added to the end of the table to update the table as a new HMVP candidate. The table is managed by applying the FIFO rule, and in addition to the basic FIFO mechanism, redundant candidates in the HMVP table are deleted first instead of the first candidate. The table is reset for each CTU row to enable parallel processing.
[0057] In the AMVP mode, the AMVR (adaptive motion vector resolution) algorithm is adopted. The AMVR tool can signal MVd at a luminance sample resolution of 1 / 4 pixel, 1 / 2 pixel, integer pixel, or 4 pixels. This can save bits when encoding MVd information. In AMVR, the resolution of the motion vector is selected at the block level.
[0058] Finally, the internal motion vector representation is realized with a luminance sample accuracy of 1 / 16 instead of 1 / 4 luminance sample accuracy in HEVC.
[0059] In VVC, in the merge mode, the motion information for defining an inter prediction block is represented in the same way as in the HEVC merge mode.
[0060] The merge candidate list is different from the merge candidate list used in HEVC.
[0061] In VVC, both the MMVD (Merge Mode with Motion Vector Difference) mode and the CIIP (Composite Intra / Inter Prediction) mode can establish a merge candidate list.
[0062] The MMVD mode can represent the motion information associated with an inter prediction block by encoding a limited motion vector difference (MVd) in the selected merge candidate. MMVD encoding is limited to four vector directions, eight amplitude values, and from 1 / 4 luminance samples to 32 luminance samples, as shown in FIG. 9. The MMVD mode can signal motion information by making an intermediate trade-off between the rate cost and the MV (motion vector) accuracy in order to provide an intermediate accuracy level. The MMVD offset can also be signaled in the bitstream. Then, the motion vector is derived by adding the MV predictor to the MMVD offset.
[0063] CIIP combines an inter prediction signal and an intra prediction signal to predict the current block of a video image. The inter prediction signal in the CIIP mode is derived using a prediction process similar to that applied to the merge mode. The intra prediction signal is derived according to a normal intra prediction process having a Planar mode. The Planar prediction mode predicts a block through a spatial interpolation process of adjacent reconstructed samples of the prediction blocks at the top and left of the block. Then, the intra and inter prediction signals are blended (composited) using a weighted average, where the calculation of the weighting value is based on the encoding modes of the adjacent blocks at the top and left as follows. PCIIP = (Wmerge × Pmerge + Wintra × Pintra + 2) >> 2 Weight W merge and W intra The sum of is 4, and these weights are constant throughout the current block.
[0064] In short, in VVC, the merge candidate list may contain the same spatial MVP candidates as HEVC, and only the two first candidates are swapped. During the construction of the merge candidate list, candidate B1 is considered before candidate A1. The merge candidate list may further include TMVP, for example, the HEVC, HMVP candidates in the VVC AMVP mode. Some HMVP candidates are inserted into the merge candidate list so that the merge candidate list reaches the maximum allowable number (-1) of MVP candidates. The merge candidate list may further include an average candidate that results in at most one pair calculated as follows. Consider the two first merge candidates present in the merge candidate list and average their motion vectors. This average value is calculated separately for each reference video picture list L0 and L1. Therefore, if both MVPs are bi-directional, the motion vectors associated with the two lists L0 and L1 are averaged. If there is only one motion vector in the reference video picture list, it is taken as it is and a paired candidate is formed. Finally, if the merge candidate list has not reached the maximum allowable number of MVP candidates, the merge candidate list is filled with zero motion vectors.
[0065] In HEVC and VVC, screen content is encoded using the Intra Block Copy (IBC) prediction mode. The IBC prediction mode is known to significantly improve the encoding efficiency of screen content materials. Since the IBC prediction mode is realized as a block-level encoding mode, block matching (BM) is performed in the encoder to find the optimal block vector (or motion vector) for each block to be predicted. The block vector indicates the displacement from the block to be predicted in the video image to the reference block, and the reference block has already been reconstructed (decoded) within the video image. The block vector of the IBC prediction block for luminance samples has integer precision. The block vector of the IBC prediction block for chrominance samples can also be rounded to integer precision. When merged with AMVR, the IBC prediction mode can switch between 1-pixel and 4-pixel motion vector precisions. The IBC prediction mode is regarded as a third prediction mode other than the intra or inter prediction mode.
[0066] At the CU level, the IBC prediction mode is signaled by a flag indicating that the IBC-AMVP mode or the IBC-skip / merge mode is being used.
[0067] In the IBC-skip / merge mode, the block vector for defining the IBC prediction block is indicated by one merge index in the merge candidate list. The merge index indicates the block vector in the merge candidate list. The merge candidate list may include spatial candidates, HMVP candidates, and paired candidates. The acquisition of the HMVP (History-based Motion Vector Prediction) candidate block vector predictor is the same as in the case of conventional inter-merge encoding. HMVP includes a motion vector predictor buffer, and when an encoding unit encodes in the IBC mode for a given image, the buffer is provided. The HMVP buffer further includes a block vector buffer for encoding or decoding in the IBC mode by providing candidates for predicting the block vector of the current encoding unit.
[0068] By averaging two IBC merge candidates, a new pair of IBC merge candidates can be formed. This means averaging two first block vector candidates in the constructed merge list to form a block vector prediction candidate for the block to be predicted, which is a so-called pair of blocks. This pair of candidates is added to the IBC merge candidate list and is located after the HMVP candidates.
[0069] For HMVP, the block vector is inserted into the history buffer for future reference.
[0070] In the IBC-AMVP mode, two block vector predictors are determined. One is from the adjacent area on the left side of the block to be predicted, and the other is from the adjacent area above the block to be predicted (when IBC is predicted). If either adjacent area is not available, the default block vector predictor is considered. The block vector predictor index is indicated by signaling a flag. In fact, since at most two block vector prediction candidates are considered in the IBC AMVP mode, the flag is sufficient to identify the block vector prediction candidate used for encoding the block vector information of the coding unit. The block vector difference is encoded in a similar way to the motion vector difference of the inter-coding unit.
[0071] Also, in ECM, the IBC-TM merge mode is defined as an IBC prediction mode combined with template matching. The IBC-TM merge list related to the IBC-TM merge mode is different from the list used for the IBC skip / merge mode. In the normal TM merge mode, candidates are selected by pruning based on the moving distance between candidates. The zero-motion candidates are replaced with (-W, 0), (0, -H), (-W, -H) MVs.
[0072] In the IBC-TM merge mode, the selected candidates are subdivided using the template matching method. Signaling the TM-merge flag indicates whether to use the template matching merge IBC mode or not.
[0073] In the IBC-TM AMVP mode, up to three candidates are selected from the IBC-TM merge list. The IBC-TM merge list is a list of block vector prediction candidates for predicting the block vector of the block to be predicted. It is structured by the blocks around the current block to be predicted, which have already been encoded and / or decoded in the IBC coding mode. Each of these candidates is subdivided according to the normal template matching method, and these candidates are sorted according to the obtained TM cost. The TM cost is calculated as the distortion between the top and left template regions of the current block to be predicted (in its decoded and reconstructed state) and the top and left corresponding to the surroundings of the candidate predicted block pointed by the candidate block vector.
[0074] When used in the IBC prediction mode, TM subdivision is performed at integer pixel positions, and in the IBC-TM AMVP mode, TM subdivision is performed at integer precision or 4-pixel precision based on the AMVR value. The subdivision is performed within the existing IBC reference region.
[0075] The IBC prediction mode cannot be used in combination with the inter-prediction tools of VVC such as CIIP and MMVD.
[0076] In VVC, different from the HEVC screen content coding extension, the current video image is not included in the reference image list 0 for IBC prediction as one of the reference images. The block vector derivation process in the IBC mode excludes all adjacent blocks in the inter mode, and vice versa.
[0077] In VVC, the IBC prediction mode shares the same process as the normal MV merge mode that includes a pair of merge candidates and a motion predictor based on history, but the TMVP and zero vector are not used because they are invalid for the IBC prediction mode.
[0078] Separate HMVP buffers (5 candidates per buffer) are used for conventional MVs and the IBC prediction mode.
[0079] On the other hand, in the current design of the IBC prediction tool, it has been found that the IBC design is more suitable for graphic and screen video content than natural video content in various aspects.
[0080] This limits the improvement in compression performance obtained when using the IBC prediction mode for video content captured by a camera.
[0081] An object of the present invention is to adapt the canonical design of the IBC encoding mode and significantly improve its performance when compressing natural video while maintaining at least the encoding performance for screen content.
[0082] The problem to be solved by the present invention is to adapt the IBC prediction mode defined in VVC to improve the compression efficiency when enabling the IBC prediction mode in video content captured by a camera, and at the same time maintain the performance of the current IBC prediction mode in at least screen content and graphic content.
[0083] For the problem to be solved, a conventional solution is to include the currently encoded or decoded video image in a reference image buffer (so-called decoded picture buffer, DPB) and use it for temporal prediction of blocks in the current video image.
[0084] Therefore, compared with VVC, the performance when encoding natural sequences in the IBC prediction mode may be improved.
[0085] However, since the IBC prediction mode adapts to motion representation and compensation for natural content and not for screen content, performance may degrade in screen content and graphic content.
[0086] In view of the above situation, at least one exemplary embodiment of the present application is designed. SUMMARY OF THE INVENTION
[0087] The following section provides a basic understanding of some aspects of the present application by showing a simplified summary of at least one exemplary embodiment. This summary is not a detailed summary of the exemplary embodiment. It does not identify critical or important elements of the exemplary embodiment. The following summary merely shows some aspects of at least one exemplary embodiment in a simplified form as a prelude to the more detailed description provided elsewhere in the document.
[0088] According to a first aspect of the present application, a method for predicting a block of a video image is provided, the method comprising: signaling a type of an intra block copy mode, the type of the intra block copy mode indicating whether the intra block copy prediction mode is set to predict video content captured by a camera; determining at least one block vector for predicting a block of the video image from at least one reference block of the video image in the intra block copy prediction mode; setting the intra block copy prediction mode to predict a block of the video image when the type of the intra block copy mode indicates that the intra block copy prediction mode is set to predict video content captured by a camera; and deriving a predicted block of the block of the video image based on the set intra block copy mode prediction mode.
[0089] In an exemplary embodiment, a block belongs to a slice of a video image, and the type of the incoming intra block copy mode is signaled at the slice level.
[0090] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera further includes signaling a syntax element (ibc_pred_idc_flag) that indicates whether bidirectional prediction is permitted, and signaling two block vectors when the syntax element indicates that bidirectional prediction is permitted, wherein the prediction of a block of the video image is the average value of the two block vectors.
[0091] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera includes signaling a syntax element (mvp_l1_flag) to identify a block vector predictor of each block vector in a block vector list associated with a decoded block adjacent to a block of the video image, further including that the adjacent block belongs to the video image.
[0092] In an exemplary embodiment, at least one block vector is represented at the sub-pixel accuracy level.
[0093] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera further includes signaling at least one block vector difference calculated between a block vector and a reference block within the video image, wherein each block vector difference is signaled at the sub-pixel accuracy level.
[0094] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera further includes encoding each block vector difference, wherein the encoding is limited to several vector directions and widths.
[0095] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera further includes signaling (step 310) a syntax element (mmvd_merge_flag) that indicates whether the encoding of the block vector difference is limited to several vector directions and widths.
[0096] In an exemplary embodiment, setting the intra block copy prediction mode to predict video content captured by a camera further includes signaling a syntax element (intra_tmp_flag) to indicate whether to derive a predicted block of a block of a video image using intra block copy prediction based on template matching, or whether to derive a predicted block of a block of a video image based on a first aspect.
[0097] In an exemplary embodiment, the step of deriving a predicted block of a block of a video image based on the set intra block copy mode prediction mode includes deriving a first predicted block of a block of the video image from the set intra block copy mode, deriving a second predicted block of a block of the video image, and blending the first predicted block and the second predicted block to derive a predicted block of a block of the video image.
[0098] In an exemplary embodiment, the second predicted block is derived based on motion compensation prediction or intra prediction of a block of the video image.
[0099] According to a second aspect of the present application, a method for encoding a block of a video image based on a predicted block derived from the method of the first aspect is provided.
[0100] According to a third aspect of the present application, there is provided a method for decoding a block of a video image based on a prediction block derived from the method of the first aspect.
[0101] According to a fourth aspect of the present application, there is provided a bitstream formatted to include encoded video image data obtained from the method of the first aspect.
[0102] According to a fifth aspect of the present application, there is provided an apparatus including means for performing one of the methods of the first, second, and / or third aspects.
[0103] According to a sixth aspect of the present application, there is provided a computer program product including instructions for causing one or more processors to execute the method according to the first, second, and / or third aspects when the program is executed by the one or more processors.
[0104] According to a seventh aspect of the present application, there is provided a non-transitory storage medium carrying instructions of program code for performing the method according to the first, second, and / or third aspects.
[0105] The specific nature of at least one embodiment in the exemplary embodiments and other objects, advantages, features, and uses of the at least one embodiment in the exemplary embodiments will become apparent from the description given with reference to the examples in combination with the following drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0106] Here, by way of example, reference is made to the drawings of the exemplary embodiments of the application.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
[0107] Similar or identical parts are represented by the same reference numerals.
DETAILED DESCRIPTION OF THE INVENTION
[0108] Hereinafter, at least one exemplary embodiment will be described more comprehensively with reference to the drawings, where an example of at least one exemplary embodiment of the exemplary embodiment is described. However, the exemplary embodiments can be implemented in many alternative forms and should not be construed as limited to the examples described herein. Therefore, it should be understood that there is no intention to limit the exemplary embodiments to the specific forms disclosed. Rather, this application intends to cover all modifications, equivalent substitutions, and alternatives within the spirit and scope of this application.
[0109] At least one of these aspects generally relates to the encoding and decoding of video images, another aspect generally relates to the transmission of the bitstream to be provided or encoded, and one of the other aspects relates to the reception / access of the decoded bitstream.
[0110] Although at least one exemplary embodiment of these exemplary embodiments is described to encode / decrypt video images, since each video image is sequentially encoded / decrypted as described below, it is also extended to the encoding / decryption of video images (image sequences).
[0111] Also, at least one exemplary embodiment is not limited to the current version of VVC. The at least one exemplary embodiment can be applied to existing or future-developed VVC and extensions of recommendations. Unless otherwise specified or technically excluded, each aspect described in this application can be used individually or in combination.
[0112] A pixel corresponds to the smallest display unit on the screen, which can be composed of one or more light sources (one for a monochrome screen and three or more for a color screen).
[0113] A video image, also referred to as a frame or an image frame, includes at least one component (also called an image component or channel) determined by a specific image / video format, and the specific image / video format specifies all information associated with pixel values and all information for displaying and / or decrypting video image data related to the video image by being used by a display unit and / or any other device.
[0114] A video image usually includes at least one component represented in the form of an array of samples.
[0115] A monochrome video image contains a single component, and a color video image may contain three components.
[0116] For example, if the image / video format is the well-known (Y, Cb, Cr) format, a color video image may contain a luminance (or luminosity) component and two chrominance components, or if the image / video format is the well-known (R, G, B) format, a color video image may contain three color components (one for red, one for green, and one for blue).
[0117] Each component of a video image may contain multiple samples with respect to the number of pixels on the screen where the video image is displayed. In a variant, the number of samples included in a component may be a multiple (or fraction) of the number of samples included in other components of the same video image.
[0118] For example, if the video format contains a luminance component and two chrominance components (e.g., the (Y, Cb, Cr) format), depending on the color format considered, the chrominance components may contain half the number of samples in terms of width and / or height compared to the luminance component.
[0119] A sample is the smallest visual information unit of a component that makes up a video image. The sample value may be, for example, a luminance or chrominance value or a color value in the (R, G, B) format.
[0120] A pixel value is the value of a screen pixel. For a monochrome video image, the pixel value can be represented by one sample, and for a color video image, the pixel value can be represented by multiple collocated samples. Collocated samples associated with a pixel mean samples corresponding to the position of the pixel on the screen.
[0121] Typically, a video image is regarded as a set of pixel values, and each pixel is represented by at least one sample.
[0122] A block of a video image is a set of samples of one component of the video image. When the image / video format is the well-known (Y, Cb, Cr) format, a block of at least one luminance sample or a block of at least one chrominance sample is considered, or when the image / video format is the well-known (R, G, B) format, a block of at least one color sample is considered.
[0123] At least one exemplary embodiment is not limited to a specific image / video format.
[0124] Generally, the present application relates to a method for a block of a video image, the method comprising the step of signaling a type of an intra-block copy mode, wherein the type of the intra-block copy mode indicates whether the intra-block copy prediction mode is set to predict video content captured by a camera, the step of determining a block vector for predicting a block of the video image from a reference block of the video image in the intra-block copy prediction mode, and when the type of the intra-block copy mode indicates that the intra-block copy prediction mode is set to predict video content captured by a camera, the step of setting the intra-block copy prediction mode to predict the block of the video image, and the step of deriving a predicted block of the block of the video image based on the set intra-block copy mode prediction mode.
[0125] According to the present invention, the intra-block copy (IBC) prediction mode is more suitable for predicting video content captured by a camera than the IBC prediction modes specified in VVC and ECM. The present invention can perform switching of video encoding / decoding between different settings of the IBC prediction mode, where one is used for predicting screen / graphic content and the other is used for predicting video content captured by a camera. The encoding efficiency of a block of a video image related to predictive encoding of a block of a video image of a predicted block derived from the present invention is compared with the encoding efficiency of the block of the IBC prediction mode specified in VVC and ECM.
[0126] Other aspects and advantages of the present application will be described in the following specific exemplary embodiments with reference to the accompanying drawings.
[0127] FIG. 10 is a block diagram of a method 300 for predicting a block of a video image according to an exemplary embodiment.
[0128] Method 300 provides a predicted block of a block of a video image. This predicted block can be used for the prediction mode selection of method 100 or the prediction mode of the prediction process of method 200.
[0129] Signaling prediction information related to method 300, that is, when method 300 is used in method 100, writing the prediction information into the bitstream, and when method 300 is used in method 200, parsing the prediction information from the bitstream.
[0130] In step 301, the type of the intra-block copy mode (abbreviated as the IBC mode type) is signaled as prediction information. The IBC mode type indicates whether the IBC prediction mode is set to predict video content captured by a camera. The IBC prediction mode is used to determine at least one block vector for predicting a block of a video image from at least one (reconstructed) reference block of the video image.
[0131] In one exemplary embodiment of step 301, the IBC mode type is signaled at the slice level.
[0132] In a variant, the IBC mode type is signaled only when the IBC prediction mode is enabled.
[0133] Table 1 shows an example of signaling the IBC prediction mode type as the flag sh_ibc_camera_content_flag within the slice header syntax defined by VVC based on an exemplary embodiment.
[0134] [Table 1]
[0135] When the flag sh_ibc_camera_content_flag is false (=0), the blocks of the video image using the IBC prediction mode in the corresponding slice are encoded for screen content using the IBC modes defined in VVC and ECM.
[0136] When the flag sh_ibc_camera_content_flag is true (=1), the blocks of the video image using the IBC prediction mode in the corresponding slice use the prediction blocks derived from step 302 of method 300 for the video content captured by the camera. As described above, step 302 adapts (configures) the IBC prediction mode defined in VVC to improve the prediction of the video content captured by the camera, thereby improving the efficiency of the predictive encoding of the video content captured by the camera.
[0137] When the IBC mode type indicates that the IBC prediction mode is set to predict video content captured by a camera, in step 302, the IBC prediction mode is set to predict video content captured by the camera, and in step 303, a predicted block is derived by predicting blocks of the video image based on the set IBC prediction mode.
[0138] In a first exemplary embodiment of step 302, setting the intra block copy prediction mode to predict video content captured by a camera (step 302) further includes signaling the syntax element ibc_pred_idc_flag as prediction information (step 304) to indicate whether bidirectional prediction is permitted.
[0139] In a variation of the first exemplary embodiment of step 302, when the syntax element ibc_pred_idc_flag indicates that bidirectional prediction is permitted, setting the intra block copy prediction mode to predict video content captured by a camera (302) further includes signaling two block vectors as prediction information (305), and the predicted block of the block of the video image may be the average value of these two block vectors (step 303).
[0140] In a second exemplary embodiment of step 302, setting the intra block copy prediction mode to predict video content captured by a camera (302) further includes signaling the syntax element mvp_l1_flag as prediction information (306) to identify the block vector predictor of each block vector of the IBC prediction mode in a block vector list associated with (decoded) blocks adjacent to the block of the video image, and the adjacent (decoded) blocks belong to the video image.
[0141] Table 2 shows the syntax of VVC modified based on the first and second exemplary embodiments and variations of step 302 related to blocks of video images.
[0142] [Table 2]
[0143] In the third exemplary embodiment of step 302, at least one block vector determined by the IBC prediction mode can be shown using a sub-pixel resolution level such as AMVR in the inter prediction mode of VVC / ECM. Therefore, a spatial interpolation filter similar to the spatial interpolation filter of the inter prediction mode of VVC applied to motion vector compensation can be applied to the at least one block vector.
[0144] In a variation of the third exemplary embodiment of step 302, setting (302) the intra block copy prediction mode to predict video content captured by a camera further includes signaling (step 307) a sub-pixel resolution level as prediction information.
[0145] In the fourth exemplary embodiment of step 302, setting (302) the intra block copy prediction mode to predict video content captured by a camera further includes signaling (step 308) at least one block vector difference as prediction information, where at least one block vector difference is calculated between the block vector determined by the IBC prediction mode and a reference block within the video image.
[0146] In a variation of the fourth exemplary embodiment of step 302, each block vector difference can be signaled at the sub-pixel accuracy level, and a spatial interpolation filter similar to the spatial interpolation filter of the inter prediction mode of VVC applied to motion vector compensation can be applied.
[0147] In a variation of the fourth exemplary embodiment of step 302, setting (302) the intra block copy prediction mode to predict video content captured by a camera further includes encoding (step 309) a block vector difference, wherein the encoding is limited to several vector directions and widths.
[0148] For example, the block vector difference is limited to four vector directions and eight width values from 1 / 4 luminance samples to 32 luminance samples, similar to the motion vector difference encoded according to MMVD. The block vector offset can be applied to the block vector predictor to reconstruct the block vector in the IBC prediction mode, which is the same as the MMVD defined in VVC being applied to the motion vector predictor to reconstruct the motion vector.
[0149] In a variation of the fourth exemplary embodiment of step 302, when the merge mode is enabled, setting (302) the intra block copy prediction mode to predict video content captured by a camera further includes signaling (step 310) the syntax element mmvd_merge_flag as prediction information to indicate whether the encoding of the block vector difference is limited to several vector directions and widths.
[0150] In a variation of the fourth exemplary embodiment of step 302, setting the intra block copy prediction mode (302) to predict video content captured by a camera further includes signaling (step 311) a syntax element mmvd_cand_flag to indicate a merge candidate block vector index in the mmvd merge candidate list, and the index is used to derive the block vector of a block of a video image to be predicted. The mmvd merge candidate list is basically a subset of the normal merge candidate list and consists of only two elements. Basically, the first two merge candidates of the normal merge candidate list are used to form the mmvd merge candidate list. In practice, a single flag mmvd_cand_flag is required to indicate which id of the two mmvd merge candidates is used to derive the motion information of the current block.
[0151] In a variation of the fourth exemplary embodiment of step 302, setting the intra block copy prediction mode (302) to predict video content captured by a camera further includes signaling (step 312) two syntax elements mmvd_distance_idx and mmvd_direction_idx to indicate a block vector offset associated with the block vector difference.
[0152] Table 3 is an example of the VVC syntax modified based on the fourth exemplary embodiment of step 302 and its variations related to merge data.
[0153]
Table 3
[0154] FIG. 11 shows an example of a method 400 for predicting a block of a video image based on the first, second, and fourth exemplary embodiments of step 302 and their variations described above.
[0155] An exemplary embodiment of step 302 obtains prediction information. The prediction information can be determined in the encoding method 100 and analyzed from the bitstream in the decoding method 200.
[0156] Based on the exemplary embodiment and variations of step 302, the prediction information may include syntax elements ibc_pred_idc_flag, mvp_l1_flag, mmvd_merge_flag, mmvd_cand_flag, mmvd_distance_idxmmvd_direction_idx, at least a block vector difference, and / or a sub-pixel resolution level.
[0157] In step 401, the method checks whether the merge mode is enabled.
[0158]
Number
[0159]
Number
[0160]
Number
[0161] In step 404, method 400 checks whether the syntax element ibc_pred_idc_flag permits bidirectional prediction.
[0162] If permitted, after step 404, step 412 follows.
[0163] If not permitted, after step 404, steps 405 and 406 follow.
[0164]
Number
[0165]
Number
[0166] After step 406 is step 412.
[0167]
Number
[0168] In step 407, method 400 checks whether the syntax element mmvd_merge_flag indicates that the encoding (decoding) of the block vector difference is limited to several vector directions and widths.
[0169]
Number
[0170]
Number
[0171] In step 409, the merge candidate block vector identified by the syntax element mmvd_cand_flag is obtained from the mmvd merge candidate list.
[0172]
Number
[0173]
Number
[0174] After step 411 is step 412.
[0175]
Number
[0176]
Number
[0177]
Number
[0178] In the modification example, each block vector difference is encoded / decoded based on the syntax element mmvd_merge_flag. The encoding is limited to several vector directions and widths.
[0179] In the fifth exemplary embodiment of step 302, setting (302) the intra-block copy prediction mode to predict the video content captured by the camera further includes signaling the syntax element intra_tmp_flag (step 313) to indicate whether to use intra-block copy prediction (IBCTMP) based on template matching to derive the prediction block of the block of the video image.
[0180] Basically, IBCTMP is similar to the IBC prediction mode, but the reference region for predicting the block of the video image is searched based on template matching (TM) instead of being signaled through block vectors as in the IBC prediction mode.
[0181] When the syntax element intra_tmp_flag is 1, the prediction block of the video image is derived using IBCTMP. Therefore, no information related to the block vector is encoded in the bitstream. Instead, it is searched synchronously in both the encoder and the decoder.
[0182] This search is based on template matching, similar to that performed by the template matching process dedicated to screen content. Also, since IBCTMP is used for content captured by a camera, in template matching, not only the integer position of the reference block but also the fractional spatial position between luminance samples is searched, similar to that performed in the template matching of inter prediction.
[0183] When the syntax element intra_tmp_flag is 0, as described above, the IBC prediction mode set according to one of the examples and variations of step 302 is used for the prediction block of the video image.
[0184] Table 4 shows the syntax structure modified according to the second and sixth exemplary embodiments and their variations based on the syntax structure of Table 2.
[0185]
Table 4
[0186] In an exemplary embodiment of step 303, method 300 further includes steps of deriving an IBC prediction block of a block of the video image from the set IBC prediction mode (step 314), deriving another prediction block of the block of the video image (315), and deriving a prediction block of the block of the video image (316) by blending the first and second prediction blocks.
[0187] In a variation of the above exemplary embodiment of step 303, the another prediction block is derived from motion compensation prediction of the block of the video image.
[0188] The advantage of this exemplary embodiment is to provide a means of bidirectional prediction for executing a block of a video image based on a reference block of a reference image of the current video image and a reconstructed block region of the current video image. This prediction mode further improves the coding efficiency based on the prior art.
[0189] In a modification of the exemplary embodiment of step 303, the other prediction block is derived from an intra prediction of a block of the video image.
[0190] For example, the intra prediction may be DC, planar, direct mode, horizontal and vertical angle prediction, ECM's DIMD (Decoder-side Intra Mode Derivation) or TIMD (Template-based Intra Mode Derivation) of VVC. These coding modes are described in the ECM file ("Description of the Algorithm of Extended Compression Model 4 (ECM4)", M. Coban, F. Le Leannec, K. Naser, J. Strom, ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29's Joint Video Experts Team (JVET), 23rd meeting, teleconference, July 7-16, 2021, Document JVET-Y202-v2, https: / / jvet-experts.org / doc_end_user / documents / 25_Teleconference / wg11 / JVET-Y2025-v2.zip).
[0191] DC is an intra prediction mode, which generates a prediction using the average sample value of the reference samples on the left and above the block.
[0192] The planar mode is the weighted average value of four reference sample values (obtained as the orthogonal projection of the samples and predicted in the reconstructed regions at the top and left).
[0193] The direct mode (DM) is a chroma intra prediction mode and corresponds to the intra prediction mode of the collocated reconstructed luminance samples.
[0194] Each of the horizontal and vertical modes predicts the sample rows and columns respectively by using a copy of the reconstructed samples on the left and above without interpolation.
[0195] In a variant of the exemplary embodiment of step 303, the blend may be the same as the CIIP mode.
[0196] Since the intra predictor is an IBC predictor optimized for video content captured by a prediction camera, the exemplary embodiment of step 303 improves the coding efficiency compared to the CIIP mode in VVC.
[0197] In one variant of the exemplary embodiment of step 303, method 300 further includes a step of signaling (317) the syntax element regular_merge_flag to the bitstream. The syntax element regular_merge_flag being 1 indicates that the prediction block of the block of the video image is derived based on one or a variant of the exemplary embodiment of step 302. The syntax element regular_merge_flag being 0 indicates that the IBC prediction block is merged with the other prediction block of the video image.
[0198] In a variant of the exemplary embodiment of step 303, the syntax element regular_merge_flag is added to Table 5, and Table 5 is an example of a VVC syntax related to merge data and modified based on the sixth exemplary embodiment of step 302.
[0199] Table 5 shows the structure of the syntax modified based on Table 3 and based on the exemplary embodiment and variants of step 303.
[0200]
Table 5
[0201] In a variant of step 303, when the IBC prediction block is determined, the block vector associated with the block of the video image is set as the first merge candidate in the merge candidate list (similar to the IBC skip / merge mode).
[0202] In another variation of step 303, a merge index can be signaled, and a block vector (similar to the IBC skip / merge mode) can be derived based on the signaled merge index.
[0203] In another variation of step 303, when the syntax element ibc_pred_idc_flag indicates that bi - directional prediction is permitted, two block vectors are signaled to identify a specific block in the block of the video picture to be predicted, and then the IBC prediction block may be the average value of these two block vectors.
[0204] FIG. 12 is an exemplary schematic block diagram of a system 600 that implements various aspects and exemplary embodiments.
[0205] System 600 may be incorporated as one or more devices including various components described below. In various exemplary embodiments, system 600 may be configured to implement one or more aspects described in the present application.
[0206] Examples of devices that make up all or part of system 600 may include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles, and related processing systems, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process the output from video decoders, pre-processors that provide input to video encoders, web servers, video servers (e.g., broadcast servers, video broadcast servers or network servers), still or video cameras, encoding or decoding chips, or any other communication device. The elements of system 600 can be implemented singly or in combination within a single integrated circuit (IC), multiple ICs and / or individual components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of system 600 can be distributed across multiple ICs and / or individual components. In various exemplary embodiments, system 600 can be communicatively coupled to other similar systems or electronic devices, for example via a communication bus or dedicated input and / or output ports.
[0207] System 600 may include at least one processor 610, and the at least one processor 610 is configured to execute instructions loaded to implement each aspect described in the present application. The processor 610 may include an embedded memory, an input / output interface, and various other circuits known in the art. System 600 may include at least one memory 620 (e.g., a volatile memory device and / or a non-volatile memory device). System 600 may include a storage device 640 including non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. By way of non-limiting example, the storage device 640 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0208] System 600 may include an encoder / decoder module 630 configured to provide encoded / decoded video image data, for example, by processing data. And the encoder / decoder module 630 may include its own processor and memory. The encoder / decoder module 630 may represent one or more modules that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding and a decoding module. Also, the encoder / decoder module 630 may be implemented as a separate element of the system 600 or may be incorporated within the processor 610 as a combination of hardware and software known to those skilled in the art.
[0209] The program code loaded into the processor 610 or the encoder / decoder 630 to execute each aspect described in this application is stored in the storage device 640, loaded into the memory 620, and executed by the processor 610. According to various exemplary embodiments, during the execution of the processes described in this application, one or more of the processor 610, the memory 620, the storage device 640, and the encoder / decoder module 630 can store one or more of various items. Such stored items include, but are not limited to, video image data, information data used for encoding / decoding video image data, bitstreams, matrices, variables, equations, mathematical formulas, operations, and intermediate or final results of operation logic processing.
[0210] In some exemplary embodiments, the memory inside the processor 610 and / or the encoder / decoder module 630 may be used to store instructions and provide a working memory for the processes executed during encoding or decoding.
[0211] However, in other exemplary embodiments, a memory external to the processing device (for example, the processing device may be the processor 610 or the encoder / decoder module 630) is used for one or more of these functions. The external memory may be the memory 620 and / or the storage device 640, and may be, for example, dynamic volatile memory and / or non-volatile flash memory. In some exemplary embodiments, the external non-volatile flash memory may be used to store the operating system of the television. In at least one exemplary embodiment, a high-speed external dynamic volatile memory such as RAM can be utilized as a working memory for video encoding and decoding operations such as, for example, Part 2 of MPEG-2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, and also known as MPEG-2 video), AVC, HEVC, EVC, VVC, AV1, etc.
[0212] As shown in block 690, inputs to the elements of system 600 can be provided via various input devices. Such input devices may include, but are not limited to, the following (i) - (v). (i) An RF section capable of receiving RF signals wirelessly transmitted, for example, by a broadcast device, etc., (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI (registered trademark) input terminal, (v) a bus such as CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802 - 3) bus when the present invention is implemented in the automotive field.
[0213] In various exemplary embodiments, the input devices of block 690 may each have associated input processing elements as known in the relevant art. For example, the RF section may be associated with elements necessary for (i) selecting a desired frequency (also called selecting a signal or restricting a signal to a frequency band), (ii) down - converting the selected signal, (iii) again band - limiting to a narrower frequency band to select a signal frequency band, called a channel in a particular exemplary embodiment, for example, (iv) demodulating the down - converted and band - limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired data packet stream. The RF section of various exemplary embodiments may include one or more elements for performing these functions, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down - converter, a demodulator, an error corrector, and a demultiplexer. The RF section down - converts the received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. It may include a tuner for performing various functions.
[0214] In one example of a set-top box, the RF portion and its associated input processing elements can receive RF signals transmitted via a wired (e.g., cable) medium. The RF portion can then perform filtering, down-conversion, and re-filtering to perform frequency selection in the desired frequency band.
[0215] In various exemplary embodiments, the order of the above (and other) elements is rearranged, some of these elements are removed, and / or other elements performing similar or different functions are added.
[0216] Adding elements includes inserting elements between conventional elements, for example, inserting amplifiers or analog or digital converters. In various exemplary embodiments, RF may include an antenna.
[0217] Also, the USB and / or HDMI terminals may include corresponding interface processors to connect the system 600 to other electronic devices via SB and / or HDMI connections. Note that, when necessary, each aspect of input processing (e.g., Reed-Solomon error correction) can be implemented within a separate input processing IC or within the processor 610. Similarly, when necessary, each aspect of USB or HDMI interface processing can be implemented within a separate interface IC or within the processor 610. Demodulation, error correction, and the multiplexed and separated streams can be provided to various processing elements, including, for example, the processor 610 and the encoder / decoder 630, which operate in combination with memory and storage elements to process the data stream and display it on an output device when necessary.
[0218] The various elements of the system 600 can be provided within an integrated housing. Within the integrated housing, appropriate connection arrangements 690 can be used, for example, internal buses (including I2C buses), wiring, and printed circuit boards known in the art to interconnect the various types of elements and transmit data therebetween.
[0219] The system 600 may include a communication interface 650 so as to be able to communicate with other devices via a communication channel 651. The communication interface 650 may include, but is not limited to, a transceiver for transmitting and receiving data on the communication channel 651. The communication interface 650 includes, but is not limited to, a modem or a network card, and the communication channel 651 can be implemented, for example, within a wired and / or wireless medium.
[0220] In various exemplary embodiments, data can be streamed to the system 600 using a Wi-Fi network such as IEEE802.11. The Wi-Fi signals of these exemplary embodiments can be received via a communication channel 651 and a communication interface 650 that are compliant with Wi-Fi communication. The communication channel 651 of these exemplary embodiments can usually be connected to an access point or a router, and the access point or router enables streaming applications and other over-the-top communications by providing access to an external network including the Internet.
[0221] Other exemplary embodiments can provide data streamed to the system 600 using a set-top box, and the set-top box distributes data via the HDMI connection of the input block 690.
[0222] Other exemplary embodiments can provide the system 600 with data streamed using the RF connection of the input block 690.
[0223] The streamed data can be used as a method for signaling information used by the system 600. The signaling information may include information on the bitstream B and / or the number of pixels of the video image and / or any encoding / decoding setting parameters.
[0224] Note that signaling can be realized in various ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. may be used to signal information to the corresponding decoder.
[0225] System 600 can provide output signals to various output devices including display 661, speaker 671, and other peripheral devices 681. In various examples of the exemplary embodiments, other peripheral devices 681 may include one or more of an independent DVR, a disk player, a stereo system, an illumination system, and other devices that provide functions based on the output of system 600.
[0226] In various exemplary embodiments, the control signal uses, for example, AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or signaling of other communication protocols that enable control between devices regardless of the presence or absence of user intervention, to communicate between system 600 and display 661, speaker 671, or other peripheral devices 681.
[0227] The output devices can be communicatively coupled to system 600 via dedicated connections through corresponding interfaces 660, 670, and 680.
[0228] Optionally, the output devices can be connected to system 600 using communication channel 651 via communication interface 650. Display 661 and speaker 671 may be coupled in a single unit with other components of system 600 in an electronic device (e.g., a television).
[0229] In various exemplary embodiments, display interface 660 may include a display driver such as a timing controller (T Con) chip.
[0230] For example, if the RF portion of input terminal 690 is part of a separate set-top box, display 661 and speaker 671 may be selectively configured separately from one or more of the other components. In various exemplary embodiments where display 661 and speaker 671 may be external components, output signals can be provided via dedicated output connections (including, for example, HDMI ports, USB ports, or COMP output sides, etc.).
[0231] In FIGS. 1 - 12, various methods are described, each method including one or more steps or operations to implement the described method. The order and / or use of specific steps and / or operations can be changed or combined, provided that a specific order of steps or operations is not required for the exact operation of the method.
[0232] Several examples of block diagrams and / or flowcharts of operations are described. Each block represents a circuit element, module, or portion of code that includes one or more executable instructions for implementing the specified (one or more) logical functions. Note that in other embodiments, the (one or more) functions shown in the blocks may not be executed in the indicated order. For example, based on the related functions, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may be executed in reverse order in some cases.
[0233] The embodiments and aspects described herein can be realized, for example, in a method or process, apparatus, computer program, data flow, bit stream, or signal. Even when described in the context of a single form of embodiment (e.g., only described as a method), the embodiments of the described features may be realized in other forms (e.g., apparatus or computer program).
[0234] The method can be implemented, for example, within a processor, which generally refers to a processing device including a computer, a microprocessor, an integrated circuit, or a programmable logic device, etc. The processor further includes a communication device.
[0235] Also, the method can be realized by instructions executed by a processor, and such instructions (and / or data values generated by the embodiments) can be stored in a computer-readable storage medium. The computer-readable storage medium can use the aspect of a computer-readable program product 54C1 that is implemented on and within one or more computer-readable media and has computer-readable program code executable by a computer. Considering the inherent ability to store information therein and the inherent ability to retrieve information therefrom, the computer-readable storage medium used herein can be regarded as a non-transitory storage medium. The computer-readable storage medium may include, for example, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or apparatuses, or any suitable combination of the foregoing. The following shows more specific examples of computer-readable storage media to which this exemplary embodiment is applicable, but it should be noted that these are merely examples and not an exhaustive list, as can be easily understood by those skilled in the art. The above computer-readable storage medium can be, for example, a portable computer floppy disk, a hard disk, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0236] The instructions can form an application program specifically implemented on a readable medium.
[0237] For example, the instructions can exist in hardware, firmware, software, or a combination. For example, the instructions can be found in an operating system, a single application, or a combination of both. Thus, a processor may be characterized as a device, such as a device configured to execute a process, or a device that includes a processor-readable medium (such as a storage device) having instructions for executing a process. Also, in addition to or instead of the instructions, the processor-readable medium can store data values generated by an embodiment.
[0238] The apparatus can be implemented, for example, in suitable hardware, software, and firmware. Examples of such an apparatus include a personal computer, a laptop, a smartphone, a tablet, a digital multimedia set-top box, a digital television receiver, a personal video recording system, connected household appliances, a head-mounted display device (HMD, see-through glasses), a projector (beamer), a "cave" (a system including multiple displays), a server, a video encoder, a video decoder, a post-processor that processes the output from the video decoder, a pre-processor that provides an input to the video encoder, a web server, a set-top box, and any other device for processing video images, or other communication devices. Note that the apparatus can be mobile and mounted on a moving vehicle.
[0239] Computer software can be implemented by a processor 610, or by hardware, or by a combination of hardware and software. As a non-limiting example, exemplary embodiments can also be implemented with one or more integrated circuits. Memory 620 can be of any type suitable for the technical environment and can be implemented by any suitable data storage technology (non-limiting examples include, for example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory). Processor 610 can be of any type applicable to the technical environment and can include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0240] As will be apparent to those skilled in the art, embodiments can generate signals formatted to carry, for example, information to be stored or transmitted. The information can include instructions for executing a method or data generated by one of the described embodiments. For example, the signal can be formatted to carry a bitstream of the exemplary embodiments described. This signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as baseband. The formatting can include encoding the data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. As is well known, signals can be transmitted via a variety of different wired or wireless links. Signals can be stored on a processor-readable medium.
[0241] The terms used in this specification are for the purpose of describing particular exemplary embodiments and are not intended to be limiting. Unless otherwise indicated from the context, the singular forms "a", "an" and "the" used in this specification are intended to include the plural forms as well. In addition, the terms "include" and / or "including" used in this specification can specify the existence of stated features, integers, steps, operations, elements and / or components, etc., but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof. Furthermore, when one element is referred to as "responding to" or "connected to" or "associated with" another element, it may respond directly to the other element, or be associated with the other element, or there may be intermediate elements. In contrast, when one element is referred to as "responding directly to" or "directly connected to" or "directly associated with" another element, there are no intermediate elements.
[0242] Note that, for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the symbols / terms " / ", "and / or", and "at least one of" is intended to include the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or the selection of only the second-listed option (B), or the selection of only the third-listed option (C), or the selection of only the first and second-listed options (A and B), or the selection of only the first and third-listed options (A and C), or the selection of only the second and third-listed options (B and C), or the selection of all three options (A, B, and C). As is well known to those skilled in the art, it can be extended by the number of listed items.
[0243] In the present application, various numerical values can be used. Specific values may be used for illustrative purposes and each described aspect is not limited to these specific values.
[0244] Note that terms such as first, second, etc. may be used herein to describe various elements, but these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, unless departing from the teachings of the present application, the first element may be referred to as the second element, and similarly, the second element may be referred to as the first element. No order is implied between the first element and the second element.
[0245] References to "exemplary embodiment", "exemplary examples", "one embodiment", "embodiment" and other variations are often used to convey that certain features, structures, characteristics, etc. (described in combination with the exemplary embodiment / example) are included in at least one exemplary embodiment / example. Thus, the appearance of the terms "in one exemplary embodiment", "in an exemplary embodiment", "in one embodiment", "in an embodiment" and any variations thereof, wherever they appear in this application, do not necessarily refer to the same exemplary embodiment.
[0246] Similarly, references to "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" and other variations are often used to convey that certain features, structures or characteristics (described in combination with the exemplary embodiment / example) may be included in at least one exemplary embodiment / example. Thus, the expressions "according to an exemplary embodiment / example / embodiment" or "in an exemplary embodiment / example / embodiment" that appear throughout this application do not necessarily refer to the same exemplary embodiment / example, and individual or alternative exemplary embodiments / examples are not necessarily mutually exclusive of other exemplary embodiments / examples.
[0247] The reference numbers recited in the claims are for illustrative purposes only and do not limit the scope of the claims. Although not explicitly described, the exemplary embodiments / examples and variations can be used in any combination or sub - combination.
[0248] It should be understood that when the drawings are represented as flowcharts, block diagrams of the corresponding devices are also provided. Similarly, when the drawings are shown as block diagrams, flowcharts of the corresponding method / process are also provided.
[0249] Some of the drawings include arrows on communication paths to indicate the main direction of communication, but it should be understood that communication may occur in the direction opposite to the drawn arrows.
[0250] Various embodiments relate to decoding. As used in this application, "decoding" includes all or part of the process of the final output that is adapted for further processing in the displayed or reconstructed video region, which is performed on the received video image (which may include a received bitstream encoding one or more video images). In various exemplary embodiments, this process includes one or more of the processes typically performed by a decoder. In various exemplary embodiments, for example, such a process can optionally include the processes performed by the decoders of the various embodiments described in this application.
[0251] As a further example, in one exemplary embodiment, "decoding" refers only to inverse quantization. In one exemplary embodiment, "decoding" can refer to entropy decoding. In another exemplary embodiment, "decoding" may refer only to differential decoding. In another exemplary embodiment, "decoding" may refer to a combination of inverse quantization, entropy decoding, and differential decoding. Based on the specifically described context, it is considered that those skilled in the art will fully understand whether the term "decoding process" refers to a subset of operations or a more extensive decoding process.
[0252] Various embodiments relate to encoding. Similar to the above description of "decoding", "encoding" as used in this application can include all or part of the process performed on the input video image to generate an output bitstream. In various exemplary embodiments, this process includes one or more of the processes typically performed by an encoder. In various exemplary embodiments, this process further includes or optionally includes the processes performed by the encoders of the various embodiments described in this application.
[0253] As a further example, in an exemplary embodiment, "encoding" may refer only to quantization; in another exemplary embodiment, "encoding" may refer only to entropy encoding; in yet another exemplary embodiment, "encoding" may refer only to differential encoding; and in other exemplary embodiments, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Based on the specifically described context, it is clear whether the term "encoding process" refers to a specific subset of operations or to a broader encoding process, and it is considered to be well understood by those skilled in the art.
[0254] Also, in this application, "acquisition" of various information is mentioned. Acquisition of information may include, for example, any one or more of estimation of information, calculation of information, prediction of information, search for information from memory, processing of information, transfer of information, copying of information, deletion of information, calculation of information, determination of information, prediction of information, or estimation of information.
[0255] Also, in this application, "receiving" of various information is mentioned. Receiving of information may include, for example, any one or more of accessing information or receiving information from a communication network.
[0256] And, as used herein, the term "signal" refers specifically to indicating something to a corresponding decoder. For example, in some exemplary embodiments, the encoder signals certain information such as encoded parameters or encoded video image data. In this way, in the exemplary embodiments, the same parameters can be used on the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicitly signal) specific parameters to the decoder, whereby the decoder can use the same specific parameters. In contrast, when the decoder has specific parameters and other parameters, by signaling (implicitly signaling) without transmission, the decoder is informed of and obtains the specific parameters. By avoiding the transmission of actual functions, bit savings are achieved in various exemplary embodiments. It should be understood that multiple ways of signaling can be accomplished. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. The foregoing relates to the verb form of the term "signal", but the term "signal" may also be used as a noun herein.
[0257] Although many embodiments have been described, it should be understood that various modifications are possible. For example, elements of different implementations can be combined, supplemented, changed, or deleted to produce other embodiments. Further, as would be understood by one of ordinary skill in the art, the disclosed structures and processes can be replaced with other structures and processes, and the resulting embodiments perform at least substantially the same function, in at least substantially the same way, and achieve at least substantially the same result as the disclosed implementations. Accordingly, these embodiments and other embodiments are contemplated by this application.
[0258] This application claims the priority of European Patent Application No. "22306003.9" filed on July 5, 2022, the entire content of which is incorporated herein by reference.
Claims
1. A method for predicting a block of a video image, comprising: - Signaling (301) a type of intra block copy mode, wherein the type of the intra block copy mode indicates whether the intra block copy prediction mode is set to predict video content captured by a camera, and the intra block copy prediction mode determines at least one block vector for predicting a block of the video image from at least one reference block of the video image; - When the type of the intra block copy mode indicates that the intra block copy prediction mode is set to predict video content captured by a camera, - Setting (302) the intra block copy prediction mode to predict a block of the video image; - Deriving (303) a predicted block of the block of the video image based on the set intra block copy mode prediction mode. A method for predicting a block of a video image.
2. The method according to claim 1, wherein the block belongs to a slice of the video image, and the type of the intra block copy mode is signaled (301) at a slice level. The method for predicting a block of a video image according to claim 1.
3. Setting (302) the intra block copy prediction mode to predict video content captured by a camera further comprises: - Signaling (304) a syntax element (ibc_pred_idc_flag) indicating whether bidirectional prediction is permitted, and when the syntax element indicates that bidirectional prediction is permitted, - Signaling (305) two block vectors, and further comprising that the prediction of the block of the video image is an average value of the two block vectors. The method for predicting a block of a video image according to claim 1 or 2.
4. Setting (302) the intra block copy prediction mode to predict video content captured by a camera further comprises: Signaling (306) a syntax element (mvp_l1_flag) to identify a block vector predictor of each block vector in a block vector list associated with a decoded block adjacent to a block of a video image, further including that the adjacent block belongs to the video image. A method for predicting a block of a video image according to claims 1 to 3. **Claim 5** At least one block vector is represented at a sub-pixel accuracy level. A method for predicting a block of a video image according to any one of claims 1 to 4. **Claim 6** Setting (302) an intra block copy prediction mode to predict video content captured by a camera is Signaling (308) at least one block vector difference calculated between a block vector and a reference block within a video image, further including that each block vector difference is signaled at a sub-pixel accuracy level. A method for predicting a block of a video image according to any one of claims 1 to 5. **Claim 7** Setting (302) an intra block copy prediction mode to predict video content captured by a camera is Encoding (309) each block vector difference, further including that the encoding is limited to several vector directions and widths. A method for predicting a block of a video image according to claim 6. **Claim 8** Setting (302) an intra block copy prediction mode to predict video content captured by a camera is Signaling (step 310) a syntax element (mmvd_merge_flag) to indicate whether the encoding of the block vector difference is limited to several vector directions and widths. A method for predicting a block of a video image according to claim 7. **Claim 9** Setting (302) an intra block copy prediction mode to predict video content captured by a camera is Signaling (313) a syntax element (intra_tmp_flag) to indicate whether to derive a predicted block of a block of a video image using intra block copy prediction based on template matching or whether to derive a predicted block of a block of a video image based on any one of claims 1 to 11. A method for predicting a block of a video image according to claim 1 or 2.
10. The step of deriving (303) a predicted block of a block of a video image based on a set intra block copy mode prediction mode includes: - a step of deriving (314) a first predicted block of the block of the video image from the set intra block copy prediction mode; - a step of deriving (315) a second predicted block of the block of the video image; - a step of blending the first predicted block and the second predicted block to derive (316) a predicted block of the block of the video image. A method for predicting a block of a video image according to claim 1 or 2.
11. The second predicted block is derived based on motion compensation prediction or intra prediction of the block of the video image. A method for predicting a block of a video image according to claim 10.
12. A method for encoding a block of a video image based on a predicted block derived from the method according to any one of claims 1 to 11.
13. A method for decoding a block of a video image based on a predicted block derived from the method according to any one of claims 1 to 11.
14. A bitstream formatted to include encoded video image data obtained from the method according to any one of claims 1 to 11.
15. An apparatus comprising means for performing the method according to any one of claims 1 to 13.
16. A computer program product comprising instructions, wherein when the program is executed by one or more processors, the instructions cause the one or more processors to perform the method according to any one of claims 1 to 13. Computer program product.
17. A non-transitory storage medium containing instructions of program code for performing the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Intra block copy merging data syntax for video coding
US20200336735A1
Coding of block vectors for intra block copy-coded blocks
WO2020238837A1