Encoding / decoding video image data

By determining and inferring intra prediction modes of different geometric transformation types in video image encoding, the problem of high signal cost in the intra prediction mode in the prior art is solved, and the encoding efficiency is improved.

CN120113239APending Publication Date: 2025-06-06BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380063500.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-04-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the existing video image encoding technology, the signal cost of the intra prediction mode is relatively high, which affects the encoding efficiency.

Method used

The signal cost is reduced by determining whether adjacent blocks are predicted according to intra prediction modes using different geometric transformation types and inferring the first geometric transformation type from these prediction modes.

Benefits of technology

The signal cost of the intra prediction mode is reduced and the encoding efficiency of video image block prediction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120113239A_ABST
    Figure CN120113239A_ABST
Patent Text Reader

Abstract

The present invention relates to a method of intra prediction of a block of a video image according to a first intra prediction mode using a first geometric transformation identified by a first geometric transformation type. The method determines (510) whether to predict at least one neighboring block according to a second intra prediction mode using a second geometric transform identified by a second geometric transform type, and inferes (520) a first geometric transform type from the second geometric transform type if at least one neighboring block is predicted according to the second intra prediction mode.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is based on and claims priority from European patent application No. “22306423.9” filed on September 27, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present application generally relates to video image encoding and decoding. In particular, but not exclusively, the technical field of the present application relates to intra-frame prediction of video image blocks. Background Art

[0004] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of at least one exemplary embodiment of the present application described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of various aspects of the present application. Therefore, it should be understood that these statements should be read in this light, and not as admissions of prior art.

[0005] In state-of-the-art video compression systems such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Generic Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level picture partitioning are provided to divide a video picture into picture regions, so-called coding tree units (CTUs), which for HEVC may typically have a size between 16×16 and 64×64 pixels, while for VVC the size of the coding tree unit (CTU) may be 32×32, 64×64 or 128×128 pixels.

[0006] The CTU partitioning of the video image forms a grid of CTUs of a fixed size, i.e., a CTU grid, in which the upper and left boundaries coincide spatially with the upper and left borders of the video image. The CTU grid represents the spatial partitioning of the video image.

[0007] In VVC and HEVC, the CTU size (CTU width and CTU height) of all CTUs of the CTU grid is equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) may be equal to 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is encoded into the bitstream, for example at the sequence level in a sequence parameter set (SPS).

[0008] The spatial position of a CTU in the CTU grid is determined by the CTU address ctuAddr, which defines the spatial position of the upper left corner of the CTU from the origin. Figure 1 As illustrated, a CTU address may define a spatial position from the upper left corner of a higher-level spatial structure S that contains the CTU.

[0009] A coding tree is associated with each CTU to determine the tree partitioning of the CTU.

[0010] like Figure 1 As shown, in HEVC, the coding tree is a quadtree partition of CTU, where each leaf is called a coding unit (CU). The spatial position of the CU in the video image is defined by the CU index cuIdx, which indicates the spatial position from the upper left corner of the CTU. The CU is spatially partitioned into one or more prediction units (PUs). The spatial position of the PU in the video image VP is defined by the PU index puIdx, which defines the spatial position from the upper left corner of the CTU, and the spatial position of the elements of the partitioned PU is defined by the PU partition index puPartIdx, which defines the spatial position from the upper left corner of the PU. Each PU is assigned some intra-frame or inter-frame prediction data.

[0011] Intra or inter coding modes are assigned on CU level. This means the same intra / inter coding mode is assigned to each PU of a CU, although the prediction parameters vary from PU to PU.

[0012] According to a quadtree called a transform tree, a CU can also be spatially partitioned into one or more transform units (TUs). A transform unit is a leaf of a transform tree. The spatial position of a TU in a video image is defined by a TU index tuIdx, which defines the spatial position from the upper left corner of the CU. Each TU is assigned some transform parameters. The transform type is assigned at the TU level, and a 2D separate transform is performed at the TU level during encoding or decoding of an image block.

[0013] The PU partition types that exist in HEVC are as follows: Figure 2As shown. They include square partitions (2N×2N and N×N), which are the only partitions used in both intra- and inter-predicted CUs; ​​symmetric non-square partitions (2N×N, N×2N, used only in inter-predicted CUs); and asymmetric partitions (used only in inter-predicted CUs). For example, PU type 2N×nU represents an asymmetric horizontal partition of a PU, where a smaller partition is located at the top of the PU. According to another example, PU type 2N×nL represents an asymmetric horizontal partition of a PU, where a smaller partition is located at the top of the PU.

[0014] like Figure 3 As shown, in VVC, the coding tree starts from the root node (i.e., CTU). Next, a quadtree (or quadtree) partitioning divides the root node into 4 nodes (solid lines) corresponding to 4 sub-blocks of equal size. Next, the quadtree (or quadtree) leaves can then be further partitioned by the so-called multi-type tree, which involves Figure 4 Binary or ternary segmentation of one of the 4 segmentation modes shown. These segmentation types are vertical and horizontal binary segmentation modes, denoted SBTV and SBTH; and vertical and horizontal ternary segmentation modes SPTTV and STTH.

[0015] In the case of a joint coding tree shared by luma and chroma components, the leaves of the coding tree of a CTU are CUs.

[0016] In contrast to HEVC, in VVC, in most cases, CU, PU, ​​and TU have equal sizes, which means that coding units are generally not partitioned into PUs or TUs except in some specific coding modes.

[0017] Figure 5 and Figure 6 An overview of video encoding / decoding methods used in current video standard compression systems (such as, for example, VVC) is provided.

[0018] Figure 5 A schematic block diagram showing the steps of a method 100 for encoding a video image VP according to the prior art is shown.

[0019] In step 110, the video picture VP is partitioned into a CTU grid and the partition information data is signaled into the bitstream B. A coding tree is associated with each CTU of the CTU grid, and each CU of the coding tree associated with each CTU is a block of samples of the video picture VP. In short, a CU of a CTU is a block.

[0020] The CTUs of the CTU grid are considered along a scan order, typically a raster scan order of a video picture. Each block of the CTU is also considered along a scan order, typically a raster scan order of the blocks of the CTU.

[0021] Each block of each CTU is then encoded using intra or inter prediction coding mode.

[0022] Intra prediction (step 120) consists in predicting the current block with the help of a predicted block based on already coded, decoded and reconstructed samples located around the current block (usually located at the top and to the left of the current block).Intra prediction is performed in the spatial domain.

[0023] In inter prediction mode, motion estimation (step 130) and motion compensation (step 135) are performed. Motion estimation searches for candidate reference blocks that are good predictors of the current block in one or more reference video images used to predictively encode the current video image. For example, a good predictor of the current block is a predictor that is similar to the current block. The output of the motion estimation step 130 is one or more motion vectors and (one or more) reference image indices associated with the current block. Next, motion compensation (step 135) obtains the predicted block with the help of the (one or more) motion vectors and (one or more) reference image indices determined by the motion estimation step 130. Basically, the block belonging to the selected reference image and pointed to by the motion vector can be used as the predicted block of the current block. In addition, since the motion vector is represented as a fraction of an integer pixel position (this is called sub-pixel precision motion vector representation), motion compensation usually involves spatial interpolation of some reconstructed samples of the reference image to calculate the predicted block samples.

[0024] The prediction information data is signaled in the bitstream B. The prediction information may include prediction mode, prediction information coding mode, intra prediction mode or motion vector(s) and reference picture index(es) and any other information for obtaining the same predicted block at the decoding side.

[0025] The method 100 optimizes the rate-distortion trade-off to select one of the intra-frame mode or the inter-frame coding mode by considering, for example, the encoding of the prediction residual block calculated by subtracting the candidate predicted block from the current block, and the signaling of the prediction information data required to determine the candidate predicted block at the decoding side.

[0026] Typically, the best prediction mode is given as the prediction mode of the best coding mode p* for the current block given by the following equation:

[0027]

[0028] Where P is the set of all candidate coding modes for the current block, p represents the candidate coding mode in the set, and RD cost(p) is the rate-distortion cost of candidate coding mode p, usually expressed as:

[0029] RD cost(p) =D(p)+λ.R(p)

[0030] D(p) is the distortion between the current block and the reconstructed block obtained after encoding / decoding the current block with candidate coding mode p, R(p) is the rate cost associated with encoding the current block with coding mode p, and λ is a Lagrangian parameter that represents the rate constraint for encoding the current block and is typically calculated based on the quantization parameter used to encode the current block.

[0031] The current block is usually encoded according to a prediction residual block PR. More precisely, the prediction residual block PR is calculated, for example, by subtracting the best predicted block from the current block. The prediction residual block PR is then transformed (step 140) by using, for example, a DCT (discrete cosine transform) or DST (discrete sine transform) type transform, and the obtained transform coefficient block is quantized (step 150).

[0032] In a variant, the method 100 may also skip the transform step 140 and apply quantization directly to the prediction residual block PR according to a so-called transform skip coding mode.

[0033] The block of quantized transform coefficients (or the block of quantized prediction residuals) is entropy encoded into the bitstream B (step 160).

[0034] Next, the quantized transform coefficient block (or quantized residual block) is dequantized (step 170) and inverse transformed (180) (or not inverse transformed) to produce a decoded prediction residual block. The decoded prediction residual block and the predicted block are then combined, usually summed, which provides a reconstructed block.

[0035] Other information data may also be entropy encoded to encode the current block of the video picture VP.

[0036] An in-loop filter (step 190) may be applied to the reconstructed image (including the reconstructed blocks) to reduce compression artifacts. After all image blocks are reconstructed, a loop filter may be applied. For example, they include a deblocking filter, a sample adaptive offset (SAO), or an adaptive loop filter.

[0037] The reconstructed block or the filtered reconstructed block forms a reference picture which can be stored into a decoded picture buffer (DPB) so that it can be used as a reference picture for encoding of the next current block of the video picture VP or the next video picture to be encoded.

[0038] Figure 6 A schematic block diagram showing the steps of a method 200 for decoding a video image VP according to the prior art is shown.

[0039] In step 210, partition information data, prediction information data and a quantized transform coefficient block (or a quantized residual block) are obtained by entropy decoding the bitstream B.

[0040] The partition information data defines a CTU grid (arrangement) on the video picture. The CTU grid divides the video picture VP into a plurality of CTUs. The CTUs of the CTU grid are considered along a scan order, typically a raster scan order of the video picture VP. The blocks of the considered CTUs are also considered along a scan order, typically a raster scan order of the blocks of the CTUs.

[0041] Other information data may also be decoded from the bitstream B for decoding a current block of a current CTU of the CTU grid.

[0042] In step 220, entropy decoding is performed on each current block of the current CTU.

[0043] Each decoded current block may be a quantized transform coefficient block or a quantized prediction residual block.

[0044] In step 230, the current block (of the current CTU) is dequantized and possibly inverse transformed (step 240) to obtain a decoded prediction residual block.

[0045] On the other hand, the prediction information data is used to predict the current block. The predicted block is obtained by its intra prediction (step 250) or its motion compensated temporal prediction (step 260). The prediction process performed at the decoding side is the same as the prediction process performed at the encoding side.

[0046] Next, the decoded prediction residual block and the predicted block are then combined, typically summed, which provides the reconstructed block.

[0047] In step 270, the in-loop filter may be applied to the reconstructed image (including the reconstructed block), and the reconstructed block or the filtered reconstructed block forms a reference image, which may be stored in a decoded picture buffer (DPB) as discussed above ( Figure 5 ).

[0048] The intra block copy (IBC) prediction mode is used for screen content coding in HEVC and VVC. It is well known that the IBC prediction mode significantly improves the coding efficiency of screen content material. Since the IBC prediction mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each block to be predicted. The block vector indicates the displacement from the block to be predicted of the video image to the reference block (i.e., the prediction block of the reconstructed area of ​​the video image), which has been reconstructed (decoded) within the video image. The block vector of the IBC prediction block of the luminance sample is integer precision. The block vector of the IBC predicted block of the chrominance sample is also rounded to integer precision. When combined with AMVR (an adaptive motion vector resolution tool that allows motion vector differences to be signaled with quarter-pixel, half-pixel, integer-pixel or 4-pixel luminance sample resolution), the IBC prediction mode can switch between 1 pixel and 4 pixel motion vector precision. The IBC prediction mode is regarded as a third prediction mode in addition to the intra or inter prediction mode.

[0049] At the CU level, the IBC prediction mode is signaled as prediction information with a flag indicating the use of IBC-AMVP mode or IBC-skip / merge mode.

[0050] In IBC skip / merge mode, the block vector used to define the IBC predicted block is represented by a merge index that indicates an element in a block vector prediction candidate list (merge candidate list). The merge candidate list can include spatial, HMVP and paired block vector candidates.

[0051] HMVP (History-based Motion Vector Prediction) involves a buffer of block vector (motion vector) candidates that is fed whenever a block of a video image is predicted by using an IBC prediction mode. The HMVP buffer is used to provide block vector candidates for predicting a current block of a video image by using an IBC prediction mode.

[0052] For HMVP, the block vectors are inserted into a history buffer for future reference.

[0053] A paired block vector candidate can be generated by averaging two IBC block vector candidates (i.e., block vectors derived from the IBC prediction mode). This means that the first two block vector candidates in the merge candidate list being constructed are averaged to form a so-called paired block vector candidate. This paired block vector candidate is added to the merge candidate list after the HMVP candidate.

[0054] In IBC-AMVP mode, two block vectors are determined, one from the left neighbor of the block to be predicted and one from the upper neighbor of the block to be predicted. When either neighbor is not available, a default block vector is considered. A flag is signaled as prediction information to indicate which block vector is used to predict the block vector of the current block. In practice, since at most 2 block vector candidates are considered in IBC-AMVP mode, the flag is sufficient to identify the block vector candidate for encoding the block vector information. The block vector difference is encoded in the same way as the motion vector difference of the inter-predicted block.

[0055] The IBC prediction mode cannot be used in combination with VVC's inter prediction tools (such as CIIP and MMVD).

[0056] The reconstruction-reordered IBC (RR-IBC) mode (Non-EE2: Reconstruction-Reordered IBC for screen content coding, Z. Deng et al., JVET-Z0159, April 2022) reorders the reconstructed blocks of IBC encoding with horizontal or vertical flipping, signaling the mode. The RR-IBC mode can be RR-IBC-AMVP mode or RR-IBC-Skip / Merge mode. The RR-IBC-AMVP mode represents the IBC-AMVP mode, in which the reconstructed blocks of IBC encoding are reordered with horizontal or vertical flipping, signaling the mode, and the RR-IBC Skip / Merge mode represents the IBC Skip / Merge mode, in which the reconstructed blocks of IBC encoding are reordered by horizontal or vertical flipping, signaling the mode.

[0057] Figure 7 A block diagram schematically illustrates a method 300 of toggle type selection for RR-IBC mode at the encoding side.

[0058] In step 310, the best flip type is selected from a set of flip types including identity (no flip), horizontal flip, and vertical flip.

[0059] For example, the best flipping type is selected by optimizing the rate-distortion tradeoff.

[0060] In step 320, the best flipping type is signaled in the bitstream.

[0061] Signaling the optimal flip type reduces the coding efficiency of blocks of the video picture.

[0062] The problem solved by the present invention is to reduce the signal cost.

[0063] At least one exemplary embodiment of the present application has been designed in view of the foregoing. Summary of the invention

[0064] The following section presents a simplified summary of at least one exemplary embodiment in order to provide a basic understanding of some aspects of the present application. This summary is not an exhaustive overview of the exemplary embodiments. It is not intended to identify the key or important elements of the exemplary embodiments. The following summary merely presents some aspects of at least one exemplary embodiment in a simplified form as a prelude to a more detailed description provided elsewhere in the document.

[0065] According to a first aspect of the present application, there is provided a method for intra-predicting a block of a video image according to a first intra-prediction mode using a first geometric transformation identified by a first geometric transformation type, wherein the method comprises:

[0066] - determining whether to predict at least one neighboring block according to a second intra prediction mode using a second geometric transform identified by a second geometric transform type; and

[0067] - if at least one neighboring block is predicted according to a second intra prediction mode, inferring the first geometric transformation type from the second geometric transformation type.

[0068] In an exemplary embodiment, the first geometric transformation type and / or the second geometric transformation type identifies at least one of the following geometric transformations or a combination of at least two of the following geometric transformations:

[0069] -Flip horizontally;

[0070] -Flip vertically;

[0071] - Horizontal and vertical flip;

[0072] - Rotation;

[0073] -identity (or sameness).

[0074] In an exemplary embodiment, when none of the neighboring blocks is predicted according to the second intra prediction mode, the first geometric transformation type is determined by the first intra prediction mode.

[0075] In an exemplary embodiment, the determined first transform type is signaled in the bitstream.

[0076] In an exemplary embodiment, the first geometric transform type is selected among the second geometric transform types, the second geometric transform type identifying a second geometric transform used by a second intra prediction mode of the neighboring block.

[0077] In an exemplary embodiment, inferring the first geometric transformation type from the second geometric transformation type comprises:

[0078] - determining a second geometric transform type having a greater number of occurrences in a second geometric transform used by a second intra prediction mode for predicting the at least one neighboring block; and

[0079] - Setting the first geometric transformation type to the second geometric transformation type having a greater number of occurrences.

[0080] In an exemplary embodiment, when predicting the single neighboring block according to a second intra prediction mode, inferring the first geometric transform type from the second geometric transform type includes setting the first geometric transform type to a second geometric transform type associated with the second intra prediction mode.

[0081] In an exemplary embodiment, when predicting at least one neighboring block according to the second intra prediction mode, the method further includes signaling an index in the bitstream indicating at least one neighboring block predicted by the second intra prediction mode using a second geometric transform type having a greater number of occurrences.

[0082] In an exemplary embodiment, if the second geometric transform type having a greater number of occurrences is not identical to the first geometric transform type derived from the first intra prediction mode, the method further comprises signaling the first geometric transform type in the bitstream.

[0083] In an exemplary embodiment, the second geometric transform type used by the second intra prediction mode can be obtained from the bitstream or derived from the second intra prediction mode.

[0084] In an exemplary embodiment, when the first geometric transform type determined from the first intra prediction mode is different from the inferred first geometric transform type, the inferred first geometric transform type is signaled in the bitstream.

[0085] In an exemplary embodiment, the first intra prediction mode is a reordered intra-bloc-copy mode using a first geometric transform type, the first geometric transform type indicating at least one of the following geometric transforms or a combination of at least two of the following geometric transforms:

[0086] -Flip horizontally;

[0087] -Flip vertically;

[0088] - Horizontal and vertical flip;

[0089] - Rotation;

[0090] - identity;

[0091] And the second intra prediction is an intra template matching prediction mode using a second geometric transformation type, the second geometric transformation type indicating at least one of the following geometric transformations or a combination of at least two of the following geometric transformations:

[0092] -Flip horizontally;

[0093] -Flip vertically;

[0094] - Horizontal and vertical flip;

[0095] - Rotation;

[0096] -Identity.

[0097] In an exemplary embodiment, the first intra prediction is an intra template matching prediction mode using a second geometric transform type, the second geometric transform type indicating at least one of the following geometric transforms or a combination of at least two of the following geometric transforms:

[0098] -Flip horizontally;

[0099] -Flip vertically;

[0100] - Horizontal and vertical flip;

[0101] - Rotation;

[0102] - identity;

[0103] And the second intra prediction mode is a reordered intra block copy mode using a first geometric transform type, the first geometric transform type indicating at least one of the following geometric transforms or a combination of at least two of the following geometric transforms:

[0104] -Flip horizontally;

[0105] -Flip vertically;

[0106] - Horizontal and vertical flip;

[0107] - Rotation;

[0108] -Identity.

[0109] According to a second aspect of the present application, there is provided a method for encoding a block of a video image based on a prediction block derived according to the method of the first aspect.

[0110] According to a third aspect of the present application, there is provided a method for decoding a block of a video image based on a prediction block derived according to the method of the first aspect.

[0111] According to a fourth aspect of the present application, an apparatus is provided, comprising means for executing one of the methods according to the first, second and / or third aspects.

[0112] According to a fifth aspect of the present application, a computer program product is provided, comprising instructions, which, when the program is executed by one or more processors, cause the one or more processors to perform the method according to the first, second and / or third aspect.

[0113] According to a sixth aspect of the present application, a non-transitory storage medium is provided, which carries instructions of a program code for executing the method according to the first, second and / or third aspect.

[0114] Specific properties of at least one of the exemplary embodiments and other objects, advantages, features and uses of at least one of the exemplary embodiments will become apparent from the following description of the examples taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Reference will now be made, by way of example, to the accompanying drawings which show exemplary embodiments of the present application, and in which:

[0116] Figure 1 An example of a coding tree unit according to HEVC is shown;

[0117] Figure 2 An example of partitioning a coding unit into prediction units according to HEVC is shown;

[0118] Figure 3 An example of CTU partitioning according to VVC is shown;

[0119] Figure 4 An example of supported partitioning modes in multi-type tree partitioning according to VVC is shown;

[0120] Figure 5 A schematic block diagram showing the steps of a method 100 for encoding a video image VP according to the prior art;

[0121] Figure 6 A schematic block diagram showing the steps of a method 200 for decoding a video image VP according to the prior art;

[0122] Figure 7 Schematically illustrates a block diagram of a method 300 of toggle type selection of RR-IBC mode at the encoding side according to the prior art;

[0123] Figure 8 A block diagram schematically illustrates a method 600 for performing intra-frame prediction on a target block based on template matching according to an exemplary embodiment;

[0124] Fig. 9 The figure illustrates a search area defined for predicting a target block TB of a video image;

[0125] Fig.10 FIG. 1 illustrates a case where the geometric transformation is horizontal flipping according to an exemplary embodiment. Figure 8 Method 400;

[0126] Fig.11 illustrates some examples of neighborhood template determination according to flipping (geometric transformation type) according to exemplary embodiments;

[0127] Fig.12 illustrates some examples of neighborhood template determination according to rotation (geometric transformation type) according to exemplary embodiments;

[0128] Fig.13 schematically illustrates a block diagram of a method 500 for intra-predicting a block of a video image according to a first intra-prediction mode using a first geometric transformation according to an exemplary embodiment;

[0129] Fig.14 illustrates an example of neighboring blocks according to an exemplary embodiment;

[0130] Fig.15 schematically illustrates a block diagram of a method 600 of intra-predicting a block of a video image at an encoding side by using a first intra-prediction mode, the first intra-prediction mode using a first geometric transform type according to an exemplary embodiment;

[0131] Fig.16 Schematically illustrates how Fig.15 a block diagram of a variant 700 of the method 600 discussed;

[0132] Fig.17 schematically illustrates a block diagram of a method 800 of intra-predicting a block of a video image by using a first intra-prediction mode, the first intra-prediction mode using a first geometric transform type according to an exemplary embodiment;

[0133] Fig.18 A block diagram schematically illustrates a method 900 of intra-predicting a block of a video image when a first intra-prediction mode is an RR-IBC mode and a second intra-prediction mode is an ITMP mode according to a first exemplary embodiment;

[0134] Fig.19 Schematically illustrates Fig.18 A block diagram of a variant 1000 of method 900;

[0135] Fig. 20A block diagram schematically illustrates a method 1100 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment;

[0136] Fig.21 A block diagram schematically illustrates a method 1200 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment;

[0137] Fig. 22 A block diagram schematically illustrates a method 1300 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment;

[0138] Fig.23 A schematic block diagram of an example of a system in which various aspects and exemplary embodiments are implemented is illustrated.

[0139] Similar or identical elements are denoted by the same reference numerals. DETAILED DESCRIPTION

[0140] At least one exemplary embodiment of the exemplary embodiments will be described more fully below with reference to the accompanying drawings, wherein an example of at least one exemplary embodiment of the exemplary embodiments is depicted. However, the exemplary embodiments can be implemented in many alternative forms and should not be interpreted as being limited to the examples set forth herein. Thus, it should be understood that the exemplary embodiments are not intended to be limited to the specific forms disclosed. On the contrary, the present application is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the present application.

[0141] At least one of these aspects generally relates to video image encoding and decoding, another aspect generally relates to transmitting a provided or encoded bitstream, and one of the other aspects relates to receiving / accessing a decoded bitstream.

[0142] At least one of these exemplary embodiments is described for encoding / decoding video images, but is extended to encoding / decoding video images (image sequences) in that each video image is sequentially encoded / decoded as described below.

[0143] In addition, at least one exemplary embodiment is not limited to the current version of VVC. The at least one exemplary embodiment may be applicable to existing or future developed proposals and VCC and proposed extensions. Unless otherwise specified or technically excluded, the aspects described in this application may be used alone or in combination.

[0144] A pixel corresponds to the smallest display unit on the screen, which can be composed of one or more light sources (1 for a monochrome screen and 3 or more for a color screen).

[0145] A video image (also called a frame or image frame) comprises at least one component (also called an image component or channel) determined by a specific image / video format, which specifies all information related to pixel values ​​and all information related to the video image data that can be used by a display unit and / or any other device to display and / or decode the video image data related to the video image.

[0146] A video image comprises at least one component which is typically represented in the shape of an array of samples.

[0147] A monochrome video image includes a single component, while a color video image may include three components.

[0148] For example, a color video image may include a luminance (or luminance) component and two chrominance components when the image / video format is the well-known (Y,Cb,Cr) format, or may include three color components (one for red, one for green, and one for blue) when the image / video format is the well-known (R,G,B) format.

[0149] Each component of the video image may comprise a number of samples relative to the number of pixels of the screen on which the video image is intended to be displayed. In a variant, the number of samples contained in a component may be a multiple (or fraction) of the number of samples contained in another component of the same video image.

[0150] For example, in case the video format includes a luma component and two chroma components (such as a (Y, Cb, Cr) format), the chroma components may contain half the number of samples in width and / or height relative to the luma components, depending on the color format considered.

[0151] A sample is the smallest visual information unit that makes up a component of a video image. A sample value can be, for example, a brightness or chrominance value or a color value in (R, G, B) format.

[0152] Reconstructed samples are samples that have been encoded according to method 100 and decoded according to method 200 .

[0153] The pixel value is the value of a screen pixel. For a monochrome video image, a pixel value can be represented by one sample, and for a color video image, a pixel value can be represented by multiple co-located samples. The co-located sample associated with a pixel refers to a sample corresponding to the position of the pixel in the screen.

[0154] A video image is usually considered as a set of pixel values, with each pixel represented by at least one sample.

[0155] A block of a video image is a group of samples of a component of a video image. When the image / video format is the well-known (Y, Cb, Cr) format, at least one block of luminance samples or at least one block of chrominance samples may be considered, or when the image / video format is the well-known (R, G, B) format, at least one block of color samples may be considered.

[0156] At least one exemplary embodiment is not limited to a particular image / video format.

[0157] In the present application, the RR-IBC mode indicates the RR-IBC mode defined by JVET-Z0159 (Non-EE2: Reconstruction-Reordered IBC for screen content coding, Z. Deng et al., April 2022) and integrated in ECM (Algorithm description of Enhanced Compression Model 6 (ECM 6), M. Coban et al., JVET-AA2025, July 2022), wherein the geometric transformation type may indicate at least one of the following geometric transformations or a combination of at least two of the following geometric transformations:

[0158] -Flip horizontally;

[0159] -Flip vertically;

[0160] - Horizontal and vertical flip;

[0161] - Rotation;

[0162] -Identity.

[0163] Then, according to the exemplary embodiment, modify Figure 7 Step 310 is replaced by step 330, in which the best geometric transform type is selected from a set of geometric transforms including a combination of at least one of the preceding geometric transforms or at least two of the following geometric transforms, and step 320 is replaced by step 340, signaling the best geometric transform type in the bitstream.

[0164] In the following, the intra template matching prediction (ITMP) mode is an intra prediction mode of a target block TB (block of a video image to be predicted). The prediction block of the target block is obtained based on the best reference block located in the search area (reconstructed samples of the video image).

[0165] A target template patch TTP is associated with a target block TB. The target template patch TTP is constructed based on the available reconstructed samples in the target block neighborhood TBN. The target block neighborhood TBN defines the locations of the available reconstructed samples used to construct the target template patch TTP.

[0166] The availability of reconstructed samples in the target block neighborhood TBN depends on the location of the target block TB relative to the video image or the search area boundary.

[0167] At both the encoder and the decoder, the best reference block is produced by a candidate template matching search, during which, for each current position in the search area, a candidate reference block RB located at the current position is considered, and a candidate reference template patch RTP is associated with the candidate reference block RB.

[0168] The cost of each reference block RB is evaluated by comparing the reconstructed samples of the reference template patch RTP with the reconstructed samples of the target template patch TTP associated with the target block TB (block of the video image to be predicted).

[0169] Then, the best reference block is the candidate reference template patch RTP that provides the minimum cost. The best reference block is the prediction block of the target block.

[0170] Figure 8 A block diagram of a method 400 for intra-predicting a target block TB based on template matching according to an exemplary embodiment is schematically illustrated.

[0171] Method 400 is implemented in the same manner as method 100 (encoding) and method 200 (decoding), thereby saving data signaling and avoiding additional complexity of control.

[0172] The method 400 is based on a template matching candidate search that compares a target template patch TTP with a candidate reference template patch RTP associated with a reference block RB located at a certain position of the search area.

[0173] Fig. 9 The diagram shows a search area defined for predicting a target block TB of a video image.

[0174] Typically, the search area includes four spatial regions R1, R2, R3 and R4. Region R1 is the current CTU including the target block TB (block to be predicted of the video image), region R2 is the upper left CTU, R3 is the upper CTU, and R4 is the left CTU.

[0175] The width SearchRange_w and height SearchRange_h of the regions R1 to R3 are set to be proportional to the target block width BlkW and the target block height BlkH to have a fixed number of comparisons per pixel. That is:

[0176] SearchRange_w=a*BlkW

[0177] SearchRange_h = a * BlkH (2)

[0178] Where "a" is a constant integer value that controls the gain / complexity trade-off. The factor "a" may depend on the target block width and height, the target template width (e.g., fixed to 4 samples), and / or the CTU size. In practice, "a" is equal to 5.

[0179] exist Figure 8 In step 410, a target block neighborhood TBN is defined according to the availability of reconstructed samples in the current CTU and surrounding CTUs including the target block TB. As discussed later, the target block neighborhood TBN defines the locations of available reconstructed samples for constructing a target template patch TTP.

[0180] For example, Fig. 9 As illustrated, if a vertical or horizontal flip occurs on the target block TB, the neighboring samples located below or to the right of the target block, respectively, are unavailable (have not been reconstructed) and cannot be used to generate the target block neighborhood TBN used in the template matching candidate search.

[0181] If reconstructed samples in the target block's L-shaped neighborhood are available, the target block neighborhood TBN is of the "L-shaped template" type (e.g. Fig. 9As shown in the figure), if only the reconstructed samples in the left part of the target block TB are available (that is, when the TB reaches the top video image border or the search area R, the upper CTU samples are usually unavailable), the target block neighborhood TBN is of the "left template" type, if only the reconstructed samples in the upper part of the target block TB are available (that is, when the target block TB reaches the left video image border or the search area R, the left CTU samples are usually unavailable), the target block neighborhood TBN is of the "upper template" type, or if no reconstructed samples are available around the target block (usually when the target block TB reaches the upper left corner of the video image or the search area R), the target block neighborhood TBN is of the "no template" type.

[0182] If the target block neighborhood TBN is of the “no template” type, then in step 330, the prediction block used to predict the target block TB is set to a DC (direct coding) value, for example, 1<<(1-bitdepth Y ), where bitdepth Y Indicates the bit depth of luma samples of a video image.

[0183] In step 430, a candidate reference block RB located at the current position is considered and a candidate reference template patch RTP is associated with the candidate reference block RB. The candidate reference template RTP is constructed from reconstructed samples in the candidate reference block neighborhood RBN.

[0184] A candidate reference block neighborhood RBN is associated with the candidate reference block RB. The candidate reference block neighborhood RBN defines the position of the reconstructed samples in the search area in the neighborhood of the candidate reference block RB. The candidate reference block neighborhood RBN is defined based on the target block neighborhood TBN, i.e. the position defined by the candidate reference block neighborhood RBN is defined based on the position of the target block neighborhood TBN and the geometric transformation type GTT (see, for example, Fig.11 and 12 ).

[0185] The geometric transformation type GTT identifies the geometric transformation GT.

[0186] A candidate reference template patch RTP associated with the candidate reference block RB is obtained. The candidate reference template patch RTP includes reconstructed samples. The reconstructed samples are obtained by applying a geometric transformation GT identified by a geometric transformation type GTT to reconstructed samples in a candidate reference block neighborhood RBN.

[0187] For example, if the target block neighborhood TBN indicates that the upper and left adjacent reconstructed samples of the target block TB are available, the target template patch TTP includes the upper and left adjacent reconstructed samples of the target block TB. If the geometric transformation GT is a horizontal flip, the transformed reconstructed sample positions correspond to the upper and right reconstructed samples in the candidate reference block neighborhood RBN.

[0188] In another example, if the target block neighborhood TBN indicates that the upper and left adjacent reconstructed samples of the target block TB are available, the target template patch TTP includes the upper and left adjacent reconstructed samples of the target block TB. If the geometric transformation GT is a vertical flip, the transformed reconstructed sample positions correspond to the bottom and left reconstructed samples in the candidate reference block neighborhood RBN.

[0189] The reconstructed samples of the target block and the reconstructed samples of the candidate reference template patch RTP do not directly correspond (ie are not co-located with respect to the reference and target block neighborhoods) because the geometric transformation GT has been applied to the reconstructed sample locations.

[0190] When the geometric transformation GT is horizontal flip, Fig.10 An example of reordered transformed samples is given in step 3 of .

[0191] In step 440 , a cost is associated with the candidate reference block RB by comparing the reconstructed samples of the candidate reference template patch RTP with the reconstructed samples of the target template patch TTP.

[0192] In one exemplary embodiment, the cost is evaluated by the sum of absolute differences (SAD) between the reconstructed samples of the reference template patch RTP and the reconstructed samples of the target template patch TTP.

[0193] In step 450, a candidate reference block is selected. It corresponds to the candidate reference template patch RTP providing the lowest cost. The selected candidate reference block RB is associated with a geometric transformation type GTT identifying a geometric transformation GT to be applied on the reconstructed sample positions.

[0194] In step 460, an intra prediction block of the target block TB is generated by applying a geometric transform GT identified by a geometric transform type GTT associated with the selected candidate reference block RB to the reconstructed samples of the selected candidate reference block RB.

[0195] In an exemplary embodiment of the method 400 , in step 470 , a geometric transform type GTT associated with the selected candidate reference block RB is signaled in the bitstream B.

[0196] In an exemplary embodiment, the geometric transformation type indicates at least one of the following geometric transformations or a combination of at least two of the following geometric transformations:

[0197] -Flip horizontally;

[0198] -Flip vertically;

[0199] - Rotation;

[0200] -Identity.

[0201] In an exemplary embodiment, the rotation angle of the rotation is a multiple of 90°.

[0202] In some embodiments (with high-end processing capabilities), combining rotation and flip transformations is advantageous because it produces geometric transformations that cannot be produced by pure combinations of flips or rotations, increasing the number of candidates, ie improving signal adaptation.

[0203] Fig.10 The method 400 is illustrated when the geometric transformation is a horizontal flip according to an exemplary embodiment.

[0204] In step 1, the target block neighborhood is analyzed to determine the space of the search region R and identify available reconstructed samples in the target block neighborhood TBN.

[0205] In the example, only reconstructed samples located in the left neighbourhood of the target block TB are available. The target block neighbourhood TBN is of the "left template" type (a strip of 4 samples next to the target block TB).

[0206] In step 2, the sample availability and the geometric transformation type GTT (see Fig.11 ) to determine the candidate reference block neighborhood RBN, so that in this example, considering the horizontal flip GTT, the candidate reference block neighborhood RBN is located on the right side of the candidate reference block RB.

[0207] In step 3, a geometric transformation GT associated with a geometric transformation type GTT (here horizontal flip) is applied to the reconstructed samples of the candidate reference block neighborhood RBN. Here, this consists of a horizontal reordering of these sample positions. The candidate reference template patch RTP comprises the reconstructed samples located at the reordered reconstructed sample positions. The output of step 3 produces the samples that constitute the reference template patch RTP.

[0208] like Fig.10 As shown, when horizontal flipping is applied to samples of a candidate reference block RB, the transformed reconstructed samples are reordered according to the distance to the center of the candidate reference block. Fig.10As shown, the 4 reconstructed sample positions with indices 1, 2, 3 and 4 (step 2) are flipped horizontally and correspond to the transformed reconstructed sample position indices 4, 3, 2 and 1 (step 3), respectively. When the geometric transformation is a rotation according to a direction (clockwise or counterclockwise), the sample positions in the candidate reference block neighborhood RBN of the candidate reference block RB samples are reordered according to the direction of the rotation.

[0209] In step 4, for a given target block TB, a cost is associated with each candidate reference block RB by comparing the reconstructed samples of the candidate reference template patch RTP with the reconstructed samples of the target template patch TTP.

[0210] In step 5, a candidate reference block RB associated with the geometric transformation type GTT is selected. The selected candidate reference block corresponds to the candidate reference block RB associated with the minimum cost. The geometric transformation GT identified by the geometric transformation type GTT associated with the selected candidate reference block RB is applied to the reconstructed samples of the selected candidate reference block RB. Here, the selected candidate reference block RB is horizontally flipped, and the intra-frame prediction block of the target block TB is the selected candidate reference block RB after horizontal flipping.

[0211] Fig.11 Some examples of neighborhood template determination according to flipping (geometric transformation type) according to exemplary embodiments are illustrated.

[0212] When the upper and left neighboring samples of the target block TB are available and the geometric transformation GT is an identity function, horizontal flipping, vertical flipping, or horizontal and vertical flipping, the candidate reference block neighborhood RBN indicates an "L-shaped template" whose layout for the candidate reference block RB depends on the flipping mode.

[0213] When only the upper neighboring samples of the target block TB are available and the geometric transformation is an identity function or a horizontal flip, the candidate reference block neighborhood RBN indicates an “upper template” whose layout for the candidate reference block RB depends on the flip mode.

[0214] When only upper neighboring samples of the target block TB are available and the geometric transformation is vertical flipping or horizontal and vertical flipping, the candidate reference block neighborhood RBN indicates a "bottom template" whose layout for the candidate reference block RB depends on the flipping mode.

[0215] When only left neighboring samples of the target block TB are available and the geometric transformation is the identity function or vertical flipping, the candidate reference block neighborhood RBN indicates a "left template" whose layout for the candidate reference block RB depends on the flipping mode.

[0216] When only left neighboring samples of the target block TB are available and the geometric transformation is horizontal flipping or horizontal and vertical flipping, the candidate reference block neighborhood RBN indicates a "right template" whose layout for the candidate reference block RB depends on the flipping mode.

[0217] Fig.12 Some examples of neighborhood template determination according to rotation (geometric transformation type) according to exemplary embodiments are illustrated.

[0218] When the upper and left neighboring samples of the target block TB are available, the candidate reference block neighborhood RBN indicates an “L-shaped template.” The “L-shaped template” layout also depends on the degree of rotation.

[0219] When only the upper neighboring samples of the target block TB are available, if the rotation degree is equal to 0, the candidate reference block neighborhood RBN is the "upper template", if the rotation degree is equal to 90, the candidate reference block neighborhood RB is the "right template", if the rotation degree is equal to 180, the candidate reference block neighborhood RB is the "bottom template", if the rotation degree is equal to 270, the candidate reference block neighborhood RB is the "left template".

[0220] When only the left neighboring samples of the target block TB are available, if the rotation degree is equal to 0, the candidate reference block neighborhood RBN is the "left template", if the rotation degree is equal to 90, the candidate reference block neighborhood RBN is the "upper template", if the rotation degree is equal to 180, the candidate reference block neighborhood RBN is the "right template", if the rotation degree is equal to 270, the candidate reference block neighborhood RBN is the "bottom template".

[0221] In general, the present invention relates to a method for intra-predicting a block of a video image according to a first intra-prediction mode using a first geometric transform identified by a first geometric transform type. The method determines whether to predict at least one neighboring block according to a second intra-prediction mode using a second geometric transform identified by a second geometric transform type, and if the at least one neighboring block is predicted according to the second intra-prediction mode, inferring the first geometric transform type from the second geometric transform type.

[0222] The present invention is advantageous because it avoids the first geometric transform type signaling, thereby improving the coding efficiency of block prediction based on the first intra prediction mode.

[0223] Fig.13 A block diagram of a method 500 of intra-predicting a block of a video image according to a first intra-prediction mode using a first geometric transformation according to an exemplary embodiment is schematically illustrated.

[0224] In step 510 , the method 500 determines whether to predict at least one neighboring block according to a second intra prediction mode using a second geometric transform identified by a second geometric transform type.

[0225] In step 520, if at least one neighboring block is predicted according to the second intra prediction mode, a first geometric transform type is inferred from a second geometric transform type that identifies a geometric transform used by the second intra prediction mode for predicting the at least one neighboring block.

[0226] In an exemplary embodiment, the first geometric transformation type and the second geometric transformation type identify at least one of the following geometric transformations or a combination of at least two of the following geometric transformations:

[0227] -Flip horizontally;

[0228] -Flip vertically;

[0229] - Horizontal and vertical flip;

[0230] - Rotation;

[0231] -Identity.

[0232] Combining rotation and flip transformations is advantageous because a combination of several flips can be replaced by a rotation, thereby reducing the complexity of the geometric transformation to reconstruct the sample.

[0233] In an exemplary embodiment, the rotation angle of the rotation is a multiple of 90°.

[0234] Fig.14 An example of neighboring blocks according to an exemplary embodiment is illustrated.

[0235] The target block C is intended to be predicted according to the first intra prediction mode.

[0236] Neighboring blocks A0, A1, B0, B1, B2, and H are examples of neighboring blocks of the checked target block C. For example, a single block B2 is predicted according to a second intra-frame prediction mode, and a second geometric transformation type used by the second intra-frame prediction mode indicates flipping. Then, a first geometric transformation type used by the first intra-frame prediction mode is inferred from the second geometric transformation type. For example, the first geometric transformation type used by the first intra-frame prediction mode for predicting the target block C is a flip type used by the second intra-frame prediction mode.

[0237] In an exemplary embodiment, a geometric transform type used for an intra prediction mode of a block of a video image is stored in a memory, and a second geometric transform type used by an intra-coded neighboring block is obtained from the memory.

[0238] Fig.15 A block diagram of a method 600 of intra-predicting a block of a video image by using a first intra-prediction mode using a first geometric transformation type according to an exemplary embodiment is schematically illustrated.

[0239] Method 600 is only used by method 100 (encoding).

[0240] In step 510, the method 600 determines whether to predict at least one neighboring block (eg, located at Fig.16 The second intra prediction mode uses a second geometric transformation identified by a second geometric transformation type.

[0241] When none of the neighboring blocks is predicted according to the second intra prediction mode, in step 610, a first geometric transformation type is determined by the first intra prediction mode, and in step 620, the determined first transformation type is signaled in the bitstream.

[0242] In an exemplary embodiment, when the first transform type is a flip type, step 610 is step 330, and when the first intra prediction mode is as described with respect to Figures 8 to 12 When the ITMP mode is discussed, step 610 is method 400.

[0243] In an exemplary embodiment of step 620, signaling the first geometric transform type in the bitstream includes context-based entropy encoding of a binary value (flag) indicating whether the first geometric transform type is signaled in the bitstream, wherein the context-based entropy encoding is performed using a context that depends on the number of neighboring blocks predicted according to the second intra-frame prediction mode.

[0244] When predicting at least one neighboring block according to a second intra-frame prediction mode, the first geometric transform type is not signaled in the bitstream, but in step 520, the first geometric transform type is inferred from a second geometric transform type, which identifies the geometric transform used by the second intra-frame prediction mode for predicting the at least one neighboring block.

[0245] In an exemplary embodiment, step 520 includes sub-steps 521 and 522 .

[0246] In sub-step 521, a second geometric transform type having a greater number of occurrences is determined in a second geometric transform used by a second intra prediction mode for predicting the at least one neighboring block.

[0247] In step 522, the first geometric transformation type is set to a second geometric transformation type having a greater number of occurrences.

[0248] In an exemplary embodiment of step 520, when predicting a single neighboring block according to a second intra prediction mode, the first geometric transformation type is set to a second geometric transformation type associated with the second intra prediction mode.

[0249] In an exemplary embodiment, when at least one neighboring block is predicted according to the second intra-frame prediction mode, in step 530, an index is signaled in the bitstream, the index indicating at least one neighboring block predicted by the second intra-frame prediction mode using the second geometric transform type having a greater number of occurrences.

[0250] exist Fig.15 In a variation of Fig.16 As illustrated, if the second geometric transform type with more occurrences (step 521) is not identical to the first geometric transform type derived from the first intra prediction mode (step 610) (step 630), then in step 620, the method 700 signals the first geometric transform type in the bitstream. If the second geometric transform type with more occurrences is identical to the first geometric transform type, the method 700 sets the first geometric transform type to the second geometric transform type with more occurrences (sub-step 522).

[0251] In an exemplary embodiment, the second geometric transform type used by the second intra prediction mode can be obtained from the bitstream or derived from the second intra prediction mode.

[0252] In an exemplary embodiment, when the first geometric transform type determined from the first intra prediction mode is different from the inferred first geometric transform type, the inferred first geometric transform type is signaled in the bitstream (step 620).

[0253] Optionally, the inferred first geometry type is never transmitted and is inferred from a geometry transform type used by a second geometry transform mode for predicting neighboring blocks.

[0254] Fig.17 A block diagram of a method 800 of intra-predicting a block of a video image by using a first intra-prediction mode using a first geometric transformation type according to an exemplary embodiment is schematically illustrated.

[0255] In step 810, if the first neighboring block (eg, Fig.14 If intra prediction is performed in step 820, the first geometry type used by the first intra prediction mode indicates the geometric transformation T1 used to predict block C.

[0256] Otherwise, in step 830, if the second neighboring block (eg, Fig.14 If intra prediction is performed by B1) in the first intra prediction mode, then in step 840, the first geometry type used by the first intra prediction mode indicates the geometric transformation T2 used to predict block C.

[0257] Otherwise, in step 850, if the first neighboring block is intra predicted according to the second intra prediction mode using the first geometric transform type TA1, if the second neighboring block is intra predicted according to the second intra prediction mode using the second geometric transform type TA2, and if the third neighboring block (e.g., Fig.14 If intra prediction is performed by the first intra prediction mode, then in step 860, the first geometric transformation type used by the first intra prediction mode indicates the geometric transformation indicated by the third geometric transformation type TB2. Otherwise, in step 870, the first geometric transformation type is determined by the first intra prediction mode and the first geometric transformation type is signaled in the bitstream, or, alternatively, a second geometric transformation type having a greater number of occurrences is determined from the second geometric transformation type used by the second intra prediction mode for predicting the at least one neighboring block, and the first geometric transformation type is inferred from the second geometric transformation type having a greater number of occurrences (steps 521, 522).

[0258] In an exemplary embodiment, the first intra prediction mode is an RR-IBC mode using a first geometric transform type as discussed above, the first geometric transform type indicating at least one of the following geometric transforms or a combination of at least two of the following geometric transforms:

[0259] -Flip horizontally;

[0260] -Flip vertically;

[0261] - Horizontal and vertical flip;

[0262] - Rotation;

[0263] - identity;

[0264] And the second intra prediction is as follows Figures 8 to 12 The ITMP model discussed.

[0265] Fig.18 A block diagram schematically illustrates a method 900 of intra-predicting a block of a video image when a first intra-prediction mode is an RR-IBC mode and a second intra-prediction mode is an ITMP mode according to a first exemplary embodiment.

[0266] In step 510, method 900 determines whether to predict at least one neighboring block (eg, located at Fig.14 Position A0, A1, B0, B1, B2 or H).

[0267] If at least one neighboring block is predicted according to the ITMP mode, then in step 910, the ITMP predicted neighboring block is determined and is intended to be used as a block vector prediction (BVP) candidate to select the RR-IBC mode for predicting the block of the video image. Therefore, the RR-IBC mode considers the ITMP predicted neighboring block as a BVP candidate outside the current RR-IBC neighboring coding block (RR-IBC merge).

[0268] In an exemplary embodiment of step 910, the ITMP predicted neighboring block is determined based on a geometric transform type that identifies a geometric transform type having a greater number of occurrences among the geometric transforms used by the ITMP intra-frame prediction mode used to predict the at least one neighboring block (step 521), and the ITMP predicted neighboring block is determined based on the ITMP intra-frame prediction mode that uses the geometric transform identified by the geometric transform type having a greater number of occurrences.

[0269] In step 920, the method 900 checks whether the selected RR-IBC mode uses the block vector prediction (BVP) candidate.

[0270] When the selected RR-IBC mode does not use the block vector prediction (BVP) candidate, in step 930, a best first geometric transform type is selected among the set of first geometric transforms, and in step 940, the best first geometric transform type is signaled in the bitstream.

[0271] In a variant, the set of first geometric transformations is restricted to ITMP geometric transformations of neighboring blocks.

[0272] The variant reduces the number of candidates, thereby reducing complexity.

[0273] When the selected RR-IBC mode uses the block vector prediction (BVP) candidate (step 910), the first geometric transform type used in the RR-IBC mode for predicting the block of the video image is set to the second geometric transform type used by the ITMP mode for predicting the ITMP-predicted neighboring block in step 950. Therefore, only a merge index is carried, and the merge index indicates both the BVP (ITMP-predicted neighboring block) and the associated geometric transform type.

[0274] Fig.19Schematically illustrates Fig.18 Block diagram of a variation 1000 of method 900 .

[0275] When the selected RR-IBC mode does not use the block vector prediction (BVP) candidate, a second geometric transform type having a greater number of occurrences is determined in a second geometric transform used by the ITMP mode for predicting the at least one neighboring block (step 521), and a first geometric transform type used in the RR-IBC mode for predicting blocks of a video image is set to the second geometric transform type having a greater number of occurrences (step 522).

[0276] In one variant, when RR-IBC merging is considered, the merge index indicates both the BVP and the associated geometry transform type, but the HMVP only stores the block vector prediction (and not the associated geometry transform type).

[0277] In a variation of the first exemplary embodiment, the merge index indicates only the BVP, and the HMVP stores both the block vector prediction and the associated geometric transform type.

[0278] In an exemplary embodiment, the first intra prediction mode is as follows Figures 8 to 12 The ITMP mode discussed, the second intra prediction mode is the RR-IBC mode using the first geometric transform type as described above, the first geometric transform type indicating at least one of the following geometric transforms or a combination of at least two of the following geometric transforms:

[0279] -Flip horizontally;

[0280] -Flip vertically;

[0281] - Horizontal and vertical flip;

[0282] - Rotation;

[0283] -Identity.

[0284] like Figures 8 to 12 As discussed in, at the decoder, the ITMP mode signals the selected geometric transform type, but increases bandwidth (signal loss), or performs a template-based matching search for different flip types, but increases decoder complexity. Inheriting the first geometric transform type from a neighboring block accumulates advantages in terms of signaling and complexity at both the encoder and the decoder.

[0285] Fig. 20 A block diagram schematically illustrates a method 1100 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment.

[0286] The approach is symmetric at both the encoder and decoder side.

[0287] In step 510 , the method 1100 determines whether at least one neighboring block (eg, in blocks at locations A0 , A1 , B0 , B1 , B2 ) is predicted according to the RR-IBC mode using a second geometric transform identified by a second geometric transform type.

[0288] If no neighboring block is predicted by RR-IBC mode, then Figures 8 to 12 The geometric transform type candidates discussed in perform a template matching search to predict the geometric transform type used by the ITMP mode for predicting a block of a video image.

[0289] In short, in steps 430 and 440, a cost is associated with each candidate reference block and each candidate reference block is associated with a geometric transformation type. In step 450, a candidate reference block is selected. It corresponds to the candidate reference template patch RTP that provides the lowest cost. The selected candidate reference block RB is associated with a geometric transformation type that identifies the geometric transformation applied to reconstruct the sample position.

[0290] If at least one of the neighboring blocks is predicted using the RR-IBC mode, then in step 1110, method 1100 obtains a geometric transform type used by the RR-IBC mode used to predict the at least one neighboring block and uses the geometric transform type as a geometric transform type candidate.

[0291] In an exemplary embodiment of step 1110, the geometric transform type used by the RR-IBC mode for predicting the at least one neighboring block is a geometric transform type with a greater number of occurrences used by the RR-IBC prediction mode for predicting the at least one neighboring block.

[0292] This exemplary embodiment is advantageous because no signaling is required about the geometric transformation type of the ITMP predicted block C, since the candidates are inferred from the neighboring blocks. Furthermore, the complexity is statistically reduced when the RR-IBC predicted neighboring blocks do not use certain flip types (e.g., the neighboring blocks at positions A1 and B1 are predicted with RR-IBC and use the horizontal flip type). Then, among the four possible flip patterns, for such located block C identities, only the identity and the horizontal flip type are checked.

[0293] In a variant, the set of candidate geometric transform types from which the geometric transform type for predicting the RR-IBC of block C is determined is restricted to the identity mode and the flipping type with a greater number of occurrences used by the predicted RR-IBC for predicting neighboring blocks.

[0294] Fig.21 A block diagram schematically illustrates a method 1200 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment.

[0295] Method 1200 is used only at the encoder.

[0296] For example Figures 8 to 12 A template matching search is performed for the candidates discussed in order to predict the type of geometric transformation used by the ITMP mode for predicting a block of a video image. In short, in steps 430 and 440, a cost is associated with each candidate reference block and each candidate reference block is associated with a geometric transformation type. In step 450, the candidate reference block associated with the geometric transformation type is selected. It corresponds to the candidate reference template patch RTP that provides the lowest cost. The selected candidate reference block RB is associated with a geometric transformation type that identifies the geometric transformation applied to the reconstructed sample position.

[0297] In step 510 , the method 1100 determines whether at least one neighboring block (eg, in a block at position A0 , A1 , B0 , B1 , B2 ) is predicted according to the RR-IBC mode using a second geometric transform identified by a second geometric transform type.

[0298] If none of the neighboring blocks is predicted by RR-IBC mode, method 1200 ends and the geometry type associated with the selected candidate reference block RB is used as the geometry transform type for RR-IBC of predicting block C.

[0299] If at least one of the neighboring blocks is predicted using the RR-IBC mode, then in step 1110, method 1200 obtains a geometric transform type used by the RR-IBC mode used to predict the at least one neighboring block, and uses the geometric transform type as a geometric transform type candidate.

[0300] In step 1210, the method 1200 checks whether the geometry type associated with the selected candidate reference block RB and the obtained geometry transform type used by the RR-IBC mode (step 1110) are the same.

[0301] If the geometry type associated with the selected candidate reference block RB and the obtained geometry transform type used by the RR-IBC mode (step 1110 ) are different, then in step 1220 , the selected candidate reference block RB is signaled in the bitstream.

[0302] If the geometry type associated with the selected candidate reference block RB is the same as the obtained geometry transformation type used by the RR-IBC mode (step 1110), the geometry transformation type used by the ITMP mode for the prediction block is inferred to be the obtained geometry transformation type used by the RR-IBC mode (step 1110).

[0303] Fig. 22 A block diagram schematically illustrates a method 1300 of intra-predicting a block of a video image when a first intra-prediction mode is an ITMP mode and a second intra-prediction mode is an RR-IBC mode according to an exemplary embodiment.

[0304] Method 1300 is used only at the encoder.

[0305] In step 1310, the method 1300 checks whether the geometric transform type of the ITMP used to predict the block of the video picture is signaled in the bitstream.

[0306] If a geometric transform type is signaled, then in step 1320, the geometric transform identified by the signaled geometric transform type is used as a candidate in the template matching search. Figures 8 to 12 A template matching search is performed for the candidate discussed in to predict the type of geometric transformation used by the ITMP mode used to predict the block of the video image (steps 430, 440, 450). In step 1330, the block of the video image is predicted according to the ITMP mode of the geometric transformation identified by the selected geometric transformation type provided by the template matching search.

[0307] If the geometric transform type is not signaled, then in step 510, method 1300 determines whether to predict at least one neighboring block (e.g., in blocks at positions A0, A1, B0, B1, B2) according to the RR-IBC mode using a second geometric transform identified by the second geometric transform type.

[0308] If no neighboring block is predicted by RR-IBC mode, then Figures 8 to 12 A template matching search is performed on the candidate discussed in order to predict the type of geometric transformation used by the ITMP mode for predicting a block of a video image (steps 430, 440, 450).

[0309] Alternatively, if none of the neighboring blocks is predicted by RR-IBC mode, the geometric transformation type indicates identity. This alternative reduces the complexity of the decoder compared to using a template matching search.

[0310] If at least one of the neighboring blocks is predicted using the RR-IBC mode, then in step 1110, method 1300 obtains the geometric transformation type used by the RR-IBC mode used to predict the at least one neighboring block, and sets the geometric transformation type used by the ITMP mode used to predict the block of the video image to the obtained geometric transformation type.

[0311] Fig.23 A schematic block diagram illustrating an example of a system 1400 in which various aspects and exemplary embodiments are implemented is shown.

[0312] The system 1400 may be embedded as one or more devices, including various components described below. In various exemplary embodiments, the system 1400 may be configured to implement one or more aspects described in this application.

[0313] Examples of equipment that may constitute all or part of system 1400 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMD, perspective glasses), projectors (projectors), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors for processing outputs from video decoders, pre-processors for providing inputs to video encoders, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or network servers), static or video cameras, encoding or decoding chips, or any other communication devices. The elements of system 1400 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one exemplary embodiment, the processing and encoder / decoder elements of system 1400 may be distributed across multiple ICs and / or discrete components. In various exemplary embodiments, system 1400 may be coupled to other similar systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports.

[0314] The system 1400 may include at least one processor 1410 configured to execute instructions loaded therein for implementing, for example, various aspects described in the present application. The processor 1410 may include embedded memory, input-output interfaces, and various other circuits known in the art. The system 1400 may include at least one memory 1420 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1400 may include a storage device 1440, which may include a non-volatile memory and / or a volatile memory, including but not limited to an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a disk drive, and / or an optical drive. As a non-limiting example, the storage device 1440 may include an internal storage device, an attached storage device, and / or a network accessible storage device.

[0315] System 1400 may include an encoder / decoder module 1430, which is configured to, for example, process data to provide encoded / decoded video image data, and the encoder / decoder module 1430 may include its own processor and memory. The encoder / decoder module 1430 may represent a (one or more) module that may be included in a device to perform encoding and / or decoding functions. As known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 1430 may be implemented as a separate element of the system 1400, or may be incorporated into the processor 1410 as a combination of hardware and software known to those skilled in the art.

[0316] Program code to be loaded onto the processor 1410 or the encoder / decoder 1430 to perform various aspects described in the present application may be stored in the storage device 1440 and subsequently loaded onto the memory 1420 to be executed by the processor 1410. According to various exemplary embodiments, during the execution of the processes described in the present application, one or more of the processor 1410, the memory 1420, the storage device 1440, and the encoder / decoder module 1430 may store one or more of various items. Such stored items may include, but are not limited to, video image data, information data for encoding / decoding video image data, bit streams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic processing.

[0317] In several exemplary embodiments, memory internal to the processor 1410 and / or encoder / decoder module 1430 may be used to store instructions and provide working memory for processes that may be performed during encoding or decoding.

[0318] However, in other exemplary embodiments, memory external to the processing device (e.g., the processing device may be the processor 1410 or the encoder / decoder module 1430) is used for one or more of these functions. The external memory may be a memory 1420 and / or a storage device 1440, such as a dynamic volatile memory and / or a non-volatile flash memory. In several exemplary embodiments, an external non-volatile flash memory is used to store an operating system for the television. In at least one exemplary embodiment, a fast external dynamic volatile memory such as RAM may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), AVC, HEVC, EVC, VVC, AV1, and the like.

[0319] As indicated in block 1490, input to the elements of system 1400 may be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that may receive an RF signal transmitted over the air, for example, by a broadcast device, (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) when the present invention is implemented in the automotive field, a bus such as a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO17458), or Ethernet (ISO / IEC 802-3) bus.

[0320] In various exemplary embodiments, the input device of block 1490 has associated corresponding input processing elements, as known in the art. For example, the RF portion may be associated with elements necessary for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) band-limiting the frequency band again to a narrower frequency band to select a signal band that may be referred to as a channel in certain exemplary embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF portion of various exemplary embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or a baseband.

[0321] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.

[0322] Various exemplary embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.

[0323] Adding components may include inserting components between existing components, such as, for example, inserting an amplifier and an analog-to-digital converter.In various exemplary embodiments, the RF portion may include an antenna.

[0324] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 1400 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or in the processor 1410 when necessary. Similarly, various aspects of USB or HDMI interface processing may be implemented in a separate interface IC or in the processor 1410 when necessary. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, the processor 1410 and the encoder / decoder 1430, which operate in conjunction with memory and storage elements to process the data streams for presentation on output devices when necessary.

[0325] The various elements of the system 1400 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and data transferred between them using a suitable connection arrangement 1490, such as an internal bus (including an I2C bus), wiring, and printed circuit boards as known in the art.

[0326] The system 1400 may include a communication interface 1450 that enables communication with other devices via a communication channel 1451. The communication interface 1450 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 1451. The communication interface 1450 may include, but is not limited to, a modem or a network card, and the communication channel 1451 may be implemented, for example, within a wired and / or wireless medium.

[0327] In various exemplary embodiments, a Wi-Fi network such as IEEE 802.11 may be used to stream data to the system 1400. The Wi-Fi signals of these exemplary embodiments may be received via a communication channel 1451 suitable for Wi-Fi communications and a communication interface 1450. The communication channel 1451 of these exemplary embodiments may typically connect to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications.

[0328] Other exemplary embodiments may provide streamed data to the system 1400 using a set top box that delivers the data through an HDMI connection to the input block 1490 .

[0329] Still other exemplary embodiments may provide streaming data to the system 1400 using an RF connection to input block 1490 .

[0330] The streamed data may be used as a means of signaling information used by the system 1400. The signaling information may include the bitstream B and / or information such as the number of pixels of a video image and / or any encoding / decoding setting parameters.

[0331] It should be appreciated that signaling may be implemented in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder.

[0332] The system 1400 may provide output signals to various output devices, including a display 1461, speakers 1471, and other peripherals 1481. In various examples of the exemplary embodiments, the other peripherals 1481 may include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 1400.

[0333] In various exemplary embodiments, control signals may be communicated between the system 1400 and the display 1461, speakers 1471, or other peripherals 1481 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.

[0334] Output devices may be communicatively coupled to system 1400 through respective interfaces 1460 , 1470 , and 1480 via dedicated connections.

[0335] Optionally, output devices may be connected to the system 1400 using the communication channel 1451 via the communication interface 1450. The display 1461 and the speaker 1471 may be integrated into a single unit with the other components of the system 1400 in an electronic device such as, for example, a television.

[0336] In various exemplary embodiments, the display interface 1460 may include a display driver such as, for example, a timing controller (T Con) chip.

[0337] For example, if the RF portion of input 1490 is part of a separate set-top box, then display 1461 and speaker 1471 may optionally be separate from one or more of the other components. In various exemplary embodiments where display 1461 and speaker 1471 may be external components, output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.

[0338] exist Figures 1 to 23 In the present invention, various methods are described herein, and each method includes one or more steps or actions to implement the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.

[0339] Some examples are described about block diagrams and / or operational flow charts. Each square block represents a portion of a circuit element, module or code, which includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that, in other embodiments, the (one or more) functions marked in the square block may not occur in the order indicated. For example, depending on the functions involved, two square blocks shown in succession can actually be executed substantially concurrently, or sometimes these square blocks can be executed in reverse order.

[0340] The embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if only discussed in the context of a single form of embodiment (e.g., discussed only as a method), the embodiments of the features discussed may also be implemented in other forms (e.g., an apparatus or a computer program).

[0341] The method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.

[0342] In addition, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values ​​generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having a computer-readable program code implemented thereon that can be executed by a computer. Considering the inherent ability to store information therein and the inherent ability to provide information retrieval therefrom, the computer-readable storage medium used herein can be considered as a non-transitory storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or device, or any suitable combination of the foregoing. It should be appreciated that although more specific examples of computer-readable storage media to which the present exemplary embodiment can be applied are provided below, as those of ordinary skill in the art will readily recognize, it is merely an illustrative and non-exhaustive list: portable computer floppy disk; hard disk; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable compact disk read-only memory (CD-ROM); optical storage device; magnetic storage device; or any suitable combination of the foregoing.

[0343] The instructions may form an application program tangibly embodied on a processor-readable medium.

[0344] For example, instructions may be in hardware, firmware, software, or a combination. For example, instructions may be found in an operating system, a separate application, or a combination of both. Thus, a processor may be characterized as, for example, a device configured to perform a process and a device including a processor-readable medium (such as a storage device) having instructions for performing the process. Additionally, in addition to or in lieu of instructions, a processor-readable medium may store data values ​​generated by an embodiment.

[0345] Device can be realized in suitable hardware, software and firmware for example.The example of such device comprises personal computer, laptop computer, smart phone, tablet computer, digital multimedia set-top box, digital television receiver, personal video recording system, connected household appliances, head-mounted display device (HMD, perspective glasses), projector (projector), "cave" (system including multiple displays), server, video encoder, video decoder, post-processor for processing the output from video decoder, pre-processor for providing input to video encoder, web server, set-top box, and any other equipment for processing video image, or other communication equipment.It should be clear that equipment can be mobile and even installed in a mobile vehicle.

[0346] The computer software may be implemented by the processor 1410 or by hardware, or by a combination of hardware and software. As a non-limiting example, the exemplary embodiments may also be implemented by one or more integrated circuits. The memory 1420 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). The processor 1410 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.

[0347] As will be apparent to one of ordinary skill in the art, embodiments may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bit stream of the described exemplary embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

[0348] The terms used herein are only used to describe the purpose of specific exemplary embodiments and are not intended to be limited. As used herein, the singular "a / kind (a)", "an / kind (an)" and "the / said (the)" may also be intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the terms "include / comprise" and / or "including / comprising" may specify the existence of stated, for example, features, integers, steps, operations, elements and / or components, but do not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. Moreover, when an element is referred to as "response" or "connection" or "associated" to another element, it may directly respond or be connected to another element or be associated with another element, or there may be an intermediate element. On the contrary, when an element is referred to as "direct response" or "direct connection" to another element or "directly associated" with another element, there is no intermediate element.

[0349] It should be appreciated that use of any of the symbols / terms " / ", "and / or", and "at least one of" may be intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B), such as in the case of "A / B", "A and / or B", and "at least one of A and B". As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This may be extended to as many items as listed, as will be apparent to one of ordinary skill in this and related arts.

[0350] Various numerical values ​​may be used in the present application. Specific values ​​may be used for example purposes and the described aspects are not limited to these specific values.

[0351] It will be understood that although the terms first, second, etc. can be used to describe various elements in this article, these elements are not limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the teaching of this application, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. There is no suggestion of sorting between the first element and the second element.

[0352] References to "an exemplary embodiment" or "exemplary embodiments" or "one implementation" or "implementation" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in conjunction with an exemplary embodiment / implementation) is included in at least one exemplary embodiment / implementation. Thus, the appearances of the phrases "in an exemplary embodiment" or "in an exemplary embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places in this application are not necessarily all referring to the same exemplary embodiment.

[0353] Similarly, references herein to "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an exemplary embodiment / example / implementation) may be included in at least one exemplary embodiment / example / implementation. Therefore, the expressions "according to an exemplary embodiment / example / implementation" or "in an exemplary embodiment / example / implementation" appearing in various places in this application do not necessarily all refer to the same exemplary embodiment / example / implementation, nor are separate or alternative exemplary embodiments / examples / implementations necessarily mutually exclusive of other exemplary embodiments / examples / implementations.

[0354] Reference numerals appearing in the claims are for illustration purposes only and have no limiting effect on the scope of the claims.The present exemplary embodiments / examples and variants may be employed in any combination or sub-combination although not explicitly described.

[0355] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.

[0356] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction to the depicted arrows.

[0357] Various embodiments relate to decoding. As used in this application, "decoding" may encompass, for example, all or part of a process performed on a received video image (which may include a received bitstream encoding one or more video images) to produce a final output suitable for display or further processing in a reconstructed video domain. In various exemplary embodiments, such a process includes one or more of the processes typically performed by a decoder. In various exemplary embodiments, for example, such a process also or alternatively includes a process performed by a decoder of the various embodiments described in this application.

[0358] As a further example, in one exemplary embodiment, "decoding" may refer only to dequantization, in one exemplary embodiment, "decoding" may refer to entropy decoding, in another exemplary embodiment, "decoding" may refer only to differential decoding, and in another exemplary embodiment, "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Whether the phrase "decoding process" may be intended to specifically refer to a subset of operations, or to generally refer to a broader decoding process will be clear, and is believed to be well understood by those skilled in the art, based on the context of the particular description.

[0359] Various embodiments are all related to encoding. In a manner similar to the above discussion about "decoding", "encoding" as used in this application can cover, for example, all or part of a process performed on an input video image to produce an output bit stream. In various exemplary embodiments, such a process includes one or more of the processes typically performed by an encoder. In various exemplary embodiments, such a process also includes or optionally includes a process performed by an encoder of the various embodiments described in this application.

[0360] As a further example, in one exemplary embodiment, "encoding" may refer only to quantization, in one exemplary embodiment, "encoding" may refer only to entropy coding, in another exemplary embodiment, "encoding" may refer only to differential coding, and in another exemplary embodiment, "encoding" may refer to a combination of quantization, differential coding, and entropy coding. Based on the context of the particular description, whether the phrase "encoding process" may be intended to specifically refer to a subset of operations, or to generally refer to a broader encoding process will be clear and is believed to be well understood by those skilled in the art.

[0361] Additionally, the present application may refer to "obtaining" various information. Obtaining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0362] Additionally, this application may refer to "receiving" various information. Receiving information may include, for example, one or more of accessing the information or receiving the information from a communication network.

[0363] Moreover, as used herein, the word "signal" especially refers to indicating something to a corresponding decoder, etc. For example, in some exemplary embodiments, the encoder signals specific information, such as encoding parameters or encoded video image data. In this way, in exemplary embodiments, the same parameter can be used on the encoder side and the decoder side. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit saving is achieved in various exemplary embodiments. It should be recognized that signaling can be completed in a variety of ways. For example, in various exemplary embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signal" is mentioned above, the word "signal" can also be used as a noun in this article.

[0364] A number of embodiments have been described. However, it should be understood that various modifications may be made. For example, elements of different embodiments may be combined, supplemented, modified, or removed to produce other embodiments. In addition, it will be understood by those of ordinary skill that other structures and processes may replace the disclosed structures and processes, and that the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) manners to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, the present application contemplates these and other embodiments.

Claims

1. A method for intra-predicting a block of a video image according to a first intra-prediction mode using a first geometric transform identified by a first geometric transform type, wherein the method include: - determining (510) whether to predict at least one neighboring block according to a second intra prediction mode using a second geometric transform identified by a second geometric transform type; as well as - if at least one neighboring block is predicted according to the second intra prediction mode, inferring (520) the first geometric transformation type from the second geometric transformation type.

2. The method according to claim 1, wherein the first geometric transformation type and / or the second geometric transformation type identifies at least one of the following geometric transformations or a combination of at least two of the following geometric transformations: -Flip horizontally; -Flip vertically; - Horizontal and vertical flip; - Rotation; -Identity.

3. The method according to claim 1 or 2, wherein when none of the neighboring blocks is predicted according to the second intra-frame prediction mode, the first geometric transformation type is determined (610) by the first intra-frame prediction mode.

4. The method of claim 3, wherein the determined first transform type is signaled (620) in a bitstream.

5. The method of claim 3, wherein the first geometric transform type is selected from among second geometric transform types, the second geometric transform type identifying a second geometric transform used by a second intra-prediction mode of a neighboring block.

6. The method according to claim 1 , wherein the first geometric transformation type is inferred ( 520 ) from the second geometric transformation type. include: - determining (521) a second geometric transform type having a greater number of occurrences in a second geometric transform used by said second intra prediction mode for predicting said at least one neighboring block; as well as - Setting (522) the first geometric transformation type to the second geometric transformation type having a greater number of occurrences.

7. The method of claim 1 , wherein when predicting a single neighboring block according to a second intra-prediction mode, inferring ( 520 ) the first geometric transformation type from the second geometric transformation type comprises setting the first geometric transformation type to the second geometric transformation type associated with the second intra-prediction mode.

8. The method according to claim 6 or 7, wherein when predicting the at least one neighboring block according to the second intra-frame prediction mode, the method further comprises signaling (620) an index in the bitstream, the index indicating at least one neighboring block predicted by the second intra-frame prediction mode using the second geometric transform type having a greater number of occurrences.

9. The method according to claims 6 to 8, wherein if the second geometric transform type having a greater number of occurrences is not the same as the first geometric transform type derived from the first intra-frame prediction mode, the method further comprises signaling (620) the first geometric transform type in the bitstream.

10. The method according to one of claims 1 to 8, wherein the second geometric transform type used by the second intra prediction mode is obtainable from a bitstream or derived from the second intra prediction mode.

11. The method of one of claims 5 to 10, when the first geometric transform type determined from the first intra prediction mode is different from the inferred first geometric transform type, signaling (620) the inferred first geometric transform type in the bitstream.

12. The method according to any one of claims 1 to 11, wherein the first intra prediction mode is a re-ordered intra block copy (RR-IBC) mode with a first geometric transform type, wherein the first geometric transform type indicates at least one of the following geometric transforms or a combination of at least two of the following geometric transforms: -Flip horizontally; -Flip vertically; - Horizontal and vertical flip; - Rotation; - identity; And the second intra prediction is an intra template matching prediction (ITMP) mode using a second geometric transform type, wherein the second geometric transform type indicates at least one of the following geometric transforms or a combination of at least two of the following geometric transforms: -Flip horizontally; -Flip vertically; - Horizontal and vertical flip; - Rotation; -Identity.

13. The method according to any one of claims 1 to 11, wherein the first intra prediction is an intra template matching prediction mode using a second geometric transformation type, wherein the second geometric transformation type indicates at least one of the following geometric transformations or a combination of at least two of the following geometric transformations: -Flip horizontally; -Flip vertically; - Horizontal and vertical flip; - Rotation; - identity; And the second intra prediction mode is a reordered intra block copy mode using a first geometric transform type, wherein the first geometric transform type indicates at least one of the following geometric transforms or a combination of at least two of the following geometric transforms: -Flip horizontally; -Flip vertically; - Horizontal and vertical flip; - Rotation; -Identity.

14. Method for encoding a block of a video image based on a prediction block of said block of a video image, wherein said prediction block is derived according to a method according to one of claims 1 to 13.

15. Method for decoding a block of a video image based on a prediction block of said block of a video image, wherein said prediction block is derived according to a method according to one of claims 1 to 13.

16. An apparatus comprising means for carrying out one of the methods according to any one of claims 1 to 13.

17. A computer program product comprising instructions which, when the program is executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 13.

18. A non-transitory storage medium carrying program code instructions for executing the method according to any one of claims 1 to 13.