Encoding / decoding video image data
By performing gradient analysis on samples in the template area around the sample block of the video image, the gradient histogram is calculated to determine the intra prediction mode, which solves the problem of high complexity in the video codec in the prior art, and realizes efficient intra prediction.
Patent Information
- Application Number
- CN202380063501.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-04-24
- Publication Date
- 2025-05-27
AI Technical Summary
When using intra prediction modes based on linear models, such as DIMD mode, it is difficult for prior art to maintain encoding efficiency and reduce its complexity in video codecs.
By performing gradient analysis on samples located in the template area defined around the sample block, a gradient histogram is calculated to determine the intra prediction mode for decoding. The method includes filtering the samples of at least one template area, selecting up to two angle intra prediction modes, and determining the intra prediction mode from them.
This method reduces the resource requirements required to export DIMD mode, reduces the complexity of the encoder and decoder, while maintaining the encoding efficiency of the video codec.
Smart Images

Figure CN120051989A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This disclosure claims priority to and the benefit of European Patent Application No. 22306424.7, filed on September 27, 2022, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to video image encoding and decoding. In particular, but not exclusively, the technical field of the present disclosure relates to intra-frame prediction blocks of video images. Background Art
[0004] This section is intended to introduce the reader to various aspects of the art that may be related to various aspects of at least one embodiment of the present disclosure described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present disclosure. Therefore, it should be understood that these statements should be read in this light, and not as admissions of prior art.
[0005] A pixel corresponds to the smallest display unit on the screen, which can be composed of one or more light sources (1 for a monochrome screen and 3 or more for a color screen).
[0006] A video image (also called a frame or image frame) comprises at least one component (also called an image component or channel) determined by a specific image / video format, which specifies all information related to pixel values and all information related to the video image data that can be used by a display unit and / or any other device to display and / or decode the video image data related to the video image.
[0007] A video image comprises at least one component which is typically represented in the shape of an array of samples.
[0008] A monochrome video image includes a single component, while a color video image may include three components.
[0009] For example, a color video image may include a luminance (or luminance) component and two chrominance components when the image / video format is the well-known (Y,Cb,Cr) format, or may include three color components (one for red, one for green, and one for blue) when the image / video format is the well-known (R,G,B) format.
[0010] Each component of the video image may comprise a number of samples relative to the number of pixels of the screen on which the video image is intended to be displayed. In a variant, the number of samples contained in a component may be a multiple (or fraction) of the number of samples contained in another component of the same video image.
[0011] For example, in case the video format includes a luma component and two chroma components (such as a (Y, Cb, Cr) format), the chroma components may contain half the number of samples in width and / or height relative to the luma components, depending on the color format considered.
[0012] A sample is the smallest visual information unit that makes up a component of a video image. A sample value can be, for example, a brightness or chrominance value or a color value in (R, G, B) format.
[0013] The pixel value is the value of a screen pixel. For a monochrome video image, a pixel value can be represented by one sample, and for a color video image, a pixel value can be represented by multiple co-located samples. The co-located sample associated with a pixel refers to a sample corresponding to the position of the pixel in the screen.
[0014] A video image is usually considered as a set of pixel values, with each pixel represented by at least one sample.
[0015] A block of a video image is a group of samples of a component of a video image. When the image / video format is the well-known (Y, Cb, Cr) format, at least one block of luminance samples (referred to as luminance block) or at least one block of chrominance samples (referred to as chrominance block) may be considered, or when the image / video format is the well-known (R, G, B) format, at least one block of color samples may be considered.
[0016] At least one embodiment is not limited to a particular image / video format.
[0017] In state-of-the-art video compression systems such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Generic Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level picture partitioning are provided to divide a video picture into picture regions, so-called coding tree units (CTUs), which for HEVC may typically have a size between 16×16 and 64×64 pixels, while for VVC the size of the coding tree unit (CTU) may be 32×32, 64×64 or 128×128 pixels.
[0018] The CTU partitioning of the video image forms a grid of CTUs of a fixed size, i.e., a CTU grid, in which the upper and left boundaries coincide spatially with the upper and left borders of the video image. The CTU grid represents the spatial partitioning of the video image.
[0019] In VVC and HEVC, the CTU size (CTU width and CTU height) of all CTUs of the CTU grid is equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) may be equal to 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is encoded into the bitstream, for example at the sequence level in a sequence parameter set (SPS).
[0020] The spatial position of a CTU in the CTU grid is determined by the CTU address ctuAddr, which defines the spatial position of the upper left corner of the CTU from the origin. Figure 1 As illustrated, a CTU address may define a spatial position from the upper left corner of a higher-level spatial structure S that contains the CTU.
[0021] A coding tree is associated with each CTU to determine the tree partitioning of the CTU.
[0022] like Figure 1 As shown, in HEVC, the coding tree is a quadtree partition of CTU, where each leaf is called a coding unit (CU). The spatial position of the CU in the video image is defined by the CU index cuIdx, which indicates the spatial position from the upper left corner of the CTU. The CU is spatially partitioned into one or more prediction units (PUs). The spatial position of the PU in the video image VP is defined by the PU index puIdx, which defines the spatial position from the upper left corner of the CTU, and the spatial position of the elements of the partitioned PU is defined by the PU partition index puPartIdx, which defines the spatial position from the upper left corner of the PU. Each PU is assigned some intra-frame or inter-frame prediction data.
[0023] Intra or inter coding modes are assigned on CU level. This means the same intra / inter coding mode is assigned to each PU of a CU, although the prediction parameters vary from PU to PU.
[0024] According to a quadtree called a transform tree, a CU can also be spatially partitioned into one or more transform units (TUs). A transform unit is a leaf of a transform tree. The spatial position of a TU in a video image is defined by a TU index tuIdx, which defines the spatial position from the upper left corner of the CU. Each TU is assigned some transform parameters. The transform type is assigned at the TU level, and a 2D separate transform is performed at the TU level during encoding or decoding of an image block.
[0025] The PU partition types in HEVC are as follows: Figure 2 As shown. They include square partitions (2N×2N and N×N), which are the only partitions used in both intra- and inter-prediction CUs; symmetric non-square partitions (2N×N, N×2N, used only in inter-prediction CUs); and asymmetric partitions (used only in inter-prediction CUs). For example, PU type 2N×nU represents an asymmetric horizontal partition of a PU, where a smaller partition is located at the top of the PU. According to another example, PU type 2N×nL represents an asymmetric horizontal partition of a PU, where a smaller partition is located at the top of the PU.
[0026] like Figure 3 As shown, in VVC, the coding tree starts from the root node (i.e., CTU). Next, a quadtree (or quadtree) partitioning divides the root node into 4 nodes (solid lines) corresponding to 4 sub-blocks of equal size. Next, the quadtree (or quadtree) leaves can then be further partitioned by the so-called multi-type tree, which involves Figure 4 Binary or ternary segmentation of one of the 4 segmentation modes shown. These segmentation types are vertical and horizontal binary segmentation modes, denoted SBTV and SBTH; and vertical and horizontal ternary segmentation modes SPTTV and STTH.
[0027] In the case of a joint coding tree shared by luma and chroma components, the leaves of the coding tree of a CTU are CUs.
[0028] In contrast to HEVC, in VVC, in most cases, CU, PU, and TU have equal sizes, which means that coding units are generally not partitioned into PUs or TUs except in some specific coding modes.
[0029] Figure 5 and Figure 6 An overview of video encoding / decoding methods used in current video standard compression systems (such as, for example, HEVC or VVC) is provided.
[0030] Figure 5 A schematic block diagram showing the steps of a method 100 for encoding a video image VP according to the prior art is shown.
[0031] In step 110, the video picture VP is partitioned into blocks of samples and the partition information data is signaled into the bitstream. Each block comprises samples of one component of the video picture VP. Thus, the blocks comprise samples defining each component of the video picture VP.
[0032] For example, in HEVC, a picture is divided into coding tree units (CTUs). Each CTU can be further subdivided using a quadtree partition, where each leaf of the quadtree represents a coding unit (CU). The partition information data may then include data describing the quadtree subdivision of the CTU and each CTU.
[0033] Then, each block of samples (block for short) may be a CU (if the CU includes a single PU) or a PU of the CU.
[0034] Each current block is encoded along a coding loop (also referred to as "in-loop") using an intra or inter prediction mode.
[0035] Intra prediction (step 120) uses intra prediction data. Intra prediction consists in predicting the current block with the help of an intra prediction block based on already coded, decoded and reconstructed samples located in a so-called L-shaped template area defined around the current block (usually located at the top, left and above the left of the current block). Intra prediction is performed in the spatial domain.
[0036] In inter prediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches for a reference block that is a good predictor of the current block in one or more reference pictures used to predictively encode the current video picture. In unidirectional motion estimation / compensation, the candidate reference blocks belong to a single reference picture of a reference picture list denoted L0 or L1, and in bidirectional motion estimation / compensation, the candidate reference blocks are derived from reference blocks of reference picture list L0 and reference blocks of reference picture list L1.
[0037] For example, a good predictor for the current block is a candidate reference block that is similar to the current block. It may also correspond to a reference block that provides a good compromise between its similarity to the current block and the rate cost of indicating its motion information required for temporal prediction of the current block.
[0038] The output of the motion estimation step 130 is inter-frame prediction data, which includes motion information associated with the current block and other information for obtaining the same prediction block at the encoding / decoding side. Typically, the motion information includes one motion vector and reference image index for unidirectional estimation / compensation, and two motion vectors and two reference image indexes for bidirectional estimation / compensation). Next, motion compensation (step 135) obtains the prediction block with the help of the (one or more) motion vectors and (one or more) reference image indexes determined by the motion estimation step 130. Basically, the reference block belonging to the selected reference image and pointed to by the motion vector can be used as the prediction block of the current block. In addition, since the motion vector is represented as a fraction of an integer pixel position (this is called sub-pixel precision motion vector representation), motion compensation usually involves spatial interpolation of some reconstructed samples of the reference image to calculate the prediction block.
[0039] The prediction information data is signaled into the bitstream.The prediction information may include the prediction mode (intra or inter or skip), intra / inter prediction data and any other information for obtaining the same prediction block at the decoding side.
[0040] The method 100 optimizes the rate-distortion trade-off to select a prediction mode (intra-frame or inter-frame prediction mode) by considering, for example, the encoding of a prediction residual block calculated by subtracting a candidate prediction block from a current block, and the signaling of the prediction information data required to determine the candidate prediction block at the decoding side.
[0041] Typically, the best prediction mode is given as the prediction mode of the best coding mode p* for the current block given by the following equation:
[0042]
[0043] Where P is the set of all candidate coding modes for the current block, p represents the candidate coding mode in the set, and RD cost (p) is the rate-distortion cost of candidate coding mode p, usually expressed as:
[0044] RD cost(p) =D(p)+λ.R(p)
[0045] D(p) is the distortion between the current block and the reconstructed block obtained after encoding / decoding the current block with candidate coding mode p, R(p) is the rate cost associated with encoding the current block with coding mode p, and λ is a Lagrangian parameter that represents the rate constraint for encoding the current block and is typically calculated based on the quantization parameter used to encode the current block.
[0046] The current block is usually encoded according to a prediction residual block PR. More precisely, the prediction residual block PR is calculated, for example, by subtracting the best prediction block from the current block. The prediction residual block PR is then transformed (step 140) by using, for example, a DCT (discrete cosine transform) or DST (discrete sine transform) type transform or any other suitable transform, and the obtained transform coefficient block is quantized (step 150).
[0047] In a variant, the method 100 may also skip the transform step 140 and apply quantization (step 150) directly to the prediction residual block PR according to a so-called transform skip coding mode.
[0048] The block of quantized transform coefficients (or the block of quantized prediction residuals) is entropy encoded into a bitstream (step 160).
[0049] Next, as part of the encoding loop, the quantized transform coefficient block (or quantized residual block) is dequantized (step 170) and inverse transformed (180) (or not inverse transformed), producing a decoded prediction residual block. The decoded prediction residual block and the prediction block are then combined, usually summed, which provides a reconstructed block.
[0050] Further information data may also be entropy encoded in step 160 to encode the current block of the video picture VP.
[0051] An in-loop filter (step 190) may be applied to the reconstructed image (including the reconstructed blocks) to reduce compression artifacts. After all image blocks are reconstructed, a loop filter may be applied. For example, they include a deblocking filter, a sample adaptive offset (SAO), or an adaptive loop filter (ALF).
[0052] The reconstructed block or the filtered reconstructed block forms a reference picture which can be stored into a decoded picture buffer (DPB) so that it can be used as a reference picture for encoding of the next current block of the video picture VP or the next video picture to be encoded.
[0053] Figure 6 A schematic block diagram showing the steps of a method 200 for decoding a video image VP according to the prior art is shown.
[0054] In step 210 , partition information data, prediction information data and quantized transform coefficient blocks (or quantized residual blocks) are obtained by entropy decoding the bit stream of the encoded video image data. For example, the bit stream is generated according to method 100 .
[0055] The other information data may also be entropy decoded to decode the current block of the video picture VP from the bitstream.
[0056] In step 220, the reconstructed image is divided into current blocks based on the partition information. Each current block is entropy decoded from the bitstream along a decoding loop (also referred to as "in-loop"). Each decoded current block is a quantized transform coefficient block or a quantized prediction residual block.
[0057] In step 230, the current block is dequantized and possibly inversely transformed (step 240) to obtain a decoded prediction residual block.
[0058] On the other hand, the prediction information data is used to predict the current block. The prediction block is obtained by its intra prediction (step 250) or its motion compensated temporal prediction (step 260). The prediction process performed at the decoding side is the same as the prediction process performed at the encoding side.
[0059] Next, the decoded prediction residual block and the prediction block are then combined, typically summed, which provides the reconstructed block.
[0060] In step 270, the in-loop filter may be applied to the reconstructed image (including the reconstructed block), and the reconstructed block or the filtered reconstructed block forms a reference image, which may be stored in a decoded picture buffer (DPB) as discussed above ( Figure 5 ).
[0061] In VVC, motion information is stored in every 4×4 block in each video image. This means that once the reference image is stored in the decoded picture buffer (DPB, Figure 5 Or in 6), the motion vector and reference image index for temporal prediction of the video image block are stored on a 4×4 block basis, which can be used as temporal prediction for motion information of subsequent inter-frame prediction video images for encoding / decoding.
[0062] In order to reduce cross-component redundancy, VVC defines the so-called linear model-based intra prediction mode (LM mode).
[0063] One of them is the so-called cross-component linear model (CCLM), which derives a linear model-based predictor (LM predictor for short) by using the following linear model consisting of co-located reconstructed luminance samples rec of the same CU: ′ L The predicted chrominance sample pred of (i,j) C (i,j):
[0064] pred C (i,j)=α·rec L ′(i,j)+β
[0065] where α and β are linear parameters of the linear model, which are derived from the reference samples (i.e., the reconstructed chrominance and luminance samples).
[0066] Reconstructed brightness samples rec ′ L (i,j) is downsampled to match the CU chroma size after filtering.
[0067] In VVC, three CCLM modes are specified, denoted as CCLM_LT, CCLM_T, and CCLM_L. These three CCLM modes differ in the location of the reference samples used for linear parameter derivation. Reference samples from the top boundary are included in the CCLM_T mode, and reference samples from the left boundary are included in the CCLM_L mode. In the CCLM_LT mode, reference samples from both the top boundary and the left boundary are used.
[0068] In general, the prediction process of CCLM mode consists of the following three steps: 1) reconstruct the brightness samples rec of the brightness block and its adjacent brightness samples rec ′ L (i, j) are downsampled to match the size of the corresponding chroma block, 2) linear parameters are derived based on the adjacent reconstructed luma samples, and 3) equation (10) is applied to generate the chroma intra-frame prediction samples (predicted chroma block).
[0069] Another LM mode is the so-called multi-model linear model (MMLM, [K. Zhan et al., "Enhanced Cross-component Linear Model Intra Prediction", JVET-D0110, San Diego, October 2016]). MMLM is an extension of CCLM because more than one linear model is used to reconstruct the luminance samples rec from the co-located ′ L (i, j). In MMLM, adjacent reconstructed luminance samples and adjacent chrominance samples are classified into several groups, and each group is used as a training set to derive linear parameters of the linear model (i.e., specific α and β are derived for a specific group). In addition, the samples of the current luminance block are also classified based on the same rules used for the classification of adjacent luminance samples. For example, adjacent samples are classified into M groups. The MMLM method of M=2 and M=3 is designed as two additional LM modes for chrominance in addition to the original LM mode (CCLM), named MMLM2 and MMLM3. The encoder selects the best LM mode in the rate / distortion optimization process and signals the best LM mode. For example, when M is equal to 2, the threshold is calculated as the average of the adjacent reconstructed luminance samples. Adjacent reconstructed samples rec' that are lower than or equal to the threshold L[x,y] is classified into group 1; and the adjacent reconstructed samples rec' L [x,y] is classified into group 2.
[0070] Then, the following two LM predictors are derived:
[0071]
[0072] MMLM is included in the Enhanced Compression Model (ECM) that explores compression performance improvements beyond VVC (M. Coban et al., "Algorithm description of Enhanced Compression Model 4 (ECM 4)", JVET-Y2025, online, July 2021), where adjacent reconstructed samples are classified into two classes using a threshold that is the average of the brightness adjacent reconstructed samples. A linear model for each class is derived using the least mean square (LMS) method.
[0073] VVC further defines a decoder-side intra mode derivation (DIMD) mode for both luma and chroma samples. The DIMD mode is not a linear model-based mode (referred to as non-LM mode), i.e., an intra prediction mode that does not involve a linear model, such as, for example, a planar mode or a direct mode (DM).
[0074] DM is an intra prediction mode that uses the average sample value of reference samples to the left and above the block for prediction.
[0075] Planar mode is a weighted average of 4 reference sample values (picked as orthogonal projections of samples for prediction in top and left reconstruction regions). Horizontal and vertical modes use copies of left and above reconstructed samples respectively without interpolation to predict sample rows and columns respectively.
[0076] For luma sample prediction, the use of DIMD luma mode is signaled in the bitstream by a single flag, and the intra predictor is not explicitly signaled in the bitstream.
[0077] Figure 7 A block diagram of a method 300 of deriving a DIMD brightness pattern according to the prior art is schematically illustrated.
[0078] Basically, the DIMD luma mode (predictor) is derived by using the gradient analysis of neighboring reconstructed luma samples (i.e., the gradient analysis of luma samples located in the L-shaped template region defined around the current block of samples). If the DIMD luma mode is not enabled, the intra prediction mode can be parsed from the bitstream as in the classic intra prediction mode. The DIMD luma mode is implicit. Therefore, the DIMD luma mode is derived identically at the encoder and decoder side during the reconstruction process.
[0079] In step 310, available upper samples in the frame and available left samples in the frame around the luminance block to be predicted are determined, thereby determining an L-shaped template area.
[0080] For example, Figure 8 As illustrated, a 3-sample wide (in width or height) L-shaped template region T consisting of left, above, and upper-left reconstructed luma samples of a reconstructed region R adjacent to a current block B (current CU) is defined.
[0081] In step 320, a Histogram of Gradients (HoG) is constructed as follows.
[0082] First, the samples of the L-shaped template region T are filtered using a filter window W centered at the middle line sample position of the L-shaped template region T. The amplitude and angle of the brightness direction (orientation) are assigned to each middle line sample of the L-shaped template region T.
[0083] For example, when an edge detection filter (3×3 horizontal and vertical Sobel filter) is used to filter samples of an L-shaped template region T, the magnitude and angle in the brightness direction are given by:
[0084] Angle = arctan(G hor / G ver )
[0085] Amplitude = |G hor |+|G hor |
[0086] Among them G hor and G ver It is the pure horizontal and vertical intensity calculated by the Sobel filter. Each pair of magnitude and angle in the luma direction corresponds to an angular intra-prediction mode.
[0087] Next, the HoG is computed, where each entry (angle) corresponds to an angular intra prediction mode, and the accumulated magnitudes are stored.
[0088] In step 330, at most two angular intra prediction modes M1 and M2 are selected from the HoG.
[0089] For example, Fig. 9 As illustrated, the two most representative angular intra prediction modes M1 and M2 have the largest histogram magnitude values.
[0090] It may happen that 0, 1 or 2 angular intra prediction modes are selected from the HoG. For example, if the accumulated magnitudes of the HoGs are not greater than a threshold, no angular intra prediction mode is selected.
[0091] When no angular intra prediction mode is selected, then in step 340, the planar mode is the intra prediction mode used as the intra predictor for the luma portion of the current block to be predicted.
[0092] When a single angular intra prediction mode is selected, then in step 350, that angular intra prediction mode is the intra prediction mode used as the intra predictor for the luma portion of the current block to be predicted.
[0093] When two angular intra prediction modes (M1 and M2) are selected, then in step 360, the DIMD brightness mode is determined by blending (mixing, fusing) three brightness predictors: two intra predictors derived from the two selected angular intra prediction modes M1 and M2, and a planar predictor (M. Abdoli et al., "Non-CE3: Decoder-side Intra Mode Derivation with Prediction Fusion Using Planar", JVET-O0449, Gothenburg, July 2019).
[0094] like Fig.10 As shown, the blend is defined as a weighted average of the three intra predictors using three weights w derived (step 361) from the ratio of the selected angular intra prediction magnitudes by: 1 、w 2 、w 3 To do:
[0095]
[0096] The weight w 1 and w 2 For two angular intra prediction modes M1 and M2, the weight w 3For planar mode, that is, 21 / 64 with 6-bit integer precision. In step 362, the intra predictor of the current luminance block is calculated based on the two selected angular intra prediction modes M1 and M2 and the planar mode. Finally, in step 363, the DIMD predictor is determined by blending the three intra predictors using weights.
[0097] Fig.11 A block diagram of a method 400 of deriving a DIMD chromaticity mode according to the prior art is schematically illustrated.
[0098] Basically, to derive the DIMD chroma pattern, a similar approach used to derive the DIMD luma pattern is applied to the co-sited reconstructed luma samples, i.e., a gradient analysis is performed on the co-sited reconstructed luma samples in an L-shaped template region defined around the current block of chroma samples.
[0099] The use of DIMD chroma mode is signaled in the bitstream by a single flag and the intra predictor is not explicitly signaled in the bitstream but is derived by gradient analysis using co-located reconstructed luma samples. The DIMD chroma mode is implicit. Hence, the intra predictor is derived from the DIMD chroma mode equally at the encoder and decoder side during the reconstruction process.
[0100] In step 410, method 400 determines whether reconstructed luma samples are available. These reconstructed luma samples are co-located with the chroma block to be predicted, or with the L-shaped template area around the chroma block. The L-shaped template area is formed from the available L-shaped template area.
[0101] In step 420, a gradient histogram HoG is constructed for the available co-sited reconstructed brightness samples in the L-shaped region, as in step 320. In particular, for each co-sited reconstructed brightness sample ( Fig.12 The gray circles in the figure are taken from JVET-Y0092) calculate the horizontal gradient and the vertical gradient, and construct the HoG based on the horizontal gradient and the vertical gradient.
[0102] In step 430, the magnitudes of the available chroma samples in an L-shaped region around the chroma block are accumulated into a HoG, where each entry corresponds to an orientation, ie, angular intra prediction mode (same as the luma HoG counter-part).
[0103] In a variant, HoG can be calculated by accumulating amplitude values over the reconstructed Y, Cb, and Cr samples of the L-shaped template region ( Fig.13 ), that is, the first HoG ( Fig.13 The left part of the L-shaped template area), the second HoG calculated based on the Cb sample ( Fig.13 The middle part of the graph) and the third HoG calculated from the adjacent reconstructed Cr samples ( Fig.13 The right part of ) is obtained by HoG.
[0104] In step 440, the angular intra prediction mode corresponding to the maximum histogram amplitude value is selected from HoG as the first DIMD chroma mode.
[0105] When the DIMD chroma mode is not the direct mode (DM), the method 400 ends, and the DIMD chroma mode is the first DIMD chroma mode.
[0106] Direct mode (DM) is a chroma intra prediction mode corresponding to an intra prediction mode that co-locates reconstructed luma samples.
[0107] When the first DIMD chroma mode is the direct mode (DM), then in step 450, the second DIMD chroma mode is selected as the angular intra prediction mode corresponding to the second largest histogram magnitude value.
[0108] When the first DIMD chromaticity mode is not equal to the second DIMD chromaticity mode, the method ends, and the DIMD chromaticity mode is the second DIMD chromaticity mode.
[0109] When the first DIMD chroma mode is equal to the second DIMD chroma mode, then in step 460, the DIMD chroma mode is direct coding (DC).
[0110] Then, the non-linear model-based predictor (referred to as non-LM predictor for short) of the chroma block is derived from the non-LM mode selected from the five default modes and the DIMD chroma mode. The chroma predictor is derived from those non-LM modes, and the selected non-LM mode corresponds to the non-LM predictor that minimizes the rate-distortion tradeoff.
[0111] The five default modes are DC, planar mode, direct mode, horizontal and vertical angle prediction.
[0112] The non-LM predictor of the chroma block derived from the selected non-LM mode is then fused (blended, mixed) with the LM predictor derived from the MMLM mode as follows:
[0113] pred=(w0*pred0+w1*pred1)>>shift
[0114] Where pred0 is the predictor obtained by applying the non-LM mode, pred1 is the predictor obtained by applying the LM mode, and pred is the final predictor of the chroma block. For I slices (intra-coded slices), the two weights w0 and w1 are determined by the intra-prediction mode of the adjacent chroma blocks, and shift is set to be equal to 2. In particular, when both the upper and left adjacent blocks are encoded in LM mode, {w0,w1}={1,3}; when both the upper and left adjacent blocks are encoded in non-LM mode, {w0,w1}={3,1}; otherwise, {w0,w1}={2,2}. For non-I slices, w0 and w1 are both set to be equal to 2.
[0115] If non-LM mode is selected, a flag is signaled to indicate whether fusion is applied.
[0116] In a variant, for I slices (i.e., intra-frame coded slices), non-LM predictors derived from DM mode, four default modes and DIMD chroma mode can be fused with the LM predictor derived from LM mode, while for non-I slices, only the non-LM predictor derived from DIMD chroma mode can be fused with the LM predictor derived from LM mode using equal weights.
[0117] Chroma Fusion mode can be used with DIMD Chroma (if selected as the best non-LM mode), but any other non-LM mode can be used instead.
[0118] Using DIMD for luma and chroma improves the coding efficiency of the video image. However, the price to pay is about 20% additional complexity at the encoder and about 5% additional complexity at the decoder for all intra conditions.
[0119] The problem to be solved by the present disclosure is to maintain the coding efficiency and reduce the complexity of the state-of-the-art video codec when DIMD is used for luminance and chrominance.
[0120] At least one embodiment of the present disclosure has been designed in view of the foregoing. Summary of the invention
[0121] The following section presents a simplified summary of at least one embodiment in order to provide a basic understanding of some aspects of the present disclosure. This summary is not an exhaustive overview of the embodiments. It is not intended to identify the key or important elements of the embodiments. The following summary only presents some aspects of at least one embodiment in a simplified form as a prelude to a more detailed description provided elsewhere in the document.
[0122] According to a first aspect of the present disclosure, a method for decoding a sample block of a video image is provided, the method comprising determining an intra prediction mode for decoding the sample block based on an analysis of gradients of samples located in at least one template area defined around the sample block by: calculating a gradient histogram by filtering samples of at least one template area, wherein each entry of the gradient histogram corresponds to an angular intra prediction mode, the filtering being performed using a filtering window centered on a midline sample position of at least one template area; selecting at most two angular intra prediction modes by comparing the magnitudes of the angular intra prediction modes in the gradient histogram; determining the intra prediction mode from the at most two selected angular intra prediction modes; wherein the integer number of midline sample positions of at least one template area around which the filtering window is centered is lower than the total integer number of midline sample positions of at least one template area.
[0123] According to a second aspect of the present disclosure, a method for encoding a sample block of a video image is provided, the method comprising determining an intra-frame prediction mode for decoding the sample block based on an analysis of gradients of samples located in at least one template area defined around the sample block by: calculating a gradient histogram by filtering samples of at least one template area, wherein each entry of the gradient histogram corresponds to an angular intra-frame prediction mode, the filtering being performed using a filtering window centered on a midline sample position of at least one template area; selecting at most two angular intra-frame prediction modes by comparing the magnitudes of the angular intra-frame prediction modes in the gradient histogram; determining the intra-frame prediction mode from the at most two selected angular intra-frame prediction modes; wherein the integer number of midline sample positions of at least one template area around which the filtering window is centered is lower than the total integer number of midline sample positions of at least one template area.
[0124] In one embodiment, a gradient histogram of the chroma component of the video image is calculated by filtering only the chroma samples of at least one template region.
[0125] In one embodiment, at least one of the at least one stencil region includes samples along a bottom frontier of an adjacent upper virtual pipeline data unit and samples along a right frontier of an adjacent left virtual pipeline data unit.
[0126] In one embodiment, at least one of the at least one stencil region includes all luminance samples on a border of adjacent virtual pipeline data units.
[0127] In one embodiment, the stencil region includes a subset of luma samples on the boundaries of adjacent virtual pipeline data units.
[0128] In one embodiment, a gradient histogram is calculated for each chrominance component of the video image.
[0129] In one embodiment, each gradient histogram of the chrominance component of the video image is calculated based on the luma samples and the chrominance samples of the template region.
[0130] In one embodiment, the chroma samples of the template are multiplied by a compensation factor that depends on the chroma subsampling defined by the video image format.
[0131] In one embodiment, the middle line sample position of at least one template region of at least one template region around which the filter window is centered is determined to position the filter window at one middle line sample position of a fourth number (N4) of middle line sample positions of the template region.
[0132] In one embodiment, the midline sample position of at least one of the at least one template regions around which the filter window is centered is determined to avoid any overlap between the filter windows.
[0133] In one embodiment, the filter windows have different sizes.
[0134] In one embodiment, the size of the filter window depends on the size of the sample block and the sample availability of the template area.
[0135] According to a third aspect of the present disclosure, there is provided an apparatus comprising means for executing one of the methods according to the first and / or second aspect of the present disclosure.
[0136] According to a fourth aspect of the present disclosure, a non-transitory storage medium is provided, which carries program code instructions for executing the method according to the first and / or second aspect of the present disclosure.
[0137] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing instructions executable by the processor. The processor is configured to execute the method according to the first and / or second aspect of the present disclosure.
[0138] According to a sixth aspect of the present disclosure, a computer program product is provided, which includes instructions. When the instructions are executed by one or more processors, the one or more processors execute the method according to the first and / or second aspect of the present disclosure.
[0139] The specific nature of at least one of the embodiments and other objects, advantages, features and uses of at least one of the embodiments will become apparent from the following description of the examples taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0140] Reference will now be made, by way of example, to the accompanying drawings which show embodiments of the present disclosure, and in which:
[0141] Figure 1 An example of a coding tree unit according to HEVC is shown;
[0142] Figure 2 An example of partitioning a coding unit into prediction units according to HEVC is shown;
[0143] Figure 3 An example of CTU partitioning according to VVC is shown;
[0144] Figure 4 An example of supported partitioning modes in multi-type tree partitioning according to VVC is shown;
[0145] Figure 5 A schematic block diagram showing the steps of a method 100 for encoding a video image VP according to the prior art;
[0146] Figure 6 A schematic block diagram showing the steps of a method 200 for decoding a video image VP according to the prior art;
[0147] Figure 7 Schematically shows a block diagram of a method 300 for deriving a DIMD brightness mode according to the prior art;
[0148] Figure 8 An L-shaped template area is illustrated;
[0149] Fig. 9 An example of a HoG is shown;
[0150] Fig.10 A method of blending multiple brightness predictors according to the prior art is schematically shown;
[0151] Fig.11 Schematically illustrates a block diagram of a method 400 for deriving a DIMD chromaticity mode according to the prior art;
[0152] Fig.12 schematically illustrates an example of co-located luma samples of a chroma block according to the prior art;
[0153] Fig.13 Schematically illustrates other examples of co-located luma samples of chroma blocks according to the prior art;
[0154] Fig.14 Schematically illustrates a block diagram of a method 500 of deriving a DIMD mode according to an embodiment;
[0155] Fig.15schematically illustrates the position of the filter window according to an embodiment (step 510);
[0156] Fig.16 schematically illustrates the position of the filter window according to an embodiment (step 510);
[0157] Fig.17 schematically illustrates the position of the filter window according to an embodiment (step 510);
[0158] Fig.18 illustrates a VPDU and an L-shaped template area around the current VPDU according to an embodiment;
[0159] Fig.19 illustrates an L-shaped template area when using VPDU according to an embodiment;
[0160] Fig. 20 illustrates an L-shaped template area when using VPDU according to an embodiment;
[0161] Fig.21 schematically illustrates a block diagram of a method 600 of constructing a HoG for chroma DIMD derivation according to an embodiment; and
[0162] Fig. 22 A schematic block diagram of an example of a system in which various aspects and embodiments are implemented is illustrated.
[0163] Similar or identical elements are denoted by the same reference numerals. DETAILED DESCRIPTION
[0164] At least one embodiment of the embodiments will be described more fully below with reference to the accompanying drawings, wherein an example of at least one embodiment of the embodiments is depicted. However, the embodiments can be implemented in many alternative forms and should not be interpreted as being limited to the examples set forth herein. Thus, it should be understood that the present disclosure is not intended to limit the embodiments to the specific forms disclosed. On the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives that fall within the spirit and scope of the present disclosure.
[0165] At least one of these aspects generally relates to video image encoding and decoding, another aspect generally relates to transmitting a provided or encoded bitstream, and one of the other aspects relates to receiving / accessing a decoded bitstream.
[0166] At least one of these embodiments is described for encoding / decoding video images, but extends to encoding / decoding video images (image sequences) in that each video image is encoded / decoded sequentially as described below.
[0167] In addition, at least one embodiment is not limited to AVC (ISO / IEC 14496-10, Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Basic Video Coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), VVC (ISO / IEC 23090-3 Universal Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), but may be applicable to other standards and recommendations such as AV1 (AOMediaVideo 1, http: / / aomedia.org / av1 / specification / ). The at least one embodiment may be applicable to existing or future developed standards and recommendations and extensions of any such standards and recommendations. Unless otherwise specified or technically excluded, the aspects described in the present disclosure may be used alone or in combination.
[0168] In general, the present disclosure relates to decoding a sample block of a video image, wherein a DIMD pattern is derived for luma samples and chroma samples of a sample block to be predicted by filtering samples of at least one template area, wherein the filtering is performed on samples of at least one shaped template area using a filtering window centered about midline sample positions of the at least one template area, wherein an integer number of midline sample positions of the at least one template area on which the filtering window is centered is lower than a total integer number of midline sample positions of the at least one template area.
[0169] This reduces the resource requirements (memory and computational power, complexity) for deriving DIMD patterns for both luma and chroma samples of a block to be predicted compared to the resources required by current DIMD pattern derivation methods.
[0170] The following embodiments are described using an L-shaped template region consisting of left, top, and top-left reconstructed samples of a reconstructed region adjacent to a current block. However, the present disclosure extends to a template region consisting only of left reconstructed samples of a reconstructed region adjacent to a current block, or a template region consisting only of top reconstructed samples of a reconstructed region adjacent to a current block.
[0171] Fig.14 A block diagram of a method 500 of deriving a DIMD mode according to an embodiment is schematically illustrated.
[0172] The method 500 may be applied to both the method 100 (encoding) and the method 200 (decoding) of intra-predicting a block.
[0173] In step 510, middle line sample positions of the L-shaped template region T1 are determined. The integer number N1 of middle line sample positions of the L-shaped template region T1 around which the filter window is centered is lower than the total number of middle line sample positions of the L-shaped template region.
[0174] In step 520, a luminance DIMD pattern is derived from method 300, wherein the filter window is centered at each of the N1 middle line sample positions of the L-shaped template region T1.
[0175] In one variation of step 520, middle line sample positions of the L-shaped template region T2 of the co-located reconstructed luma samples are determined. Middle line sample positions of the L-shaped template region T3 of the available chroma samples are determined. The number N2 or N3 of middle line sample positions of the L-shaped template region T2 or T3 on which the filter window is centered is lower than the total number of middle line sample positions of the L-shaped template region T2 or T3. The chroma DIMD pattern is derived from method 400, wherein the filter window is centered on each of the N2 middle line sample positions of the L-shaped template region T2 and is centered on each of the N3 middle line sample positions of the L-shaped template region T3.
[0176] In one embodiment of step 510, the midline sample position of at least one L-shaped template area (T1, T2, T3) around which the filter window is centered is determined to position the filter window at one midline sample position among the number N4 (N4>=2) of midline sample positions of the L-shaped template area T1, T2 or T3.
[0177] Fig.15 An example of the position of a filter window centered on the determined mid-line sample position (white circle) is schematically illustrated when N1 = 2. The filter window is indicated by a dotted line.
[0178] In one embodiment of step 510, the midline sample position of at least one of the at least one L-shaped template regions (T1, T2, T3) around which the filter window is centered is determined to avoid any overlap between the filter windows.
[0179] Fig.16 An example of the position of a filter window centered about the determined mid-line sample position (white circle) is schematically illustrated. The filter window is indicated by a dotted line.
[0180] In one embodiment of step 510, the middle line sample position of at least one L-shaped template region of at least one L-shaped template region (T1, T2, T3) around which the filter window is centered is determined to allow the filter window to have a common boundary. These middle line sample positions of the L-shaped template region can then be determined in the upper left, upper, and top positions of the L-shaped template region.
[0181] Fig.17 An example of the position of a filter window centered about the determined mid-line sample position (white circle) is schematically illustrated. The filter window is indicated by a dotted line.
[0182] In one embodiment of the method 500, the size of the filter window is larger than 3×3 to compensate for the reduction in the number of filtered samples (ie, the reduction in the number of HoG entries). For example, a filter of size 5×5 is used.
[0183] In one embodiment of method 500, the filter windows have different sizes.
[0184] In one embodiment of the method 500, the size of the filter window depends on the size of the block to be predicted and the sample availability of the L-shaped template region.
[0185] For example, depending on sample availability, the size of the filter window is increased for L-shaped template regions in larger sized blocks (perhaps when the block size exceeds a predetermined size, such as 8×8).
[0186] The previous embodiments for determining the midline sample position of the L-shaped template region around which the filter window is centered and the filter window size may be combined.
[0187] For example, for samples in an L-shaped template region adjacent to a longer block size, the filter window size is 5×5, and the authorized mid-line sample positions are at specific locations, either at evenly spaced locations, or at locations where the filter windows do not overlap.
[0188] Since the number of intermediate line sample positions is reduced (step 510) and may be less representative of directions that exist in the vicinity of the current block to be predicted, in one embodiment, the blending weight derivation (for luma) is modified so that instead of calculating the weight w by equation (1) 1 and w 2 , but the weight w is calculated as follows 1 and w 2 :
[0189] If the magnitude of the first angle intra prediction mode M1 is greater than twice the magnitude of the second angle intra prediction mode M2, the second weight w 2 is equal to 0, for example and w 3 =21 / 64.
[0190] In one variation, the third weight w 3 is equal to 0 and
[0191] If the magnitude of the first angular intra prediction mode M1 is lower than or equal to twice the magnitude of the second angular intra prediction mode M2, the second weight w 2 Not equal to 0.
[0192] For example,
[0193] In one variant, the third weight is equal to 0, and for example, and w 3 =32 / 64.
[0194] As discussed in the introduction above, chroma samples in the L-shaped template region and co-sited reconstructed luma samples in the L-shaped template are used to construct the HoG. This requires that the luma samples be encoded and decoded first before the chroma is encoded using the DIMD chroma tool, which increases the latency of the encoding / decoding methods 100 and 200. In addition, in the case where the DIMD luma mode will not be enabled to predict the luma block, the method for deriving the DIMD chroma mode will calculate the HoG from the co-sited reconstructed luma samples in either way to obtain the DIMD chroma predictor. Depending on the hardware architecture design, the calculation of the HoG from the co-sited reconstructed luma samples may be repeated and redundant between the derivation of the DIMD luma mode and the DIMD chroma mode.
[0195] In one embodiment of the method 500, to derive the DIMD pattern of the chroma components of the video image, the HoG is computed based only on the filtered chroma samples of the L-shaped template region.
[0196] This embodiment is applicable only to derived DIMD chrominance modes.
[0197] This embodiment is advantageous because it avoids encoding and decoding luma samples before encoding chroma samples using the DIMD chroma mode. This embodiment is also advantageous because it has low complexity, low latency, and requires less computing power and memory than methods that use co-location to reconstruct luma and chroma samples.
[0198] In VVC, a virtual pipeline data unit (VPDU) is a concept of non-overlapping units of samples, which are typically 64×64 in size for luma and 32×32 in size for chroma. The goal is to ensure that a VPDU is fully processed before starting to process the next VPDU, so that the memory footprint of the hardware implementation remains reasonable. VPDUs can have a significant impact on tool design.
[0199] Therefore, when VPDU is considered in the design of chroma DIMD, and when the construction of the HoG used to derive the DIMD predictor uses not only chroma samples present in the available template area, but also luma samples (such as discussed above), these luma samples are taken from the current VPDU (VPDUc) so that the luma samples are fully processed and available.
[0200] In one embodiment of method 500, at least one L-shaped template region is used, wherein the at least one L-shaped template region includes a VPDU ( Fig.18 The luminance samples along the bottom VPDU edge of the VPDU1 in the figure and along the adjacent left VPDU ( Fig.18 The luminance samples (e.g., at least 3 samples wide) of the right VPDU edge of VPDU2 in FIG.
[0201] This embodiment reduces the latency of method 500 .
[0202] In one variant, at least one of the at least one L-shaped template regions includes all (i.e., 64×2 in VVC design or 64×3 with width 3) luma samples on the boundaries of adjacent VPDUs (VPDU1 and VPDU2) ( Fig.19 ). All of these brightness samples are used to construct the HoG, rather than using adjacent co-located brightness samples.
[0203] In one variant, a subset of luma samples on the boundary of adjacent VPDUs (VPDU1 and VPDU2) is used ( Fig. 20 ). These brightness samples are used to construct the HoG instead of using neighboring co-located brightness samples.
[0204] In one variant, if a subset of luma samples belonging to a 3-sample-wide boundary of adjacent VPDUs is used, then this subset corresponds to luma samples co-located with (a portion of) the orthogonal projection of a chroma block onto the adjacent VPDU, e.g. Fig.18 , 19 and 20 as shown.
[0205] Only the luma samples of the current VPDU (VPDUc) are used to construct the HoG to derive the chroma DIMD predictor. If the neighboring VPDU does not exist, the HoG is constructed using only the chroma samples present in the L-shaped template region.
[0206] In one embodiment of method 500, one HoG is calculated for each chrominance component of a video image.
[0207] This embodiment allows parallelisation of the calculation of the HoG for each chroma component and thus improves the latency of the method 500 .
[0208] Furthermore, this embodiment improves signal adaptation by separating the HoG construction for chroma components and thus independently assigning the final intra prediction mode to each chroma component when chroma DIMD is enabled.
[0209] Since the final intra prediction mode selected for each chrominance component can be different, this embodiment further provides additional flexibility in the design. This embodiment does not require other coding tools to act on the chrominance coding blocks or components separately. The advantage is that no additional signaling (beyond the DIMD chrominance flag) is required for the modes of both components.
[0210] Fig.21 A block diagram schematically illustrates a method 600 of constructing a HoG for chroma DIMD derivation according to an embodiment.
[0211] In this embodiment, the available chroma samples in the L-shaped template area and the available reconstructed co-located luma samples in the L-shaped template area are used to construct each chroma component HoG.
[0212] However, since in non-4:4:4 signals there are fewer chroma samples than luma samples, a compensation coefficient is introduced in the HoG construction.
[0213] In step 610, the current filtering window W is considered.
[0214] In step 620, the horizontal Dx and vertical Dy gradients are calculated by applying horizontal and vertical edge detection filters, such as horizontal and vertical 3x3 Sobel filters.
[0215] In step 635, compensation coefficients are determined based on the signal chroma subsampling defined by the video image format.
[0216] For example, for a 4:2:0 signal, the compensation factor is equal to 2, for a 4:2:2 signal, the compensation factor is equal to 1.5, and for a 4:4:4 signal, the compensation factor is equal to 1.
[0217] In one variant, for a 4:2:0 signal the compensation factor is a multiple of 2, for a 4:2:2 signal the compensation factor is a multiple of 1.5, and for a 4:4:4 signal the compensation factor is a multiple of 1.
[0218] Alternatively, for a 4:2:0 signal, one of the two (evenly spaced) luma samples of the L-shaped template region is used as the center of the filter window for constructing the HoG (e.g. Fig.15 ), so that the number of processed intermediate line luma samples is equal to the number of processed chroma samples (for each chroma component).
[0219] exist Fig.21 In a variation of this embodiment, only the available chroma samples in the L-shaped template region are used in constructing each chroma component HoG. Then, filter window positions are added to compensate for the lack of samples. Typically, the third row or column of chroma samples is also used as the filter window position in the center for constructing each chroma HoG.
[0220] Fig. 22 A schematic block diagram illustrating an example of a system 700 in which various aspects and embodiments are implemented is shown.
[0221] System 700 may be embedded as one or more devices, including various components described below. In various embodiments, system 700 may be configured to implement one or more aspects described in the present disclosure.
[0222] Examples of equipment that may constitute all or part of system 700 include personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, connected vehicles and their associated processing systems, head-mounted display devices (HMD, perspective glasses), projectors (projectors), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors for processing outputs from video decoders, pre-processors for providing inputs to video encoders, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or network servers), static or video cameras, encoding or decoding chips, or any other communication devices. The elements of system 700 may be implemented individually or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 700 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 700 may be coupled to other similar systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports.
[0223] The system 700 may include at least one processor 710 configured to execute instructions loaded therein for implementing, for example, various aspects described in the present disclosure. The processor 710 may include embedded memory, input-output interfaces, and various other circuits known in the art. The system 700 may include at least one memory 720 (e.g., a volatile memory device and / or a non-volatile memory device). The system 700 may include a storage device 740, which may include a non-volatile memory and / or a volatile memory, including but not limited to an electrically erasable programmable read-only memory (EEPROM), a read-only memory (ROM), a programmable read-only memory (PROM), a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a flash memory, a disk drive, and / or an optical drive. As a non-limiting example, the storage device 740 may include an internal storage device, an attached storage device, and / or a network accessible storage device.
[0224] The system 700 may include an encoder / decoder module 730, which is configured to, for example, process data to provide encoded / decoded video image data, and the encoder / decoder module 730 may include its own processor and memory. The encoder / decoder module 730 may represent a module (one or more) that may be included in a device to perform encoding and / or decoding functions. As known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 730 may be implemented as a separate element of the system 700, or may be incorporated into the processor 710 as a combination of hardware and software known to those skilled in the art.
[0225] Program code to be loaded onto the processor 710 or the encoder / decoder 730 to perform various aspects described in the present disclosure may be stored in the storage device 740 and subsequently loaded onto the memory 720 to be executed by the processor 710. According to various embodiments, during the execution of the processes described in the present disclosure, one or more of the processor 710, the memory 720, the storage device 740, and the encoder / decoder module 730 may store one or more of various items. Such stored items may include, but are not limited to, video image data, information data for encoding / decoding video image data, bit streams, matrices, variables, and intermediate or final results of equations, formulas, operations, and operation logic processing.
[0226] In several embodiments, memory internal to the processor 710 and / or encoder / decoder module 730 may be used to store instructions and provide working memory for processes that may be performed during encoding or decoding.
[0227] However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 710 or the encoder / decoder module 730) is used for one or more of these functions. The external memory may be a memory 720 and / or a storage device 740, such as a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM may be used as working memory for video encoding and decoding operations, such as for MPEG-2 Part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), AVC, HEVC, EVC, VVC, AV1, etc.
[0228] As indicated in block 790, input to the elements of system 700 may be provided through various input devices. Such input devices include, but are not limited to, (i) an RF section that may receive an RF signal transmitted over the air, for example, by a broadcast device, (ii) a composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) when the present disclosure is implemented in the automotive field, a bus such as a CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate), FlexRay (ISO 17458), or Ethernet (ISO / IEC 802-3) bus.
[0229] In various embodiments, the input device of block 790 has associated corresponding input processing elements, as known in the art. For example, the RF portion may be associated with elements necessary for: (i) selecting a desired frequency (also referred to as selecting a signal, or limiting the signal band to within a band), (ii) down-converting the selected signal, (iii) limiting the band again to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a desired packet stream. The RF portion of various embodiments may include one or more elements that perform these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or a baseband.
[0230] In one set-top box embodiment, the RF section and its associated input processing elements can receive RF signals transmitted over a wired (e.g., cable) medium. The RF section can then perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band.
[0231] Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions.
[0232] Adding an element may include inserting an element between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter.In various embodiments, the RF portion may include an antenna.
[0233] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting the system 700 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or in the processor 710 when necessary. Similarly, various aspects of USB or HDMI interface processing may be implemented in a separate interface IC or in the processor 710 when necessary. The demodulated, error-corrected, and demultiplexed streams may be provided to various processing elements, including, for example, the processor 710 and the encoder / decoder 730, which operate in conjunction with memory and storage elements to process the data streams for presentation on output devices when necessary.
[0234] The various elements of system 700 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and data transferred between them using a suitable connection arrangement 790, such as an internal bus (including an I2C bus), wiring, and printed circuit boards known in the art.
[0235] The system 700 may include a communication interface 750 that enables communication with other devices via a communication channel 751. The communication interface 750 may include, but is not limited to, a transceiver configured to send and receive data over the communication channel 751. The communication interface 750 may include, but is not limited to, a modem or a network card, and the communication channel 751 may be implemented, for example, within a wired and / or wireless medium.
[0236] In various embodiments, a Wi-Fi network such as IEEE 802.11 may be used to stream data to the system 700. The Wi-Fi signals of these embodiments may be received via a communication channel 751 suitable for Wi-Fi communications and a communication interface 750. The communication channel 751 of these embodiments may typically connect to an access point or router that provides access to external networks including the Internet to allow streaming applications and other over-the-top communications.
[0237] Other embodiments may provide streaming data to system 700 using a set top box that delivers the data through an HDMI connection to input block 790 .
[0238] Still other embodiments may use the RF connection of input block 790 to provide streaming data to system 700 .
[0239] The streamed data may be used as a means of signaling information used by the system 700. The signaling information may include the bitstream B and / or information such as 7a the number of pixels of a video image and / or any encoding / decoding setting parameters.
[0240] It should be appreciated that signaling may be implemented in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder.
[0241] The system 700 may provide output signals to various output devices, including a display 761, speakers 771, and other peripherals 781. In various examples of embodiments, the other peripherals 781 may include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functionality based on the output of the system 700.
[0242] In various embodiments, control signals may be communicated between the system 700 and the display 761, speakers 771, or other peripherals 781 using signaling such as AV.Link (Audio / Video Link), CEC (Consumer Electronics Control), or other communication protocols that enable device-to-device control with or without user intervention.
[0243] Output devices may be communicatively coupled to system 700 through respective interfaces 760 , 770 , and 780 via dedicated connections.
[0244] Optionally, output devices may be connected to the system 700 using a communication channel 751 via the communication interface 750. The display 761 and the speaker 71 may be integrated into a single unit with the other components of the system 700 in an electronic device such as, for example, a television.
[0245] In various embodiments, the display interface 760 may include a display driver, such as, for example, a timing controller (TCon) chip.
[0246] For example, if the RF portion of input 790 is part of a separate set-top box, then display 761 and speaker 771 may optionally be separate from one or more of the other components. In various embodiments where display 761 and speaker 771 may be external components, output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.
[0247] exist Figures 1 to 22 In the present invention, various methods are described herein, and each method includes one or more steps or actions to implement the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions can be modified or combined.
[0248] Some examples are described about block diagrams and / or operational flow charts. Each square block represents a portion of a circuit element, module or code, which includes one or more executable instructions for implementing (one or more) specified logical functions. It should also be noted that, in other embodiments, the (one or more) functions marked in the square block may not occur in the order indicated. For example, depending on the functions involved, two square blocks shown in succession can actually be executed substantially concurrently, or sometimes these square blocks can be executed in reverse order.
[0249] The embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if only discussed in the context of a single form of embodiment (e.g., discussed only as a method), the embodiments of the features discussed may also be implemented in other forms (e.g., an apparatus or a computer program).
[0250] The method may be implemented in, for example, a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device.
[0251] In addition, the method can be implemented by instructions executed by a processor, and such instructions (and / or data values generated by the implementation) can be stored on a computer-readable storage medium. The computer-readable storage medium can take the form of a computer-readable program product implemented in one or more computer-readable media and having a computer-readable program code implemented thereon that can be executed by a computer. Considering the inherent ability to store information therein and the inherent ability to provide information retrieval therefrom, the computer-readable storage medium used herein can be considered as a non-transient storage medium. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. It should be appreciated that although more specific examples of computer-readable storage media to which the present embodiment can be applied are provided below, as those of ordinary skill in the art will readily recognize, it is merely an illustrative and non-exhaustive list: portable computer floppy disk; hard disk; read-only memory (ROM); erasable programmable read-only memory (EPROM or flash memory); portable compact disk read-only memory (CD-ROM); optical storage device; magnetic storage device; or any suitable combination of the foregoing.
[0252] The instructions may form an application program tangibly embodied on a processor-readable medium.
[0253] For example, instructions may be in hardware, firmware, software, or a combination. For example, instructions may be found in an operating system, a separate application, or a combination of both. Thus, a processor may be characterized as, for example, a device configured to perform a process and a device including a processor-readable medium (such as a storage device) having instructions for performing the process. Additionally, in addition to or in lieu of instructions, a processor-readable medium may store data values generated by an embodiment.
[0254] Device can be realized in suitable hardware, software and firmware for example.The example of such device comprises personal computer, laptop computer, smart phone, tablet computer, digital multimedia set-top box, digital television receiver, personal video recording system, connected household appliances, head-mounted display device (HMD, perspective glasses), projector (projector), "cave" (system including multiple displays), server, video encoder, video decoder, post-processor for processing the output from video decoder, pre-processor for providing input to video encoder, web server, set-top box, and any other equipment for processing video image, or other communication equipment.It should be clear that equipment can be mobile and even installed in a mobile vehicle.
[0255] Computer software may be implemented by processor 710 or by hardware, or by a combination of hardware and software. As a non-limiting example, embodiments may also be implemented by one or more integrated circuits. Memory 720 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology (such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples). Processor 710 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture, as non-limiting examples.
[0256] As will be apparent to one of ordinary skill in the art, embodiments may generate various signals formatted to carry information that may be stored or transmitted, for example. The information may include, for example, instructions for executing a method or data generated by one of the described embodiments. For example, a signal may be formatted to carry a bit stream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.
[0257] The term used herein is only used to describe the purpose of specific embodiments and is not intended to be limited. As used herein, the singular form "one / kind (a)", "one / kind (an)" and "the / said (the)" may also be intended to include plural forms, unless the context clearly indicates otherwise. It will be further understood that, when used in this specification, the term "include / comprise" and / or "include / comprising (including / comprising)" may specify the existence of stated, for example, features, integers, steps, operations, elements and / or components, but does not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. Moreover, when an element is referred to as "response" or "connection" or "associated" to another element, it may directly respond or be connected to another element or be associated with another element, or there may be an intermediate element. On the contrary, when an element is referred to as "direct response" or "direct connection" to another element or "directly associated" with another element, there is no intermediate element.
[0258] It should be appreciated that use of any of the symbols / terms " / ", "and / or", and "at least one of" may be intended to encompass selection of only the first listed option (A), or only the second listed option (B), or both options (A and B), such as in the case of "A / B", "A and / or B", and "at least one of A and B". As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to encompass selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). This may be extended to as many items as listed, as will be apparent to one of ordinary skill in this and related arts.
[0259] Various numerical values may be used in the present disclosure. Specific values may be used for example purposes and the described aspects are not limited to these specific values.
[0260] It will be understood that although the terms first, second, etc. can be used to describe various elements in this article, these elements are not limited by these terms. These terms are only used to distinguish one element from another element. For example, without departing from the teaching of the present disclosure, the first element can be referred to as the second element, and similarly, the second element can be referred to as the first element. There is no suggestion of sorting between the first element and the second element.
[0261] References to "one embodiment" or "an embodiment" or "one implementation" or "an implementation" and other variations thereof are frequently used to convey that a particular feature, structure, characteristic, etc. (described in conjunction with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in an implementation" or "in an implementation" and any other variations appearing in various places in this disclosure are not necessarily all referring to the same embodiment.
[0262] Similarly, references herein to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations thereof are frequently used to convey that a particular feature, structure, or characteristic (described in conjunction with an embodiment / example / implementation) may be included in at least one embodiment / example / implementation. Therefore, the expressions "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" appearing in various places in this disclosure do not necessarily all refer to the same embodiment / example / implementation, nor are separate or alternative embodiments / examples / implementations necessarily mutually exclusive of other embodiments / examples / implementations.
[0263] Reference numerals appearing in the claims are for illustration only and have no limiting effect on the scope of the claims.The present embodiments / examples and variants may be employed in any combination or sub-combination although not explicitly described.
[0264] When a figure is presented as a flow chart, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow chart of the corresponding method / process.
[0265] While some diagrams include arrows on communication paths to illustrate a primary direction of communication, it should be understood that communication can occur in the opposite direction to the depicted arrows.
[0266] Various embodiments relate to decoding. As used in this application, "decoding" may encompass, for example, all or part of a process performed on a received video image (which may include a received bitstream encoding one or more video images) to produce a final output suitable for display or further processing in a reconstructed video domain. In various embodiments, such a process includes one or more of the processes typically performed by a decoder. In various embodiments, for example, such a process also or alternatively includes a process performed by a decoder of the various embodiments described in this disclosure.
[0267] As a further example, in one embodiment "decoding" may refer only to dequantization, in one embodiment "decoding" may refer to entropy decoding, in another embodiment "decoding" may refer only to differential decoding, and in another embodiment "decoding" may refer to a combination of dequantization, entropy decoding, and differential decoding. Based on the context of the specific description, whether the phrase "decoding process" may be intended to specifically refer to a subset of operations, or generally refer to a broader decoding process will be clear and is believed to be well understood by those skilled in the art.
[0268] Various embodiments all relate to encoding. In a manner similar to the above discussion about "decoding", "encoding" used in the present disclosure may encompass, for example, all or part of a process performed on an input video image to produce an output bitstream. In various embodiments, such a process includes one or more of the processes typically performed by an encoder. In various embodiments, such a process also includes or optionally includes a process performed by an encoder of the various embodiments described in the application.
[0269] As a further example, in one embodiment "encoding" may refer only to quantization, in one embodiment "encoding" may refer only to entropy coding, in another embodiment "encoding" may refer only to differential coding, and in another embodiment "encoding" may refer to a combination of quantization, differential coding, and entropy coding. Based on the context of the particular description, whether the phrase "encoding process" may be intended to specifically refer to a subset of operations, or to generally refer to a broader encoding process will be clear and is believed to be well understood by those skilled in the art.
[0270] Additionally, the present disclosure may refer to "obtaining" various information. Obtaining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0271] Additionally, the application may refer to "receiving" various information. Receiving information may include, for example, one or more of accessing the information or receiving the information from a communication network.
[0272] Moreover, as used herein, the word "signal" especially refers to indicating something to a corresponding decoder, etc. For example, in some embodiments, the encoder signals specific information, such as encoding parameters or encoded video image data. In this way, in an embodiment, the same parameter can be used on the encoder side and the decoder side. Therefore, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters and other parameters, signaling can be used without transmission (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit saving is achieved in various embodiments. It should be recognized that signaling can be completed in a variety of ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to the corresponding decoder. Although the verb form of the word "signal" is mentioned above, the word "signal" can also be used as a noun in this article.
[0273] A number of embodiments have been described. However, it should be understood that various modifications may be made. For example, elements of different embodiments may be combined, supplemented, modified, or removed to produce other embodiments. In addition, it will be understood by those of ordinary skill that other structures and processes may replace the disclosed structures and processes, and that the resulting embodiments will perform at least substantially the same (one or more) functions in at least substantially the same (one or more) manners to achieve at least substantially the same (one or more) results as the disclosed embodiments. Thus, the application contemplates these and other embodiments.
Claims
1. A method for decoding a block of samples of a video image, the method comprising determining an intra prediction mode for decoding the block of samples from an analysis of gradients of samples located in at least one template region (T, T1, T2, T3) defined around the block of samples, by: - calculating (320, 430) a gradient histogram by filtering samples of the at least one template region, wherein each entry of the gradient histogram corresponds to an angular intra prediction mode, the filtering being performed using a filtering window (W) centered at a midline sample position of the at least one template region; - selecting (330, 440) at most two angular intra prediction modes by comparing the magnitudes of the angular intra prediction modes in the gradient histogram; - determining (360, 450) the intra prediction mode from at most two selected angular intra prediction modes; The integer number of the middle line sample positions of the at least one template region around which the filter window is centered is lower than the total integer number of the middle line sample positions of the at least one template region.
2. A method for encoding a block of samples of a video image, the method comprising determining an intra prediction mode for decoding the block of samples from an analysis of gradients of samples located in at least one template region (T, T1, T2, T3) defined around the block of samples, by: - calculating (320, 430) a gradient histogram by filtering samples of the at least one template region, wherein each entry of the gradient histogram corresponds to an angular intra prediction mode, the filtering being performed using a filtering window (W) centered at a midline sample position of the at least one template region; - selecting (330, 440) at most two angular intra prediction modes by comparing the magnitudes of the angular intra prediction modes in the gradient histogram; - determining (360, 450) the intra prediction mode from at most two selected angular intra prediction modes; The integer number of the middle line sample positions of the at least one template region around which the filter window is centered is lower than the total integer number of the middle line sample positions of the at least one template region.
3. The method according to claim 1 or 2, wherein the gradient histogram of the chrominance component of the video image is calculated by filtering only the chrominance samples of the at least one template area.
4. The method according to claim 1 or 2, wherein at least one of the at least one template area includes samples along a bottom edge of an adjacent upper virtual pipeline data unit (VPDU1) and samples along a right edge of an adjacent left virtual pipeline data unit (VPDU2).
5. The method according to claim 4, wherein at least one of the at least one stencil area comprises all luma samples on the boundaries of adjacent virtual pipeline data units (VPDU1, VPDU2).
6. The method of claim 4, wherein the stencil region comprises a subset of luma samples on the boundary of adjacent virtual pipeline data units (VPDU1, VPDU2).
7. The method according to claim 1 or 2, wherein a gradient histogram is calculated for each chrominance component of the video image.
8. The method according to claim 7, wherein each gradient histogram of the chrominance component of the video image is calculated based on the luminance samples and the chrominance samples of the template area.
9. The method of claim 8, wherein the chroma samples of the template are multiplied by a compensation coefficient, the compensation coefficient depending on the chroma subsampling defined by a video image format.
10. The method according to any one of claims 1 to 9, wherein the midline sample position of at least one template area in the at least one template area on which the filter window is centered is determined to position the filter window at one midline sample position among a fourth number (N4) of midline sample positions of the template area.
11. The method according to any one of claims 1 to 9, wherein the midline sample position of at least one of the at least one template area around which the filter window is centered is determined to avoid any overlap between the filter windows.
12. The method according to any one of claims 1 to 11, wherein the filter windows have different sizes.
13. The method according to any one of claims 1 to 12, wherein the size of the filtering window depends on the size of the sample block and the sample availability of the template area.
14. An apparatus comprising means for carrying out one of the methods according to any one of claims 1 to 13.
15. A non-transitory storage medium carrying program code instructions for executing the method according to any one of claims 1 to 13.
16. An electronic device, include: processor; and a memory for storing instructions executable by the processor; The processor is configured to perform the method according to any one of claims 1 to 13.
17. A computer program product comprising instructions, which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 13.