Video picture data encoding / decoding

By analyzing gradients in a template area to determine intra prediction modes, the complexity of video codecs is reduced while maintaining coding efficiency for luma and chroma blocks, addressing the high complexity issue in modern video codecs.

JP2025532217AActive Publication Date: 2025-09-29DOLBY INTERNATIONAL AB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025517840
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-27
Filing Date
2023-04-24
Publication Date
2025-09-29
Estimated Expiration
2043-04-24

AI Technical Summary

Technical Problem

The complexity of modern video codecs is high when using Decoder-side Intra Mode Derivation (DIMD) for luma and chroma, which affects coding efficiency.

Method used

A method for determining intra prediction modes by analyzing gradients of samples in a template area around the block, selecting at most two angular intra prediction modes, and calculating a histogram of gradients using a filtering window centered on a sample position, reducing the complexity while maintaining coding efficiency.

Benefits of technology

Reduces complexity in video codecs by up to 20% in the encoder and 5% in the decoder without compromising coding efficiency for luma and chroma blocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532217000001_ABST
    Figure 2025532217000001_ABST
Patent Text Reader

Abstract

This disclosure relates to encoding / decoding a block of samples of a video picture from which a DIMD is derived, where luma samples and chroma samples of the sample block are predicted by filtering samples of at least one template area, the filtering filtering comprising filtering samples in at least one shaped template area using a filtering window centered on a sample position of an intermediate line of the at least one template area, wherein the integer number of sample positions of the intermediate line of the at least one template area on which the filtering window is centered is less than the total integer number of sample positions of the intermediate line of the at least one template area.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates generally to encoding and decoding video pictures. In particular, but not exclusively, the technical field of this disclosure relates to intra-predicted blocks of video pictures. [Background technology]

[0002] This section is intended to introduce the reader to various aspects of art that may be related to various aspects of at least one embodiment of the present disclosure, as described and / or claimed below. This discussion is believed to be helpful in providing the reader with background information to facilitate a better understanding of the various aspects of the present disclosure. Accordingly, it should be understood that these statements are to be read in this light, and not as admissions of prior art.

[0003] A pixel corresponds to the smallest display unit of the screen, which can consist of one or more light sources (one for monochrome screens, three for color screens).

[0004] A video picture, also called a frame or picture frame, includes at least one component (also called a picture component or channel) determined by a particular picture / video format that specifies all information related to pixel values ​​and all information that can be used by a display unit and / or other device to display and / or decode video picture data related to the video picture.

[0005] A video picture comprises at least one component that is usually represented in the form of an array of samples.

[0006] A monochrome video picture contains a single component, and a color video picture may contain three components.

[0007] For example, a color video picture may have a luma (or luminance) component and two chroma components if the picture / video format is the well-known (Y,Cb,Cr) format, or it may have three color components (one for red, one for green, and one for blue) if the picture / video format is the well-known (R,G,B) format.

[0008] Each component of a video picture may have a number of samples that corresponds to the number of pixels of the screen on which the video picture is intended to be displayed. In a variant, the number of samples contained in one component may be a multiple (or fraction) of the number of samples contained in other components of the same video picture.

[0009] For example, in the case of a video format having a luma component and two chroma components, such as the (Y,Cb,Cr) format, the chroma components may have half the number of samples in width and / or height as the luma component, depending on the color format considered.

[0010] A sample is the smallest visual information unit of the components that make up a video picture. Sample values ​​may be, for example, luma or chroma values, or color values ​​in (R,G,B) format.

[0011] A pixel value is the value of a pixel on the screen. A pixel value can be represented by one sample for a monochrome video picture or by multiple co-located samples for a color video picture. The co-located samples associated with a pixel refer to the samples that correspond to the pixel's location in the screen.

[0012] It is common to think of a video picture as being a set of pixel values, with each pixel represented by at least one sample.

[0013] A block of a video picture is a set of samples of one component of the video picture. A block of at least one luma sample, luma block for short, or a block of at least one chroma sample, chroma block for short, can be considered if the picture / video format is in the well-known (Y,Cb,Cr) format, or a block of at least one color sample, R,G,B for short.

[0014] At least one embodiment is not limited to a particular picture / video format.

[0015] In modern video compression systems such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), low-level and high-level picture partitioning is provided to divide a video picture into picture areas such as so-called coding tree units (CTUs), whose size is typically between 16x16 and 64x64 in the case of HEVC, and can be 32x32, 64x64, or 128x128 in the case of VVC.

[0016] The CTU division of a video picture forms a grid of fixed-size CTUs, or CTU grid, whose top and left boundaries are spatially coincident with the top and left boundaries of the video picture. The CTU grid represents a spatial partition of the video picture.

[0017] In VVC and HEVC, the CTU size (CTU width and CTU height) of all CTUs in a CTU grid is equal to the same default CTU size (default CTU width CTU DW and default CTU height CTU DH). For example, the default CTU size (default CTU height, default CTU width) is equal to 128 (CTU DW = CTU DH = 128). The default CTU size (height, width) is coded in the bitstream, for example, at the sequence level in the sequence parameter set (SPS).

[0018] The spatial location of a CTU within the CTU grid is determined from the CTU address ctuAddr, which defines the spatial location of the CTU's top left corner relative to the origin. As shown in Figure 1, the CTU address may define the spatial location relative to the top left corner of a higher-level spatial structure S that contains the CTU.

[0019] A coding tree is associated with each CTU to determine the tree partitioning of the CTU.

[0020] As shown in Figure 1, in HEVC, the coding tree is a quadtree division of CTUs, and each leaf is called a coding unit (CU). The spatial location of a CU within a video picture is defined by a CU index cuIdx, which indicates the spatial location from the top-left corner of the CTU. A CU is spatially partitioned into one or more prediction units (PUs). The spatial location of a PU within a video picture VP is defined by a PU index puIdx, which defines the spatial location from the top-left corner of the CTU, and the spatial location of an element of a partitioned PU is defined by a PU partition index puPartIdx, which defines the spatial location from the top-left corner of the PU. Each PU is assigned some intra- or inter-predicted data.

[0021] The intra or inter coding mode is assigned at the CU level, that is, the prediction parameters may vary from PU to PU, but each PU of a CU is assigned the same intra / inter coding mode.

[0022] A CU is also spatially partitioned into one or more transform units (TUs) according to a quadtree called a transform tree. A transform unit is a leaf of the transform tree. The spatial location of a TU within a video picture is defined by a TU index tuIdx, which defines the spatial location from the top-left corner of the CU. Each TU is assigned some transform parameters. A transform type is assigned at the TU level, and a 2D separate transform is performed at the TU level during encoding or decoding of a picture block.

[0023] The PU partition types present in HEVC are shown in Figure 2. They include square partitions (2Nx2N and NxN), which are the only ones used for both intra-predicted and inter-predicted CUs, non-square partitions (2NxN, Nx2N, used only for inter-predicted CUs), and asymmetric partitions (used only for inter-predicted CUs). For example, PU type 2NxnU represents an asymmetric horizontal partitioning of a PU, with the smaller partition above the PU. According to another example, PU type 2NxnL represents an asymmetric vertical partitioning of a PU, with the smaller partition to the left of the PU.

[0024] As shown in Figure 3, in VVC, the coding tree starts with a root node, i.e., CTU. Then, a quad-tree (or quaternary tree) partition divides the root node into four nodes corresponding to four equally sized sub-blocks (solid lines). The quad-tree leaves can then be further partitioned by so-called multi-type trees. A multi-type tree includes a binary or ternary tree partition according to one of four partition modes, as shown in Figure 4. These partition types are a vertical binary partition mode, denoted SBTV, and a horizontal binary partition mode, denoted SBTH, as well as a vertical third partition mode STTV and a horizontal third partition mode STTH.

[0025] The leaves of the coding tree of a CTU are CUs in the case of a joint coding tree shared by the luma and chroma components.

[0026] In contrast to HEVC, in VVC, in most cases, CUs, PUs, and TUs have equal size, that is, coding units are generally not partitioned into PUs or TUs except in some specific coding modes.

[0027] 5 and 6 provide an overview of the video encoding / decoding methods used in current video standard compression systems, such as HEVC or VVC.

[0028] FIG. 5 shows a schematic block diagram of the steps of a method 100 for encoding a video picture VP according to the prior art.

[0029] In step 110, the video picture VP is partitioned into blocks of samples, and partitioning information data is signaled in the bitstream. Each block contains samples of one component of the video picture VP. Thus, a block contains samples of each component that defines the video picture VP.

[0030] For example, in HEVC, a picture is divided into coding tree units (CTUs). Each CTU can be further subdivided using quadtree partitioning, with each leaf of the quadtree being represented as a coding unit (CU). The partitioning information data may then include data describing the CTUs and the quadtree partitioning of each CTU.

[0031] Each block of samples, or block for short, can then be either a CU (if the CU has a single PU) or a PU of a CU.

[0032] Each current block is coded along a coding loop, also called "in loop", in either intra or inter prediction mode.

[0033] Intra prediction (step 120) uses intra prediction data. Intra prediction consists of predicting the current block using intra prediction blocks based on already coded, decoded, and reconstructed samples located around the current block, typically in a so-called L-shaped template area defined above, to the left, and above-left of the current block. Intra prediction is performed in the spatial domain.

[0034] In inter-prediction mode, motion estimation (step 130) and motion compensation (step 135) are performed. Motion estimation searches for reference blocks that are good predictors of the current block among one or more reference pictures used to predictively code the current video picture. In unidirectional motion estimation / compensation, candidate reference blocks belong to a single reference picture in a reference picture list, denoted as L0 or L1, while in bidirectional motion estimation / compensation, candidate reference blocks are derived from reference blocks in reference picture list L0 and reference blocks in reference picture list L1.

[0035] For example, a suitable predictor for a current block is a candidate reference block that is similar to the current block, and may also correspond to a reference block that provides a suitable trade-off between its similarity to the current block and the rate cost of the motion information required to direct its use for temporal prediction of the current block.

[0036] The output of the motion estimation step 130 is inter-prediction data, which includes motion information related to the current block and other information used to obtain the same prediction block on the encoding / decoding side. Typically, the motion information includes one motion vector and a reference picture index in the case of unidirectional estimation / compensation, and two motion vectors and two reference picture indexes in the case of bidirectional estimation / compensation. Next, motion compensation (step 135) obtains a prediction block using the motion vector and reference picture index determined by the motion estimation step 130. Essentially, a reference block belonging to the selected reference picture and indicated by the motion vector can be used as the prediction block for the current block. Furthermore, since motion vectors are expressed as fractions of integer pixel positions (known as sub-pel precision motion vector representation), motion compensation generally involves spatial interpolation of several reconstructed samples of the reference picture to calculate the prediction block.

[0037] Prediction information data is signaled in the bitstream and may include prediction mode (intra or inter or skip), intra / inter prediction data, and other information used to obtain the same prediction block on the decoding side.

[0038] The method 100 selects one prediction mode (intra or inter prediction mode) by optimizing the rate-distortion trade-off, taking into account the encoding of a prediction residual block, calculated for example by subtracting a candidate prediction block from a current block, and the signaling of prediction information data required to determine the candidate prediction block on the decoding side.

[0039] Usually, the best prediction mode is the best coding mode p for the current block. * is given as the prediction mode of

number

number

[0040] The current block is usually coded from a prediction residual block PR. More precisely, the prediction residual block PR is calculated, for example, by subtracting the best prediction block from the current block. The prediction residual block PR is then transformed (step 140), for example, by using a DCT (Discrete Cosine Transform) or DST (Discrete Sine Transform) type transform or any other suitable transform, and the obtained transformed coefficient block is quantized (step 150).

[0041] In a variant, the method 100 may skip the transformation step 140 and apply quantization (step 150) directly to the prediction residual block PR, according to a so-called transform skip coding mode.

[0042] The quantized transform coefficient block (or quantized prediction residual block) is entropy coded into a bitstream (step 160).

[0043] The quantized transform coefficient block (or quantized residual block) is then inverse quantized (step 170) and inverse transformed (step 180) as part of the coding loop to obtain a decoded prediction residual block. The decoded prediction residual block and the prediction block are then combined, typically added together, to provide a reconstructed block.

[0044] To encode the current block of the video picture VP, other information data may also be entropy encoded in step 160 .

[0045] To reduce compression artifacts, an in-loop filter (step 190) may be applied to the reconstructed picture (including the reconstructed blocks). Loop filters may be applied after all picture blocks have been reconstructed. For example, they may consist of a deblocking filter, a Sample Adaptive Offset (SAO) or an Adaptive Loop Filter (ALF).

[0046] The reconstructed block or the filtered reconstructed block forms a reference picture that can be stored in a Decode Picture Buffer (DPB) so that it can be used as a reference picture for encoding the next current block of the video picture VP or of the next video picture to be encoded.

[0047] FIG. 6 is a schematic block diagram of the steps of a method 200 for decoding a video picture VP according to the prior art.

[0048] At step 210, partitioning information data, prediction information data and quantized transform coefficient blocks (or quantized residual blocks) are obtained by entropy decoding a bitstream of coded video picture data, for example, the bitstream generated according to method 100.

[0049] To decode the current block of the video picture VP from the bitstream, other information data may also be entropy decoded.

[0050] In step 220, the reconstructed picture is divided into current blocks based on the partitioning information. Each current block is entropy decoded from the bitstream along a decoding loop, also called "in-loop." Each decoded current block is either a quantized transform coefficient block or a quantized prediction residual block.

[0051] In step 230, the current block is dequantized and possibly inverse transformed (step 240) to obtain a decoded prediction residual block.

[0052] On the other hand, the prediction information data is used to predict the current block. The prediction block is obtained through its intra prediction (step 250) or its motion compensated temporal prediction (step 260). The prediction process performed on the decoding side is the same as that on the encoding side.

[0053] The decoded prediction residual block and the prediction block are then combined, typically added together, to provide a reconstructed block.

[0054] In step 270, an in-loop filter may be applied to the reconstructed picture (including the reconstructed block), and the reconstructed block and the filtered reconstructed block form a reference picture that may be stored in the decoded picture buffer (DPB), as described above (FIG. 5).

[0055] In VVC, motion information is stored for each 4x4 block in each video picture. This means that when a reference picture is stored in the decoded picture buffer (DPB in FIG. 5 or FIG. 6), the motion vectors and reference picture indices used for temporal prediction of video picture blocks are stored in 4x4 block units. They can be the temporal prediction of motion information for encoding / decoding subsequent inter-predicted video pictures.

[0056] To reduce cross-component redundancy, VVC defines a so-called Linear-Model based intra-prediction mode (LM mode).

[0057] One of them is the so-called Cross-Component Linear Model (CCLM), which derives a linear model-based predictor, or LM predictor for short. It uses a linear model to predict the co-located reconstructed luma samples rec' of the same CU. L Based on (i,j), the predicted chroma sample pred C (i,j) contains:

number

[0058] Reconstructed luma samples rec' L (i,j) is downsampled after filtering to fit the CU chroma size.

[0059] VVC specifies three CCLM modes, denoted CCLM_LT, CCLM_T, and CCLM_L. These three CCLM modes differ in the location of the reference samples used for linear parameter derivation. Reference samples from the upper boundary are included in CCLM_T mode, and reference samples from the left boundary are included in CCLM_L mode. In CCLM_LT mode, reference samples from both the upper and left boundaries are used.

[0060] Generally, the CCLM mode prediction process consists of three steps: 1) The luma block and its neighboring reconstructed luma samples rec' to fit the size of the corresponding chroma block. L Downsampling of (i,j); 2) Linear parameter derivation based on neighboring reconstructed luma samples; 3) Application of Equation (10) to generate chroma intra prediction samples (prediction chroma blocks).

[0061] Another LM mode is the so-called Multi-Model Linear Model (MMLM) [K. Zhan et al., “Enhanced "Cross-component Linear Model Intra Prediction," JVET-D0110, San Diego, October 2016. MMLM is a method for predicting the co-located reconstructed luma samples using more than one model. L(i,j) is derived from (i,j), so it is an extension of CCLM. In MMLM, neighboring reconstructed luma samples and neighboring chroma samples are classified into several groups, and each group is used as a training set to derive linear parameters of a linear model (i.e., a specific α and β are derived for a specific group). Furthermore, samples of the current luma block are also classified based on the same rules for classifying neighboring luma samples. Neighboring samples are classified into, for example, M groups. MMLM methods with M=2 and M=3 are designed as two additional ML modes for chroma, named MMLM2 and MMLM3, in addition to the original LM mode (CCLM). The encoder selects the optimal LM mode in the rate / distortion optimization process and signals the best LM mode. For example, when M is equal to 2, the threshold is calculated as the average value of neighboring reconstructed luma samples. Neighboring reconstructed samples rec' that are less than or equal to the threshold are classified into M groups. L [x,y] is classified into group 1, while the adjacent reconstructed sample rec' is larger than the threshold. L [x,y] is classified into group 2.

[0062] The two LM predictors are then derived as follows:

number

[0063] VVC further defines decoder-side intra mode derivation (DIMD) modes for both luma and chroma samples. DIMD modes are not linear model-based modes, or non-LM modes for short, that is, intra prediction modes that do not reference a linear model, such as planar or direct modes.

[0064] DM is an intra prediction mode that uses the average sample value of reference samples above and to the left of the block for prediction generation.

[0065] The planar mode is a weighted average of four reference sample values ​​(picked as orthogonal projections of the sample to be predicted in the upper and left reconstructed areas). The horizontal and vertical modes use copies of the left and upper reconstructed samples, respectively, without interpolation to predict the sample row and sample column, respectively.

[0066] For luma sample prediction, the use of DIMD luma mode is signaled in the bitstream by a single flag, and the intra predictor is not explicitly signaled in the bitstream.

[0067] FIG. 7 schematically depicts a block diagram of a method 300 for deriving a DIMD luma mode according to the prior art.

[0068] Essentially, the DIMD luma mode (predictor) is derived by using gradient analysis of neighboring reconstructed luma samples, i.e., luma samples located in an L-shaped template area defined around the current block of samples. If the DIMD luma mode is not enabled, the intra-prediction mode can be parsed from the bitstream similarly to a conventional intra-prediction mode. The DIMD luma mode is implicit. Therefore, the DIMD luma mode is derived during the reconstruction process at the encoder and decoder sides alike.

[0069] In step 310, the intra-top and intra-left available samples around the luma block to be predicted are determined, thus determining the L-shaped template area.

[0070] For example, as shown in FIG. 8, an L-shaped template area T is defined that is three samples wide (width or height) and consists of the reconstructed luma samples on the left, top, and top left of the reconstruction area R adjacent to the current block B (current CU).

[0071] In step 320, a Histogram of Gradients (HoG) is constructed as follows:

[0072] First, the samples of the L-shaped template area T are filtered, using a filtering window W centered on a sample position of an intermediate line of the L-shaped template area T. An amplitude and angle of intensity direction (orientation) is assigned to each intermediate line sample of the L-shaped template area T.

[0073] For example, if an edge detection filter (3x3 horizontal and vertical Sobel filter) is used for filtering the samples of an L-shaped template area T, the amplitude and angle of the luminance direction are given as:

number

[0074] Then, the HoG is calculated, where each entry (angle) corresponds to an angular intra prediction mode, and the cumulative amplitude is stored.

[0075] In step 330, at most two angular intra-prediction modes M1 and M1 are selected from the HoG.

[0076] For example, as shown in FIG. 9, the two most commonly represented angular intra-prediction modes M1 and M2 have the largest histogram amplitude values.

[0077] Zero, one, or two angular intra prediction modes may be selected from the HoG. For example, if none of the cumulative amplitudes of the HoG exceeds a threshold, no angular intra prediction mode is selected.

[0078] If no angular intra-prediction mode is selected, then in step 340, planar mode is the intra-prediction mode used as the intra-predictor for the luma portion of the current block to be predicted.

[0079] If a single angular intra-prediction mode is selected, then in step 350, this angular intra-prediction mode is the intra-prediction mode that is used as the intra-predictor for the luma portion of the current block to be predicted.

[0080] If two angular intra prediction modes are selected (M1 and M2), then in step 360, a DIMD luma mode is determined by blending (mixing, fusing) three luma predictors: two intra predictors derived from the two selected angular intra prediction modes M1 and M2, and a planar predictor [M. Abdoli et al, “Non-CE3: Decoder-side Intra Mode Derivation with Prediction Fusion Using Planar”, JVET-O0449, Gothenburg, July 2019].

[0081] As depicted in FIG. 10, blending is defined as a weighted average of these three intra predictors using three weights w1, w2, w3, which are derived from the ratio of the selected angular intra prediction amplitudes as follows (step 361):

number

[0082] FIG. 11 schematically depicts a block diagram of a method 400 for deriving DIMD chroma modes according to the prior art.

[0083] Essentially, to derive the DIMD chroma modes, a similar method used to derive the DIMD luma modes is applied to gradient analysis of collocated reconstructed luma samples, i.e., collocated reconstructed luma samples within an L-shaped template area defined around the current block of chroma samples.

[0084] The use of DIMD chroma mode is signaled in the bitstream by a single flag, and the intra predictor is not explicitly signaled in the bitstream but is derived by using gradient analysis of co-located reconstructed luma samples. DIMD chroma mode is implicit. Thus, the intra predictor is derived from DIMD chroma mode during the reconstruction process at the encoder and decoder sides alike.

[0085] At step 410, method 400 determines whether reconstructed luma samples are available. These reconstructed luma samples are collocated with the chroma block to be predicted or with an L-shaped template area around the chroma block. The L-shaped template area is formed from the available L-shaped template areas.

[0086] In step 420, a histogram of gradients (HoG) is constructed for the available collocated reconstructed luma samples in the L-shaped area, similar to step 320. Specifically, horizontal and vertical gradients are calculated for each collocated reconstructed luma sample (gray circles in FIG. 12 extracted from JVET-Y0092), and a HoG is constructed from the horizontal and vertical gradients.

[0087] In step 430, the amplitudes of available chroma samples in an L-shaped area around the chroma block are accumulated into a HoG, with each entry corresponding to an orientation, i.e., angular intra prediction mode (same as its counterpart in the luma HoG).

[0088] In a variant, the HoG may be calculated by accumulating amplitude values ​​over the reconstructed Y, Cb, and Cr samples of an L-shaped template area (FIG. 13). That is, the HoG is derived by fusing a first HoG calculated from adjacent collocated luma samples of the L-shaped template area (left part of FIG. 13), a second HoG calculated from Cb samples of the L-shaped template area (center part of FIG. 13), and a third HoG calculated from adjacent reconstructed Cr samples (right part of FIG. 13).

[0089] In step 440, the angular intra prediction mode corresponding to the largest histogram amplitude value is selected from HoG as the first DIMD chroma mode.

[0090] If the DIMD chroma mode is not direct mode (DM), the method 400 ends and the DIMD chroma mode is the first DIMD chroma mode.

[0091] Direct mode (DM) is a chrominance intra-prediction mode that corresponds to the intra-prediction mode of co-located reconstructed luma samples.

[0092] If the first DIMD chroma mode is direct mode (DM), then in step 450, a second DIMD chroma mode is selected as the angular intra prediction mode corresponding to the second largest histogram amplitude value.

[0093] If the first DIMD chroma mode is not equal to the second DIMD chroma mode, the method 400 ends and the DIMD chroma mode is the second DIMD chroma mode.

[0094] If the first DIMD chroma mode is equal to the second DIMD chroma mode, then in step 460, the DIMD chroma mode is Direct Coding (DC).

[0095] A nonlinear model-based predictor, or non-LM predictor for short, for a chroma block is then derived from a non-LM mode selected from among the five default modes and the DIMD chroma mode, where the chroma predictor is derived from these non-LM modes, and the selected non-LM mode corresponds to the non-LM predictor that minimizes the rate-distortion tradeoff.

[0096] The five default modes are DC, planar, direct mode, horizontal angle predictor, and vertical angle predictor.

[0097] The non-LM predictors of the chroma blocks derived from the selected non-LM mode are then blended (mixed) with the LM predictors derived from the MMLM mode as follows:

number

[0098] If the non-LM mode is selected, one flag is signaled to indicate whether fusion is applied or not.

[0099] In a variant, for I slices, i.e., intra-coded slices, non-LM predictors derived from DM mode, the four default modes, and DIMD chroma mode can be fused with LM predictors derived from LM mode, while for non-I slices, only non-LM predictors derived from DIMD chroma mode can be fused with LM predictors derived from LM mode using equal weights.

[0100] Chroma blend mode can be used with DIMD chroma (if selected as the best non-LM mode), but other non-LM modes may be used instead.

[0101] Using DIMD for luma and chroma improves the coding efficiency of video pictures, but incurs approximately 20% additional complexity in the encoder and 5% in the decoder in all intra conditions. Summary of the Invention [Problem to be solved by the invention]

[0102] The problem to be solved by this disclosure is to reduce the complexity of modern video codecs while maintaining their coding efficiency when DIMD is used for luma and chroma.

[0103] At least one embodiment of the present disclosure has been conceived with the above in mind. [Means for solving the problem]

[0104] The following section presents a simplified summary of at least one embodiment in order to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the embodiments. It is not intended to identify key or key elements of the embodiments. The following summary merely presents some aspects of at least one embodiment in a simplified form as a prelude to the more detailed description provided elsewhere herein.

[0105] According to a first aspect of the present disclosure, there is provided a method for decoding a block of samples of a video picture, the method comprising the step of determining an intra prediction mode for decoding the block of samples according to an analysis of gradients of samples located in at least one template area defined around the block of samples; The step of determining an intra-prediction mode includes: calculating a histogram of gradients by filtering samples of the at least one template area, each entry of the histogram of gradients corresponding to an angular intra-prediction mode, the filtering using a filtering window centered on a sample position of a mid-line of the at least one template area; selecting at most two angular intra prediction modes by comparing amplitudes of angular intra predictions of said histograms of gradients; determining the intra prediction mode from the at most two selected angular intra prediction modes; The integer number of sample positions of the intermediate line of the at least one template area around which the filtering window is centered is less than the total integer number of sample positions of the intermediate line of the at least one template area.

[0106] According to a second aspect of the present disclosure, there is provided a method for encoding a block of samples of a video picture, the method comprising the step of determining an intra prediction mode for decoding the block of samples according to an analysis of gradients of samples located in at least one template area defined around the block of samples, The step of determining an intra-prediction mode includes: calculating a histogram of gradients by filtering samples of the at least one template area, each entry of the histogram of gradients corresponding to an angular intra-prediction mode, the filtering using a filtering window centered on a sample position of a mid-line of the at least one template area; selecting at most two angular intra prediction modes by comparing amplitudes of angular intra predictions of said histograms of gradients; determining the intra prediction mode from the at most two selected angular intra prediction modes; The integer number of sample positions of the intermediate lines of the at least one template area on which the filtering window is centered is less than the total integer number of sample positions of the intermediate lines of the at least one template area.

[0107] In one embodiment, the histogram of gradients for chroma components of the video picture is calculated by filtering only chroma samples of the at least one template area.

[0108] In one embodiment, at least one of the at least one template area includes samples along a bottom edge of an adjacent upper Virtual Pipeline Data Unit and samples along a right edge of an adjacent left Virtual Pipeline Data Unit.

[0109] In one embodiment, at least one of the at least one template area includes all luma samples on a boundary of the virtual pipeline data unit in the vicinity.

[0110] In one embodiment, the template area includes a subset of luma samples on a boundary of the virtual pipeline data unit in the vicinity.

[0111] In one embodiment, one histogram of gradients is calculated for each chroma component of the video picture.

[0112] In one embodiment, each histogram of gradients for a chroma component of the video picture is calculated from luma samples and chroma samples of a template area.

[0113] In one embodiment, the chroma samples of the template area are multiplied by a compensation factor that depends on the chroma subsampling defined by the video picture format.

[0114] In one embodiment, the sample position of the intermediate line of at least one of the at least one template area in which the filtering window is centered is determined to place the filtering window at one of a fourth number (N4) of sample positions of the intermediate line of the template area.

[0115] In one embodiment, sample positions of the intermediate line of at least one of the at least one template area in which the filtering window is centered are determined such that there is no overlap between the filtering windows.

[0116] In one embodiment, the filtering windows have different sizes.

[0117] In one embodiment, the size of the filtering window depends on the size of the block of samples and the availability of samples in the template area.

[0118] According to a third aspect of the present disclosure, there is provided an apparatus having means for carrying out one of the methods according to the first and / or second aspects of the present disclosure.

[0119] According to a fourth aspect of the present disclosure, there is provided a non-transitory storage medium bearing program code instructions for performing a method according to the first and / or second aspects of the present disclosure.

[0120] According to a fifth aspect of the present disclosure, there is provided an electronic device including a processor and a memory storing instructions executable by the processor, the processor being configured to perform a method according to the first and / or second aspect of the present disclosure.

[0121] According to a sixth aspect of the present disclosure, there is provided a computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to the first and / or second aspect of the present disclosure.

[0122] At least one particular feature of the embodiment, as well as at least one other object, advantage, feature, and use of the embodiment, will become apparent from the following description of the example, read in conjunction with the accompanying drawings. [Brief explanation of the drawings]

[0123] [Figure 1] 1 shows an example of a coding tree unit according to HEVC. [Figure 2] 1 illustrates an example of partitioning coding units into prediction units according to HEVC. [Figure 3] An example of CTU division according to VVC is shown below. [Figure 4] Examples of partitioning modes supported by multi-type tree partitioning according to VVC are shown below. [Figure 5] 1 shows a schematic block diagram of the steps of a method 100 for encoding a video picture VP according to the prior art. [Figure 6] 2 shows a schematic block diagram of the steps of a method 200 for decoding a video picture VP according to the prior art. [Figure 7] 3 shows a block diagram of a method 300 for deriving a DIMD luma mode according to the prior art. [Figure 8] Represents an L-shaped template area. [Figure 9] This represents an example of HoG. [Figure 10] 1 illustrates a schematic diagram of a method for blending multiple luma predictions according to the prior art. [Figure 11] 4 shows a schematic block diagram of a method 400 for deriving DIMD chroma modes according to the prior art. [Figure 12] 1 illustrates schematically an example of co-located luma samples of a chroma block according to the prior art; [Figure 13] 10A and 10B schematically illustrate another example of co-located luma samples of a chroma block according to the prior art; [Figure 14] 5A and 5B illustrate a block diagram of a method 500 for deriving a DIMD mode according to an embodiment. [Figure 15] 5A and 5B schematically illustrate the position of the filtering window (step 510) according to an embodiment. [Figure 16] 5A and 5B schematically illustrate the position of the filtering window (step 510) according to an embodiment. [Figure 17] 5A and 5B schematically illustrate the position of the filtering window (step 510) according to an embodiment. [Figure 18] 1 illustrates an L-shaped template area around a VPDU and the current VPDU according to an embodiment. [Figure 19]10 illustrates an L-shaped template area when VPDUs are used according to an embodiment. [Figure 20] 10 illustrates an L-shaped template area when VPDUs are used according to an embodiment. [Figure 21] 6 illustrates a block diagram of a method 600 for constructing HoG for chroma DIMD derivation according to an embodiment. [Figure 22] 1 depicts a schematic block diagram of an example system in which various aspects and embodiments may be implemented. DETAILED DESCRIPTION OF THE INVENTION

[0124] Reference will now be made to the accompanying drawings which illustrate, by way of example, embodiments of the present disclosure, in which similar or identical elements are referred to with the same reference numerals, and in which:

[0125] At least one embodiment will be described more fully hereinafter with reference to the accompanying drawings, in which at least one example of an embodiment is shown. However, embodiments may be embodied in many alternate forms and should not be construed as limited to the examples set forth herein. Accordingly, it should be understood that there is no intention to limit the embodiments to the particular forms disclosed. On the contrary, the present disclosure is intended to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure.

[0126] At least one of the aspects relates generally to the encoding and decoding of video pictures, one other aspect relates generally to the provision or transmission of encoded bitstreams, and another aspect relates to the reception / access of decoded bitstreams.

[0127] At least one of the embodiments is described for encoding / decoding video pictures, but extends to encoding / decoding video pictures (a sequence of pictures), as each video picture is encoded / decoded sequentially as described below.

[0128] Furthermore, at least one embodiment may be adapted to MPEG standards such as, but not limited to, AVC (ISO / IEC 14496-10 Advanced Video Coding for generic audio-visual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en), and VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-I / en), for example, AV1 (Audio Media Video 1, http: / / aomedia.org / av1 / specification / ). At least one embodiment may apply to existing or future-developed extensions of such standards and recommendations. Unless otherwise indicated or technically prohibited, aspects described in this disclosure may be used independently or in combination.

[0129] Generally speaking, the present disclosure relates to decoding a block of samples of a video picture in which a DIMD mode is derived by filtering samples of at least one template area for luma samples and chroma samples of a sample block to be predicted, the filtering filtering comprising filtering samples in at least one shaped template area using a filtering window centered on a sample position of an intermediate line of the at least one template area, wherein the integer number of sample positions of the intermediate line of the at least one template area on which the filtering window is centered is less than the total integer number of sample positions of the intermediate line of the at least one template area.

[0130] This reduces the resource requirements (memory and computational power and complexity) for deriving DMID modes for both the luma and chroma samples of the block to be predicted compared to the resources required by current DMID mode derivation methods.

[0131] The following embodiments are described using an L-shaped template area consisting of reconstructed samples to the left, above, and to the top left of the reconstructed area adjacent to the current block, but the present disclosure extends to a template area consisting only of reconstructed samples to the left of the reconstructed area adjacent to the current block, or to a template area consisting only of reconstructed samples above the reconstructed area adjacent to the current block.

[0132] FIG. 14 schematically illustrates a block diagram of a method 500 for deriving a DIMD mode according to an embodiment.

[0133] The method 500 can be applied in both the method 100 (encoding) and the method 200 (decoding) for intra prediction of a block.

[0134] In step 510, the sample positions of the intermediate line of the L-shaped template area T1 are determined. The integer number N1 of sample positions of the intermediate line of the L-shaped template area T1 on which the filtering window is centered is less than the total number of sample positions of the intermediate line of the L-shaped template area.

[0135] In step 520, a luma DIMD mode is derived from method 300 with filtering windows centered on each of the N1 intermediate line sample locations of the L-shaped template area T1.

[0136] In one variation of step 520, sample locations of intermediate lines of L-shaped template area T2 of collocated reconstructed luma samples are determined. Sample locations of intermediate lines of L-shaped template area T3 of available chroma samples are determined. The number N2 or N3 of sample locations of intermediate lines of L-shaped template area T2 or T3 around which the filtering window is centered is less than the total number of sample locations of intermediate lines of L-shaped template area T2 or T3. A chroma DIMD mode is derived from method 400 in which the filtering window is centered on each of the sample locations of the N2 intermediate lines of L-shaped template area T2 and each of the sample locations of the N3 intermediate lines of L-shaped template area T3.

[0137] In one embodiment of step 510, the sample position of the intermediate line of at least one of the at least one L-shaped template areas (T1, T2, T3) in which the filtering window is centered is determined to be located at one of the number (N4>=2) of sample positions of the intermediate line of the L-shaped template area T1, T2, or T3.

[0138] 15 shows a schematic example of the position of the filtering window centered on the determined intermediate line sample position (open circle) when N1 = 2. The filtering window is represented by a dashed line.

[0139] In one embodiment of step 510, sample positions of the intermediate line of at least one of the at least one L-shaped template area (T1, T2, T3) in which the filtering window is centered are determined such that the filtering windows do not overlap.

[0140] 16 shows a schematic example of the position of the filtering window centered on the determined intermediate line sample position (open circle). The filtering window is represented by a dashed line.

[0141] In one embodiment of step 510, sample positions of at least one intermediate line of at least one L-shaped template area (T1, T2, T3) in which a filtering window is centered are determined so that the filtering windows can have a common boundary. The sample positions of these intermediate lines of the L-shaped template area are then determined from among the top left, upper side, and left side positions of the L-shaped template area.

[0142] 17 shows a schematic example of the position of the filtering window centered on the determined midline sample position (open circle), the filtering window being represented by a dashed line.

[0143] In one embodiment of method 500, to compensate for the reduced number of filtered samples (i.e., the reduced number of HoG entries), the filtering window size is larger than 3x3. For example, a filter of size 5x5 is used.

[0144] In one embodiment of the method 500, the filtering windows have different sizes.

[0145] In one embodiment of the method 500, the filtering window size depends on the block size to be predicted and the availability of samples in the L-shaped template area.

[0146] For example, depending on sample availability, the filtering window size may be increased for L-shaped template areas in blocks of larger dimensions (possibly when the block size exceeds a predetermined size, eg, 8x8).

[0147] The above embodiment for determining the sample position of the middle line of the L-shaped template area in which the filtering window is centered and the filtering window size may be combined.

[0148] For example, the filtering window may be 5x5 in size for samples within an L-shaped template area adjacent to the longer block size, and the allowed intermediate line sample positions may be at specific locations, or may be spaced at equally spaced locations, or may be located at locations where the filtering windows do not overlap.

[0149] Because the reduced number of sample positions for the intermediate lines (step 510) may be less representative of the directions present in the neighborhood of the current block to be predicted, in one embodiment, the blending weight derivation (for luma) is modified so that instead of calculating weights w1 and w2 according to equation (1), weights w1 and w2 are calculated as follows:

[0150] If the amplitude of the first angular intra prediction mode M1 is greater than twice the amplitude of the second angular intra prediction mode M2, the second weight w2 is equal to 0, for example, w1=43 / 64 and w3=21 / 64.

[0151] In one variant, the third weight w3 is equal to 0 and w1 is 64 / 64.

[0152] The second weight w2 is not equal to 0 if the amplitude of the first angular intra prediction mode M1 is less than or equal to twice the amplitude of the second angular intra prediction mode M2.

[0153] For example, w1=w2=w3=21 / 64.

[0154] In one variant, the third weight is equal to 0, for example w1=32 / 64 and w3=32 / 64.

[0155] As discussed above in the introduction, the chroma samples within the L-shaped template area and the co-located reconstructed luma samples within the L-shaped template area are used to construct the HoG. This requires that the luma samples be coded first before coding the chroma with the DIMD chroma tool, adding latency to the encoding method 100 and the decoding method 200. Furthermore, if a DIMD luma mode is not enabled to predict the luma block, the method for deriving the DIMD chroma mode must calculate the HoG from the co-located reconstructed luma samples to obtain the DIMD chroma predictor anyway. Depending on the hardware architecture design, the calculation of the HoG from the co-located reconstructed luma samples may be repeated between the derivation of the DIMD luma mode and the DIMD chroma mode, which may be redundant.

[0156] In one embodiment of the method 500, to derive the DIMD modes for the chroma components of a video picture, the HoG is calculated only according to the filtered chroma samples of the L-shaped template area.

[0157] This embodiment is only applicable to deriving DIMD chroma modes.

[0158] This embodiment is advantageous because it avoids encoding and decoding luma samples before coding chroma samples by using DIMD chroma mode. This embodiment is also advantageous in its lower complexity, lower latency, computational power, and memory requirements compared to methods that use both co-located reconstructed luma and chroma samples.

[0159] In VVC, a Virtual Pipeline Data Unit (VPDU) is a concept of a non-overlapping unit of samples, typically sized 64x64 for luma and 32x32 for chroma. The goal is to ensure that a VPDU is fully processed before starting processing of the next VPDU so that the memory footprint of the hardware implementation is kept reasonable. VPDUs have a strong influence on tool design.

[0160] Therefore, if a VPDU is considered in the design of a chroma DIMD, and if the construction of the HoG for deriving a DIMD predictor uses luma samples as well as chroma samples present in the available template area (e.g., as discussed above), these luma samples are obtained from the VPDUc, which is now a VPDU, so that the luma samples are fully processed and available.

[0161] In one embodiment of method 500, at least one of at least one L-shaped template area is used that includes luma samples along the lower VPDU boundary of the adjacent top VPDU (VPDU1 in FIG. 18) and luma samples along the right VPDU boundary of the adjacent left VPDU (VPDU2 in FIG. 18).

[0162] This embodiment reduces the latency of the method 500 .

[0163] In one variation, at least one of the at least one L-shaped template area includes all luma samples (i.e., 64x2 in the VVC design or 64x3 when the width is 3) on the boundary of neighboring VPDUs (VPDU1 and VPDU2) (Figure 19). All these luma samples are used in constructing the HoG instead of using neighboring co-located luma samples.

[0164] In one variation, a subset of luma samples on the boundaries of neighboring VPDUs (VPDU1 and VPDU2) are used (FIG. 20). These luma samples are used in constructing the HoG instead of using adjacent co-located luma samples.

[0165] In one variant, a subset of luma samples belonging to a three-sample wide boundary of an adjacent VPDU is used, where this subset corresponds to luma samples that are collocated with (part of) the orthogonal projection of a chroma block on the adjacent VPDU, as shown in Figures 18, 19, and 20.

[0166] Only the luma samples of the current VPDU, VPDUc, are used in constructing the HoG to derive the chroma DIMD predictor. In the absence of neighboring VPDUs, only the chroma samples present in the L-shaped template area are used to construct the HoG.

[0167] In one embodiment of the method 500, one HoG is calculated for each chroma component of a video picture.

[0168] This embodiment improves the latency of the method 500, as it allows for parallelization to compute the HoG for each chroma component.

[0169] Furthermore, this embodiment improves signal adaptation by separating the HoG construction for chroma components and, consequently, the final intra prediction mode that is assigned independently to each chroma component when chroma DIMD is enabled.

[0170] This embodiment also provides additional flexibility in design, since the final intra-prediction mode selected for each chroma component can be different. This embodiment does not require a chroma coding block or other coding tools that operate on the components individually. An advantage is that no extra signaling (beyond the DIMD chroma flag) is required for the modes of both components.

[0171] FIG. 21 schematically depicts a block diagram of a method 600 for constructing HoG for chroma DIMD derivation according to an embodiment.

[0172] In this embodiment, the available chroma samples within the L-shaped template area and the available reconstructed co-located luma samples within the L-shaped template area are used in constructing each chroma component HoG.

[0173] However, since non-4:4:4 signals have fewer chroma samples than luma samples, a compensation factor is introduced into the construction of the HoG.

[0174] In step 610, the current filtering window W is considered.

[0175] In step 620, horizontal and vertical gradients Dx and Dy are calculated by applying horizontal and vertical edge detection filters, such as horizontal and vertical 3x3 Sobel filters.

[0176] In step 635, a compensation factor is determined according to the signal chroma subsampling defined by the video picture format.

[0177] For example, the compensation factor is equal to 2 for a 4:2:0 signal, 1.5 for a 4:2:2 signal, and 1 for a 4:4:4 signal.

[0178] In one variant, this compensation factor is a multiple of 2 for 4:2:0 signals, 1.5 for 4:2:2 signals and 1 for 4:4:4 signals.

[0179] Alternatively, for a 4:2:0 signal, one of the two (equally spaced) luma samples of the L-shaped template area is used as the center of the filtering window to construct the HoG (similar to Figure 15), so that the number of luma samples of the intermediate line processed is equal to the number of chroma samples processed (for each chroma component).

[0180] In one variation of this embodiment of Figure 21, only the available chroma samples within the L-shaped template area are used in constructing each chroma component HoG. Filtering window positions are then added to make up for the missing samples. Typically, the third row or column of chroma samples is also used as the center filtering window position to construct each chroma HoG.

[0181] FIG. 22 depicts a schematic block diagram of an example system 700 in which various aspects and embodiments may be implemented.

[0182] System 700 may be embodied as one or more devices including various components described below. In various embodiments, system 700 may be configured to implement one or more of the aspects described in this disclosure.

[0183] Examples of devices that may form all or part of system 700 include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital televisions, personal video recording systems, Internet appliances, Internet vehicles and their associated processing systems, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process output from video decoders, pre-processors that provide input to video encoders, web servers, video servers (e.g., broadcast servers, video-on-demand servers, or web servers), still or video cameras, encoding or decoding chips, or other communications devices. The elements of system 700, singly or in combination, may be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 700 may be distributed across multiple ICs and / or discrete components. In various embodiments, system 700 may be communicatively coupled to other similar systems or to other electronic devices, for example, via a communication bus or through dedicated input and / or output ports.

[0184] System 700 may include at least one processor 710 configured to execute loaded instructions to implement various aspects described in this disclosure, for example. Processor 710 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 700 may include at least one memory 720 (e.g., a volatile memory device and / or a non-volatile memory device). System 700 may also include a storage device 740, which may include non-volatile and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drives, and / or optical disk drives. Storage device 740 may include, by way of non-limiting example, an internal storage device, an external storage device, and / or a network-accessible storage device.

[0185] System 700 may include an encoder / decoder module 730 configured to process data to provide, for example, encoded / decoded video picture data, which may include its own processor and memory. Encoder / decoder module 730 may correspond to a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 730 may be implemented as a separate element of system 700 or may be incorporated within a processor as a combination of hardware and software, as is known to those skilled in the art.

[0186] Program code loaded into processor 710 or encoder / decoder module 730 to perform various aspects described in this disclosure may be stored in storage device 740 and subsequently loaded into memory 720 for execution by processor 710. According to various embodiments, one or more of processor 710, memory 720, storage device 740, and encoder / decoder module 730 may hold one or more of various items during execution of processes described in this disclosure. Such held items may include, but are not limited to, video picture data, information data used to encode / decode video picture data, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and computational logic.

[0187] In some embodiments, memory within the processor 710 and / or encoder / decoder module 730 may be used to store instructions and to provide working memory for processing that may occur during encoding or decoding.

[0188] However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 710 or the encoder / decoder module 730) may be used for one or more of these functions. The external memory may be memory 720 and / or storage device 740, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory may be used to store the television's operating system. In at least one embodiment, high-speed external dynamic non-volatile memory such as RAM may be used as working memory for video encoding and decoding operations, such as in MPEG-2 part 2 (also known as MPEG-2 Video, also known as ISO / IEC 13818-2 and ITU-T Recommendation H.262), AVC, HEVC, EVC, VVC, AV1, etc.

[0189] Inputs to the elements of system 700 may be provided through various input devices, as indicated by block 790. Such input devices include, but are not limited to, (i) an RF section that may receive, for example, an RF signal transmitted wirelessly by a broadcaster, (ii) a composite input, (iii) a USB input, (iv) an HDMI® input, and (v) if the present disclosure is used in the automotive field, a bus such as a Controller Area Network (CAN), Controller Area Network Flexible Data-Rate (CAN FD), FlexRay (ISO 17458), or Ethernet® (ISO / IEC 802-3) bus.

[0190] In various embodiments, the input devices of block 790 may have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements necessary for (i) selecting a desired frequency (e.g., also referred to as selecting a signal or bandlimiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower band of frequencies to select a signal frequency band, which in certain embodiments may be referred to as a channel (e.g.,), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments may include one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs various of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband.

[0191] In one set-top box embodiment, the RF section and its associated input processing elements may receive an RF signal transmitted over a wired (e.g., cable) medium. The RF section may then perform frequency selection by filtering to a desired frequency band, down-conversion, and saturation filtering.

[0192] Various embodiments rearrange the order of the above-described (and other) elements, delete some of these elements, and / or add other elements that perform similar or different functions.

[0193] Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters, etc. In various embodiments, the RF section may include an antenna.

[0194] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 700 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 710, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, within a separate interface IC or within processor 710, as desired. The demodulated, error corrected, and demultiplexed stream may be provided to various processing elements, including processor 710 and encoder / decoder module 730, which operate in conjunction with memory and storage devices to process the data stream as desired for presentation on an output device.

[0195] The various elements of system 700 may be provided within a single unitary housing, where the various elements may be interconnected and data may be transmitted therebetween using suitable connection arrangements 790, such as internal buses, wiring, and printed circuit boards known in the art, including an I2C bus.

[0196] System 700 may include a communication interface 750 that enables communication with other devices over a communication channel 751. Communication interface 750 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 751. Communication interface 750 may include, but is not limited to, a modem or a network card, and communication channel 751 may be implemented within a wired and / or wireless medium, for example.

[0197] Data may, in various embodiments, be streamed to system 700 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments may be received via communication channel 751 and communication interface 750 adapted for Wi-Fi communication. Communication channel 751 in these embodiments may typically be connected to an access point or router that provides access to an external network, including interfaces to enable streaming applications and other over-the-top communications.

[0198] Other embodiments may provide streaming data to system 700 using a set-top box that delivers data via an HDMI connection in input block 790 .

[0199] Yet another embodiment may provide streaming data to the system 700 using the RF connection of the input block 790 .

[0200] Streaming data may be used as a method for signaling information used by system 700. The signaling information may include information such as the bitstream B, and / or the number of pixels in a video picture and / or any encoding / decoding setup parameters.

[0201] It should be understood that signaling can be achieved in various ways, for example, one or more syntax elements, flags, etc. may be used to signal information to a corresponding decoder in various embodiments.

[0202] System 700 may provide output signals to various output devices, including a display 761, speakers 771, and other peripherals 781. Other peripherals 781 may, in various example embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 700.

[0203] In various embodiments, control signals may be exchanged between the system 700 and the display 761, speaker 771, or other peripheral device 781 using signaling such as, for example, Audio / Video Link (AV.Link), Consumer Electronics Control (CEC), or other communication protocols that enable inter-device control with or without user intervention.

[0204] Output devices may be communicatively coupled to system 700 via dedicated connections through respective interfaces 760, 770 and 780.

[0205] Alternatively, output devices may be connected to system 700 using communication channel 751 via communication interface 750. Display 761 and speakers 771 may be integrated into a single unit with other components of system 700 in an electronic device such as a television.

[0206] In various embodiments, the display interface 760 may include a display driver, such as, for example, a timing controller (T Con) chip.

[0207] The display 761 and speakers 771 may alternatively be separate from one or more of the other components, for example, if the RF portion of the input block 790 is part of a separate set-top box. In various embodiments in which the display 761 and speakers 771 may be external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.

[0208] 1-22, various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be varied or combined.

[0209] Some examples are described with reference to block diagrams and / or operational flowcharts. Each block represents a portion of code, a circuit element, or a module that includes one or more executable instructions for implementing a particular logical function. It should also be noted that in other implementations, the functions noted in the blocks may occur out of the order shown. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved.

[0210] Implementations and aspects described herein may be embodied in, for example, a method or process, an apparatus, a computer program, a data stream, a bit stream, or a signal. Even if only a single embodiment is discussed (e.g., only as a method), implementations of the discussed features may also be embodied in other forms (e.g., an apparatus or a computer program).

[0211] The methods may be implemented, for example, in a processor, which refers generally to processing devices, including, for example, computers, microprocessors, integrated circuits, or programmable logic devices. A processor also includes a communications device.

[0212] Furthermore, the method may be implemented by instructions being executed by a processor, and such instructions (and / or data values ​​produced by the implementation) may be stored on a computer-readable storage medium. The computer-readable storage medium may take the form of a computer-readable program product having computer-readable program code embodied in one or more computer-readable media and executable by a computer. As used herein, a computer-readable storage medium may be considered a non-transitory storage medium, given its inherent ability to store information as well as retrieve information therefrom. The computer-readable storage medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. The following provides more specific examples of computer-readable storage media to which the present embodiment may be applied; however, as will be readily understood by those skilled in the art, these examples are illustrative and not comprehensive. For example, a portable computer diskette, a hard disk, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0213] The instructions may form an application program tangibly embodied on a processor-readable medium.

[0214] The instructions may be, for example, hardware, firmware, software, or a combination. The instructions may be found, for example, in an operating system, a separate application, or a combination of the two. A processor may thus be characterized, for example, as both a device configured to execute a process and a device that includes a processor-readable medium (e.g., a storage device) having instructions for executing the process. Furthermore, the processor-readable medium may store data values ​​produced by the execution in addition to or in place of instructions.

[0215] The apparatus may be implemented, for example, in appropriate hardware, software, and firmware. Examples of such apparatus include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, Internet appliances, head-mounted display devices (HMDs, see-through glasses), projectors (beamers), "caves" (systems including multiple displays), servers, video encoders, video decoders, post-processors that process output from video decoders, pre-processors that provide input to video encoders, web servers, set-top boxes, and other devices that process video pictures or other communications devices. Obviously, the apparatus may be mobile and installed in a mobile vehicle.

[0216] The computer software may be implemented by the processor 710 or by hardware, or by a combination of hardware and software. By way of non-limiting example, embodiments may also be implemented by one or more integrated circuits. The memory 720 may be of any type suitable for the technology environment and may be implemented using any suitable storage technology, by way of non-limiting example, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, removable memory, etc. The processor 710 may be of any type suitable for the technology environment and may include, by way of non-limiting example, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0217] Those skilled in the art will appreciate that implementations may generate a variety of signals formatted to carry information that may be, for example, stored or transmitted. Information may include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

[0218] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms (a, an, and the) are intended to include the plural forms unless the context clearly dictates otherwise. Furthermore, it should be understood that the terms "comprise" and "have," when used herein, may specify the presence of stated features, integers, steps, operations, elements, and / or components, for example, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, when an element is described as "responsive to," "connected to," or "associated with" another element, it may be directly responsive to, connected to, or associated with the other element, or there may be intervening elements. In contrast, when an element is described as "directly responsive to," "directly connected to," or "directly associated with," there are no intervening elements.

[0219] The use of the letters and terms " / ," "and / or," and "at least one of" may be intended to encompass the selection of only the first listed alternative (A), the selection of only the second listed alternative (B), or the selection of both alternatives (A and B), such as in "A / B," "A and / or B," and "at least one of A and B." As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), the selection of only the second listed alternative (B), the selection of only the third listed alternative (C), the selection of only the first and second listed alternatives (A and B), the selection of only the first and third listed alternatives (A and C), the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This may be expanded as would be apparent to one of ordinary skill in the art for multiple listed items.

[0220] Various numerical values ​​may be used in this disclosure, and specific values ​​are for illustrative purposes only and the described aspects are not limited to such specific values.

[0221] Although terms such as "first," "second," and the like may be used herein to describe various elements, it should be understood that these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a "second element," and similarly, a second element may be referred to as a "first element" without departing from the teachings of the present disclosure. No ordering between a first element and a second element is implied.

[0222] References to "one embodiment" or "embodiment" or "one implementation" or "implementation," as well as other variations thereof, are often used to convey that a particular feature, structure, characteristic, etc. (described in connection with an embodiment / implementation) is included in at least one embodiment / implementation. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation," and any other variations thereof, appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment.

[0223] Similarly, references herein to "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" and other variations thereof are often used to convey that a particular feature, structure, characteristic, etc. (described in connection with an embodiment / implementation) may be included in at least one embodiment / example / implementation. Thus, appearances of the phrase "according to an embodiment / example / implementation" or "in an embodiment / example / implementation" appearing in various places throughout this disclosure are not necessarily all referring to the same embodiment / example / implementation, nor are they necessarily separate or alternative embodiments / examples / implementations that are necessarily mutually exclusive of other embodiments / examples / implementations.

[0224] Reference numerals appearing in the claims are by way of illustration only and shall have no effect limiting the scope of the claims. Even if not explicitly stated, the embodiments / examples and variations may be used in any combination or sub-combination.

[0225] Where a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.

[0226] Although some of the figures include arrows on communication paths to indicate the primary direction of communication, it should be understood that communication may occur in the opposite direction to that shown.

[0227] Various implementations include decoding. As used herein, "decoding" may include all or part of the processes performed on received video pictures (which may include a received bitstream encoding one or more video pictures), for example, to generate a final output suitable for display or further processing in the reconstructed video domain. In various embodiments, such processes include one or more of the processes typically performed by a decoder. In various embodiments, such processes may also or alternatively include processes performed by decoders of various implementations described in this disclosure, for example.

[0228] As a further example, in one embodiment, "decoding" may refer to inverse quantization only, in one embodiment, "decoding" may refer to entropy decoding, in another embodiment, "decoding" may refer to differential decoding only, and in another embodiment, "decoding" may refer to a combination of inverse quantization, entropy decoding, and differential decoding. Whether the phrase "decoding process" may be intended to refer specifically to a subset of operations or generally to a lower decoding process will be clear based on the context of the specific description and is believed to be properly understood by one of ordinary skill in the art.

[0229] Various implementations include encoding. Similar to the above discussion of "decoding," "encoding," as used in this disclosure, may include all or part of the processes performed, for example, on input video pictures, to generate an output bitstream. In various embodiments, such processes may include one or more of the processes typically performed by an encoder. In various embodiments, such processes may also, or alternatively, include processes performed by the encoders of the various implementations described herein.

[0230] As a further example, in one embodiment, "encoding" may refer to quantization only, in one embodiment, "encoding" may refer to entropy encoding only, in another embodiment, "encoding" may refer to differential encoding only, and in another embodiment, "encoding" may refer to a combination of quantization, differential encoding, and entropy encoding. Whether the term "encoding process" may be intended to refer specifically to a subset of operations or generally to lower encoding processes will be clear based on the context of the specific description and is believed to be appropriately understood by one of ordinary skill in the art.

[0231] Additionally, this disclosure may refer to "obtaining" various information. Obtaining information may include, for example, one or more of estimating information, calculating information, predicting information, or reading information from memory, processing information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0232] Additionally, this application may refer to "receiving" various information. Receiving information may include, for example, one or more of: accessing the information or receiving the information from a communications network.

[0233] Also, as used herein, the term "signaling" refers, among other things, to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals certain information, such as coding parameters or encoded video picture data. In this way, in embodiments, the same parameters may be used on both the encoder and decoder sides. Thus, for example, an encoder may transmit certain parameters to a decoder (explicit signaling), thereby allowing the decoder to use the same certain parameters. Conversely, if the decoder already has certain parameters and other parameters, signaling may be used without transmission (implicit signaling) simply to allow the decoder to know and select certain parameters. By avoiding the transmission of any actual functions, bit savings are achieved in various embodiments. Of course, signaling may be achieved in various ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. Although the above relates to the verb form of the word "signaling," the word "signaling" is sometimes used herein as a noun (in which case it is also called a "signal").

[0234] A number of implementations have been described. Nevertheless, it will be understood that various modifications may be made. For example, elements of different implementations may be combined, substituted, modified, or eliminated to produce other implementations. Moreover, as will be understood by those skilled in the art, other structures and processes may be substituted for those disclosed, with the resulting implementations performing at least substantially the same functions, in at least substantially the same ways, to achieve at least substantially the same results as the disclosed implementations. Accordingly, these and other implementations are contemplated by this application.

[0235] This disclosure claims priority to and benefits from European Patent Application No. 22306424.7, filed September 27, 2022, the entire contents of which are incorporated herein by reference.

Claims

1. 1. A method for decoding blocks of samples of a video picture, comprising: determining an intra prediction mode for decoding said block of samples according to an analysis of gradients of samples located in at least one template area defined around said block of samples, The step of determining an intra-prediction mode includes: calculating a histogram of gradients by filtering samples of the at least one template area, each entry of the histogram of gradients corresponding to an angular intra-prediction mode, the filtering using a filtering window centered on a sample position of a mid-line of the at least one template area; selecting at most two angular intra prediction modes by comparing amplitudes of angular intra predictions of the histograms of gradients; determining the intra prediction mode from the at most two selected angular intra prediction modes; A method wherein the integer number of sample positions of intermediate lines of the at least one template area about which the filtering window is centered is less than the total integer number of sample positions of intermediate lines of the at least one template area.

2. 1. A method for encoding blocks of samples of a video picture, comprising the steps of: determining an intra prediction mode for decoding said block of samples according to an analysis of gradients of samples located in at least one template area defined around said block of samples, The step of determining an intra-prediction mode includes: calculating a histogram of gradients by filtering samples of the at least one template area, each entry of the histogram of gradients corresponding to an angular intra-prediction mode, the filtering using a filtering window centered on a sample position of a mid-line of the at least one template area; selecting at most two angular intra prediction modes by comparing amplitudes of angular intra predictions of the histograms of gradients; determining the intra prediction mode from the at most two selected angular intra prediction modes; A method wherein the integer number of sample positions of intermediate lines of the at least one template area about which the filtering window is centered is less than the total integer number of sample positions of intermediate lines of the at least one template area.

3. the histogram of gradients for chroma components of the video picture is calculated by filtering only chroma samples of the at least one template area.

3. The method according to claim 1 or 2.

4. at least one of the at least one template area includes samples along a bottom edge of an adjacent upper virtual pipeline data unit and samples along a right edge of an adjacent left virtual pipeline data unit; 3. The method according to claim 1 or 2.

5. at least one of the at least one template area includes all luma samples on a boundary of the virtual pipeline data unit in a vicinity; The method of claim 4.

6. the template area includes a subset of luma samples on a boundary of the virtual pipeline data unit adjacent to the template area; The method of claim 4.

7. one histogram of gradients is calculated for each chroma component of the video picture; 3. The method according to claim 1 or 2.

8. each histogram of gradients for a chroma component of the video picture is calculated from luma samples and chroma samples of a template area; The method of claim 7.

9. the chroma samples of the template area are multiplied by a compensation factor that depends on the chroma subsampling defined by the video picture format; The method of claim 8.

10. a sample position of the intermediate line of at least one of the at least one template area in which the filtering window is centered is determined to place the filtering window at one of a fourth number of sample positions of the intermediate line of the template area.

3. The method according to claim 1 or 2.

11. a sample position of the intermediate line of at least one of the at least one template area in which the filtering window is centered is determined so that there is no overlap between the filtering windows; 3. The method according to claim 1 or 2.

12. the filtering windows have different sizes; 3. The method according to claim 1 or 2.

13. the size of the filtering window depends on the size of the block of samples and the availability of samples in the template area; 3. The method according to claim 1 or 2.

14. Apparatus having means for carrying out the method according to claim 1 or 2.

15. A non-transitory storage medium carrying program code instructions for carrying out the method of claim 1 or 2.

16. a processor; a memory that stores instructions executable by the processor; The processor is configured to perform the method according to claim 1 or 2. Electronic devices.

17. A computer program comprising instructions that, when executed by one or more processors, cause said one or more processors to perform the method of claim 1 or 2.

Citation Information

Patent Citations

  • Method for image processing and apparatus for implementing the same

    US20220070451A1