Video image data encoding / decoding
Patent Information
- Application Number
- BR122026018566
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Publication Date
- 2026-09-01
Smart Images

Figure 00000000_0000_ABST
Description
1 / 48 “VIDEO IMAGE DATA ENCODING / DECODING” Divided from BR1120250054570, deposited on April 24, 2023. CROSS-REFERENCE TO RELATED REQUEST
[001] This disclosure claims priority and benefits of European Patent Application No. 22306424.7, filed on September 27, 2022, the full content of which is incorporated herein by reference. FIELD
[002] This disclosure generally relates to the encoding and decoding of video images. Particularly, but not exclusively, the technical field of this disclosure relates to intraprediction blocks of a video image. BACKGROUND
[003] The purpose of this section is to introduce the reader to various aspects of the art, which may be related to several aspects of at least one embodiment of the present disclosure that is described and / or claimed below. It is believed that this discussion will be useful in providing the reader with background information to facilitate a better understanding of the various aspects of the present disclosure. Consequently, it should be understood that these statements should be read in this light, and not as admissions of prior art.
[004] A pixel corresponds to the smallest display unit on a screen, which can be composed of one or more light sources (1 for a monochrome screen or 3 or more for color screens).
[005] A video image, also called a frame or picture frame, comprises at least one component (also called an image component or channel) determined by a specific image / video format that specifies all information relating to pixel values and all information that can be used by a display unit and / or any other device to display and / or decode video image data relating to said image. Petition 870260074236, dated 07 / 24 / 2026, page 10 / 80 2 / 48 video.
[006] A video image comprises at least one component, usually expressed in the form of a sample matrix.
[007] A monochrome video image comprises a single component, and a color video image may comprise three components.
[008] For example, a color video image may comprise one luma (or luminance) component and two chroma components when the image / video format is the well-known format (Y, Cb, Cr) or it may comprise three color components (one for red, one for green and one for blue) when the image / video format is the well-known format (R, G, B).
[009] Each component of a video image may comprise a number of samples relative to a number of pixels on a screen on which the video image is to be displayed. In variants, the number of samples comprised in one component may be a multiple (or fraction) of the number of samples comprised in another component of the same video image.
[010] For example, in the case of a video format comprising a luma component and two chroma components such as the (Y, Cb, Cr) format, depending on the color format considered, the chroma component may contain half the number of samples in width and / or height, in relation to the luma component.
[011] A sample is the smallest unit of visual information of a component that makes up a video image. A sample value can be, for example, a luma or chroma value or a color value of a format (R, G, B).
[012] A pixel value is the value of a pixel on a screen. A pixel value can be represented by a sample for monochrome video images and by multiple colocalized samples for color video images. Colocalized samples associated with a pixel mean samples corresponding to the location of a pixel on the screen. Petition 870260074236, dated 07 / 24 / 2026, page 11 / 80 3 / 48
[013] It is common to consider a video image as a set of pixel values, each pixel being represented by at least one sample.
[014] A video image block is a set of samples of a component of the video image. A block of at least one luma sample, in short a luma block, or a block of at least one chroma sample, in short a chroma block, may be considered when the image / video format is the well-known format (Y, Cb, Cr), or a block of at least one color sample when the image / video format is the well-known format (R, G, B).
[015] At least one modality is not limited to a particular image / video format.
[016] In next-generation video compression systems, such as HEVC (ISO / IEC 23008-2 High Efficiency Video Coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265-202108-P / en) or VVC (ISO / IEC 23090-3 Versatile Video Coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / TREC-H.266-202008-I / en), low- and high-level image partitioning is provided to divide a video image into image areas, called Coding Tree Units (CTUs), whose size can typically be between 16x16 and 64x64 pixels for HEVC and 32x32, 64x64 or 128x128 pixels for VVC.
[017] The CTU division of a video image forms a grid of fixed size CTUs, that is, a CTU grid, whose upper and left limits coincide spatially with the upper and left edges of the video image. The CTU grid represents a spatial partition of the video image.
[018] In VVC and HEVC, the CTU size (CTU width and CTU height) of all CTUs in a CTU grid is equal to the same standard CTU size (standard CTU width CTU DW and standard CTU height CTU DH). For example, the standard CTU size (standard CTU height, standard CTU width) might be equal to 128 (CTU DW=CTU DH=128). A standard CTU size (height, width) is encoded in the bitstream, for example, at a sequence level in the Set of Petition 870260074236, dated 07 / 24 / 2026, page 12 / 80 4 / 48 Sequence Parameters (SPS).
[019] The spatial position of a CTU in a CTU grid is determined from a CTU address ctuAddr defining a spatial position of the upper left corner of a CTU from an origin. As illustrated in Figure 1, the CTU address can define the spatial position of the upper left corner of a higher-level spatial structure S containing the CTU.
[020] A coding tree is associated with each CTU to determine a CTU tree split.
[021] As illustrated in Figure 1, in HEVC, the coding tree is a quaternary tree division of a CTU, where each leaf is called a Coding Unit (CU). The spatial position of a CU in the video image is defined by a CU index cu Idx indicating a spatial position from the upper left corner of the CTU. A CU is spatially partitioned into one or more Prediction Units (PU). The spatial position of a PU in the VP video image is defined by a PU index puldx defining a spatial position from the upper left corner of the CTU, and the spatial position of an element of a partitioned PU is defined by a PU partition index pu Partldx defining a spatial position from the upper left corner of a PU. Each PU receives some inter- or intraprediction data.
[022] The intra or inter coding mode is assigned at the CU level. This means that the same intra / inter coding mode is assigned to each PU of a CU, although the prediction parameters vary from PU to PU.
[023] A CU can also be spatially partitioned into one or more Transform Units (TU), according to a quaternary tree called a transform tree. Transform Units are the leaves of the transform tree. The spatial position of a TU in the video image is defined by a TU index tu Idx defining a spatial position of the upper left corner of a CU. Each TU receives some transform parameters. The type of transform is Petition 870260074236, dated 07 / 24 / 2026, page 13 / 80 5 / 48 assigned at the TU level, and the separate 2D transform is performed at the TU level during the encoding or decoding of an image block.
[024] The existing PU Partition types in HEVC are illustrated in Figure 2. They include square partitions (2Nx2N and NxN), which are the only ones used in Ima and Inter prediction CUs, non-square symmetric partitions (2NxN, Nx2N, used only in interprediction CUs) and asymmetric partitions (used only in interprediction CUs). For example, the 2NxnU PU type represents an asymmetric horizontal partitioning of the PU, where the smaller partition is on top of the PU. According to another example, the 2NxnL PU type represents an asymmetric horizontal partitioning of the PU, where the smaller partition is on top of the PU.
[025] As illustrated in Figure 3, in VVC, the coding tree starts from a root node, i.e., the CTU. Then, a quaternary tree split (or quaternary tree) divides the root node into 4 nodes corresponding to 4 sub-blocks of equal size (solid lines). Then, the leaves of the quaternary tree (or quaternary tree) can be further partitioned by a so-called multi-type tree, which involves a binary or ternary split according to one of the 4 split modes illustrated in Figure 4. These split types are the vertical and horizontal binary split modes, noted SBTV and SBTH, and the vertical and horizontal ternary split modes SPTTV and STTH.
[026] The leaves of a CTU coding tree are CU in the case of a joint coding tree shared by luma and chroma components.
[027] Unlike HEVC, in VVC, in most cases, CU, PU and TU have equal size, which means that the encoding units are generally not partitioned into PU or TU, except in some specific encoding modes.
[028] Figures 5 and 6 provide an overview of the video encoding / decoding methods used in current video standard compression systems, such as HEVC or VVC, for example.
[029] Figure 5 shows a schematic block diagram of the steps of a Petition 870260074236, dated 07 / 24 / 2026, page 14 / 80 6 / 48 method 100 of encoding a VP video image according to the previous technique.
[030] In step 110, a VP video image is partitioned into sample blocks and the partitioning information data is signaled in a bitstream. Each block comprises samples of a component of the VP video image. The blocks, therefore, comprise samples of each component that defines the VP video image.
[031] For example, in HEVC, an image is divided into Coding Tree Units (CTUs). Each CTU can be further subdivided using a quaternary tree division, where each leaf of the quaternary tree is denoted as a Coding Unit (CU). Partitioning information data can then comprise data describing the CTU and the quaternary tree subdivision of each CTU.
[032] Each short block of samples can then be a CU (if the CU comprises a single PU) or a PU of a CU.
[033] Each current block is encoded along an encoding loop also called “in loop” using an inter or intraprediction mode.
[034] Intraprediction (step 120) used intraprediction data. Intraprediction consists of predicting a current block by means of an intraprediction block based on already encoded, decoded and reconstructed samples located in an L-shaped model area defined around the current block, typically at the top left and top right of the current block. Intraprediction is performed in the spatial domain.
[035] In interprediction mode, motion estimation (step 130) and motion compensation (135) are performed. Motion estimation searches, in one or more reference images used to predictively encode the current video image, for a reference block that is a good predictor of the current block. In one-way motion estimation / compensation, a reference block Petition 870260074236, dated 07 / 24 / 2026, page 15 / 80 7 / 48 candidate belongs to a single reference image from a list of reference images denoted L0 or L1, and in bidirectional motion estimation / compensation, the candidate reference block is derived from a reference block from the L0 reference image list and a reference block from the L1 reference image list.
[036] For example, a good predictor of the current block is a candidate reference block that is similar to the current block. It may also match a reference block that provides a good trade-off between its similarity to the current block and the cost of the rate of movement information required to indicate its use for the temporal prediction of the current block.
[037] The output of motion estimation step 130 is interprediction data comprising motion information associated with the current block and other information used to obtain the same prediction block on the encoding / decoding side. Typically, the motion information comprises a motion vector and a reference image index for one-way estimation / compensation and two motion vectors and two reference image indices for two-way estimation / compensation. Motion compensation (step 135) then obtains a prediction block using the motion vector(s) and reference image index(es) determined by motion estimation step 130. Basically, the reference block belonging to a selected reference image and pointed to by a motion vector can be used as the prediction block of the current block.Furthermore, since motion vectors are expressed as fractions of whole pixel positions (what is known as sub-pel precision motion vector representation), motion compensation usually involves spatial interpolation of some reconstructed samples from the reference image to calculate the prediction block.
[038] Predictive information data are signaled in the bitstream. As Petition 870260074236, dated 07 / 24 / 2026, page 16 / 80 8 / 48 Prediction information may include prediction mode (intra or inter or jump), intra / inter prediction data, and any other information used to obtain the same prediction block on the decoding side.
[039] Method 100 selects a prediction mode (the intra- or interprediction mode) by optimizing a rate-distortion trade-off by taking into account the encoding of a calculated residual prediction block, for example, by subtracting a candidate prediction block from the current block, and the signaling of prediction information data needed to determine said candidate prediction block on the decoding side.
[040] Typically, the best prediction mode is given as the prediction mode of a best encoding mode p* for a given current block: P * = Argmm{7?DC0St(p)} pep where P is the set of all candidate encoding modes for the current block, p represents a candidate encoding mode in that set, RDcost(p) is a rate distortion cost of the candidate encoding mode p, typically expressed as: RDcostíp = D(p+ + M(p)
[041] D(p) is the distortion between the current block and a reconstructed block obtained after encoding / decoding the current block with the candidate encoding mode p, R(p) is a rate cost associated with encoding the current block with encoding mode p, and λ is the Lagrange parameter that represents the rate constraint for encoding the current block and is normally computed from a quantization parameter used to encode the current block.
[042] The current block is usually encoded from a residual PR prediction block. More precisely, a residual PR prediction block is calculated, for example, by subtracting the best prediction block from the current block. The residual PR prediction block is then transformed (step 140) using, for example, a DCT (discrete cosine transform) or DST (discrete thoracic transform) type transform. Petition 870260074236, dated 07 / 24 / 2026, page 17 / 80 9 / 48 discrete sine), or any other appropriate transform, and the resulting transformed coefficient block is quantized (step 150).
[043] In a variant, method 100 can also skip the transform step 140 and apply quantization (step 150) directly to the PR prediction residual block, according to the so-called jump-transform coding mode.
[044] The quantized transform coefficient block (or quantized prediction residual block) is entropy encoded in the bitstream (step 160).
[045] Next, the quantized transform coefficient block (or the quantized residual block) is dequantized (step 170) and inversely transformed (180) (or not) as part of the encoding loop, leading to a decoded prediction residual block. The decoded prediction residual block and the prediction block are then combined, usually summed, which gives the reconstructed block.
[046] Other information data can also be entropy encoded in step 160 to encode a current block of the VP video image.
[047] Loop filters (step 190) can be applied to a reconstructed image (comprising reconstructed blocks) to reduce compression artifacts. Loop filters can be applied after all image blocks have been reconstructed. For example, they consist of unlock filter, Adaptive Shift Sample (SAO) or Adaptive Loop Filter (ALF).
[048] The reconstructed blocks or the filtered reconstructed blocks form a reference image that can be stored in a decoded image buffer (DPB) so that it can be used as a reference image for encoding a next current block of the VP video image, or a next video image to be encoded.
[049] Figure 6 shows a schematic block diagram of steps in a method 200 for decoding a VP video image according to the previous technique. Petition 870260074236, dated 07 / 24 / 2026, page 18 / 80 10 / 48
[050] In step 210, partitioning information data, prediction information data, and quantized transform coefficient block (or quantized residual block) are obtained by entropy decoding of a bitstream of encoded video image data. For example, this bitstream was generated according to method 100.
[051] Other information data can also be decoded by entropy for decoding the bit stream of a current block of the VP video image.
[052] In step 220, a reconstructed image is divided into current blocks based on partitioning information. Each current block is entropy-decoded from the bitstream along a decoding loop also called an “in loop”. Each decoded current block is a quantized transform coefficient block or quantized prediction residual block.
[053] In step 230, the current block is dequantized and possibly inversely transformed (step 240), to obtain a decoded prediction residual block.
[054] On the other hand, prediction information data is used to predict the current block. A prediction block is obtained through its intraprediction (step 250) or its motion-compensated temporal prediction (step 260). The prediction process performed on the decoding side is identical to that on the encoding side.
[055] Next, the decoded prediction residual block and the prediction block are then combined, usually summed, which provides a reconstructed block.
[056] In step 270, loop filters can be applied to a reconstructed image (comprising reconstructed blocks) and the reconstructed blocks or the filtered reconstructed blocks form a reference image that can be stored in a decoded image buffer (DPB) as discussed above (Figure 5).
[057] In VVC, movement information is stored in 4x4 blocks Petition 870260074236, dated 07 / 24 / 2026, page 19 / 80 11 / 48 in each video image. This means that once a reference image is stored in the decoded image buffer (DPB, Figure 5 or 6), motion vectors and reference image indices used for temporal prediction of video image blocks are stored on a 4x4 block basis. They can serve as temporal prediction of motion information for encoding / decoding of a subsequent interprediction video image.
[058] To reduce redundancy between components, VVC defines the so-called Linear Model-based intraprediction mode (LM modes).
[059] One of them is the so-called Crossed Component Linear Model (CCLM) which derives a linear model-based predictor, in short, LM predictor, comprising predicted chroma samples predc(i, j) based on co-located reconstructed luma samples recL(i,;) from the same CU using a linear model: predc(i, j) = α · recL Ό, j) + β where α and β are linear parameters of the linear model that are derived from reference samples, i.e., reconstructed chroma and luma samples.
[060] The reconstructed luma samples recL(i,j) are reduced after filtering to match the CU chroma size.
[061] In VVC, three CCLM modes, denoted CCLM_LT, CCLM_T, and CCLM_L, are specified. These three CCLM modes differ in the location of the reference samples that are used for linear parameter derivation. Upper bound reference samples are involved in the CCLM_T mode, and left bound reference samples are involved in the CCLM_L mode. In the CCLM_LT mode, both upper and left bound reference samples are used.
[062] Overall, the CCLM mode prediction process consists of three steps: 1) Subsampling the luma block and its reconstructed neighboring luma samples recL(i,jj) to match the size of the corresponding chroma block, 2) derivation of linear parameters based on the reconstructed neighboring luma samples, 3) Application of equation (10) to generate the intraprediction samples of Petition 870260074236, dated 07 / 24 / 2026, page 20 / 80 12 / 48 chroma (predicted chroma block).
[063] Another LM mode is the so-called Multimodel Linear Model (MMLM, [K. Zhan et al, “Enhanced Cross-component Linear Model Intra Prediction”, JVETD0110, San Diego, October 2016]). MMLM is an extension of CCLM because more than one linear model is derived from reconstructed, colocalized luma samples recL(ij). In MMLM, neighboring reconstructed luma samples and neighboring chroma samples are classified into several groups, each group being used as a training set to derive linear parameters of a linear model (i.e., particular α and β are derived for a particular group). In addition, samples from a current luma block are also classified based on the same rule for classifying neighboring luma samples. Neighboring samples are, for example, classified into M groups. The MMLM method with M=2 and M=3 is designed as two attached LM modes for chroma called MMLM2 and MMLM3, in addition to the original LM mode (CCLM).The encoder selects the ideal LM mode in a Rate / Distortion Optimization process and signals the best LM mode. For example, when M equals 2, a threshold is calculated as the average value of the reconstructed neighboring luma samples. A reconstructed neighboring sample rec'L[x,y] less than or equal to the threshold is classified in group 1; while a reconstructed neighboring sample rec'L[x,y] greater than the threshold is classified in group 2.
[064] The two predictors of LM are then derived as: (Predc[x,y] = «ix rec 'i[x,y] + βι if rec 'i[x,y] < Limit tPredc[x,y] = a2x rec 'i[x,y] + β2if rec 'i[x,y] > Limit
[065] MMLM is included in the Enhanced Compression Model (ECM) (M. Coban et al, “Algorithm description of Enhanced Compression Model 4 (ECM 4)”, JVETY2025, Online, July 2021) exploring improved compression performance beyond VVC, where reconstructed neighboring samples are classified into two classes using a threshold that is the average of the reconstructed neighboring luma samples. The linear model of each class is derived using the Least Squares method. Petition 870260074236, dated 07 / 24 / 2026, page 21 / 80 13 / 48 (LMS).
[066] The VVC further defines a Decoder-side Intra Mode Derivation (DIMD) mode for luma and chroma samples. The DIMD mode is not a linear model-based mode, in short, a non-LM mode, i.e., an intraprediction mode that does not refer to a linear model such as, for example, planar mode or direct mode (DM).
[067] DM is an intraprediction mode that uses the sample mean of the reference samples to the left and above the block for prediction generation.
[068] The planar mode is a weighted average of 4 reference sample values (captured as orthogonal projections of the sample to predict in the reconstructed upper and left areas). The horizontal and vertical modes use, respectively, a copy of the reconstructed samples on the left and above without interpolation to predict rows and columns of samples, respectively.
[069] For luma sample prediction, the use of luma DIMD mode is signaled in a bitstream by a single flag and the intra predictor is not explicitly signaled in the bitstream.
[070] Figure 7 schematically illustrates a block diagram of a 300 method for derivating the luma DIMD mode according to the prior art.
[071] Basically, a luma DIMD (predictor) mode is derived using a gradient analysis of neighboring reconstructed luma samples, i.e., a gradient analysis of luma samples located in an L-shaped model area defined around a current block of samples. If the luma DIMD mode is not enabled, the intraprediction mode can be analyzed from the bitstream as in the classic intraprediction mode. The luma DIMD mode is implicit. Thus, the luma DIMD mode is derived during the reconstruction process identically on the encoder and decoder sides.
[072] In step 310, intra-above and intra-left samples available around a luma block to be predicted are determined and a model area in the shape Petition 870260074236, dated 07 / 24 / 2026, page 22 / 80 14 / 48 of L is then determined.
[073] For example, as illustrated in Figure 8, an L-shaped model area T of 3 samples wide (in width or height) composed of luma samples reconstructed to the left, above and above the left of the reconstructed area R adjacent to a current block B (current CU), is defined.
[074] In step 320, a Histogram of Gradients (HoG) is constructed as follows.
[075] First, samples of the L-shaped model's T-area are filtered, said filtering using W-filter windows centered on midline sample positions of the L-shaped model's T-area. An amplitude and angle of a luminance direction (orientation) are assigned to each midline sample of the L-shaped model's T-area.
[076] For example, when edge detection filters (3x3 horizontal and vertical Sobel filters) are used to filter samples from the T-area of an L-shaped model, the amplitude and angle of a luminance direction are given by: angle = arctan (Ghor / Gver) amplitude = | Ghor| + | Ghor| with Ghor and Gver being the intensity of pure horizontal and vertical directions as calculated by Sobel filters. Each pair of amplitude and angle of a luminance direction corresponds to an angular intraprediction mode.
[077] Next, the HoG is calculated where each input (angle) corresponds to an angular intraprediction mode and accumulated amplitudes are stored.
[078] In step 330, at most two angular intraprediction modes M1 and M2 are selected from the HoG.
[079] For example, as illustrated in Figure 9, the two most represented angular intraprediction modes M1 and M2 have the largest histogram amplitude values.
[080] It may happen that 0, 1 or 2 modes of angular intraprediction are Petition 870260074236, dated 07 / 24 / 2026, page 23 / 80 15 / 48 selected from the HoG. For example, none of the angular intraprediction modes are selected if none of the cumulative HoG amplitudes are greater than a threshold.
[081] When none of the angular intraprediction modes are selected, then in step 340, the planar mode is the intraprediction mode used as the intrapredictor of the luma part of the current block to be predicted.
[082] When a single angular intraprediction mode is selected, then in step 350, this angular intraprediction mode is the intraprediction mode used as the intrapredictor of the luma portion of the current block to be predicted.
[083] When two angular intraprediction modes are selected (M1 and M2), then in step 360, the luma DIMD mode is determined by mixing (blending, merging) three luma predictors: two intrapredictors derived from the two selected angular intraprediction modes M1 and M2, and one planar predictor (M. Abdoli et al, “Non-CE3: Decoder-side Intra Mode Derivation with Prediction Fusion Using Planar”, JVET-00449, Gothenburg, July 2019).
[084] As illustrated in Figure 10, the mixture is defined as a weighted average of these three intrapredictors using three derived weights wi, W2, W3 (step 361) of the ratio of the selected angular intraprediction amplitudes by: _ 43 ampl(M1)164 ampl(M1)+ampl(M2) ampl(M2) ampl(M1)+ampl(M2) (1) IV·! = —364
[085] With weights wi and W2 for the two angular intraprediction modes M1 and M2 and weight W3 for the planar mode, i.e., 21 / 64 with 6-bit integer precision. In step 362, the intrapredictors of the current luma block are calculated from the two selected angular intraprediction modes M1 and M2 and the planar mode. Finally, in step 363, the DIMD predictor is determined by mixing the three intrapredictors using the weights.
[086] Figure 11 schematically illustrates a block diagram of a 400 method for deriving a DIMD chroma mode according to the prior art. Petition 870260074236, dated 07 / 24 / 2026, page 24 / 80 16 / 48
[087] Basically, to derive the DIMD chroma mode, a similar method used to derive the DIMD luma mode is applied to colocalized reconstructed luma samples, i.e., a gradient analysis of colocalized reconstructed luma samples in an L-shaped model area defined around the current block of chroma samples.
[088] The use of the DIMD chroma mode is signaled in a bitstream by a single flag and the intrapredictor is not explicitly signaled in the bitstream, but derived using gradient analysis of co-located reconstructed luma samples. The DIMD chroma mode is implicit. Thus, the intrapredictor is derived from a DIMD chroma mode during the reconstruction process identically on the encoder and decoder sides.
[089] In step 410, method 400 determines whether reconstructed luma samples are available. These reconstructed luma samples are colocalized in the chroma block to be predicted or in an L-shaped model area around the chroma block. An L-shaped model area is formed from the available L-shaped model area.
[090] In step 420, a HoG gradient histogram is constructed as in step 320 for colocalized reconstructed luma samples available in the L-shaped area. Specifically, a horizontal gradient and a vertical gradient are calculated for each colocalized reconstructed luma sample (gray circles in Figure 12, extracted from JVET-Y0092) and the HoG is constructed from the horizontal and vertical gradients.
[091] In step 430, the amplitudes of available chroma samples in an L-shaped area around the chroma block are accumulated for the HoG, where each entry corresponds to an orientation, i.e., an angular intraprediction mode (the same as for the HoG luma counterpart).
[092] In one variant, the HoG can be calculated by cumulative amplitude values over Y, Cb, and Cr samples reconstructed from model areas in the form Petition 870260074236, dated 07 / 24 / 2026, page 25 / 80 17 / 48 of L (Figure 13), that is, a HoG derived by merging a first HoG calculated from neighboring colocalized luma samples of an L-shaped model area (left part of Figure 13), a second HoG calculated from Cb samples of an L-shaped model area (middle part of Figure 13), and a third HoG calculated from neighboring reconstructed Cr samples (right part of Figure 13).
[093] In step 440, an angular intraprediction mode corresponding to the largest amplitude value of the histogram is selected from HoG as the first DIMD chroma mode.
[094] When the DIMD chroma mode is not a Direct (DM) mode, then method 400 terminates and the DIMD chroma mode is the first DIMD chroma mode.
[095] Direct mode (DM) is a chroma intraprediction mode corresponding to the intraprediction mode of colocalized reconstructed luma samples.
[096] When the first DIMD chroma mode is a Direct Mode (DM), then at step 450 a second DIMD chroma mode is selected as an angular intraprediction mode corresponding to the second largest amplitude value of the histogram.
[097] When the first DIMD chroma mode is not equal to the second DIMD chroma mode, then the method terminates and the DIMD chroma mode is the second DIMD chroma mode.
[098] When the first DIMD chroma mode is the same as the second DIMD chroma mode, then in step 460, the DIMD chroma mode is Direct Encoding (DC).
[099] A nonlinear model-based predictor, in short, a non-LM predictor, of the chroma block is then derived from a non-LM mode selected from five standard modes and the DIMD chroma mode. The chroma predictors are derived from these non-LM modes and the selected non-LM mode corresponds to the non-LM predictor that minimizes a rate distortion trade-off.
[0100] The five standard modes are DC, Planar, Direct Mode, predictions Petition 870260074236, dated 07 / 24 / 2026, page 26 / 80 18 / 48 horizontal and vertical angles.
[0101] The non-LM predictor of the chroma block derived from the selected non-LM mode is then merged (mixed, blended) with the LM predictor derived from the MMLM mode as follows: pred = (w0 * predO + wl * predl) >> deviation where predO is the predictor obtained by applying the non-LM mode, predl is the predictor obtained by applying the LM mode, and pred is the final predictor of the chroma block. For I slices (Intra-code slices), the two weights, w0 and wl, are determined by the intraprediction mode of adjacent chroma blocks, and deviation is defined as 2. Specifically, when adjacent blocks above and to the left are both encoded with LM modes, {w0, w1}={1, 3}; when adjacent blocks above and to the left are both encoded with non-LM modes, {w0, w1}={3, 1}; otherwise, {w0, w1}={2, 2}. For non-I slices, w0 and w1 are both defined as equal to 2.
[0102] If a non-LM mode is selected, a flag is displayed to indicate whether merging is applied.
[0103] In a variant, for I slices, i.e., intracoded slice, non-LM predictors derived from DM mode, the four standard modes and the DIMD chroma mode can be merged with an LM predictor derived from LM mode, whereas for non-I slices, only the non-LM predictor derived from DIMD chroma mode can be merged with the LM predictor derived from LM mode using equal weights.
[0104] Chroma fusion mode can be used with chroma DIMD (if selected as the best non-LM mode), but any other non-LM mode can be used instead.
[0105] Using DIMD for luma and chroma improves video image encoding efficiency. However, the price to pay is an additional complexity of about 20% in the encoder and about 5% in the decoder for all intra conditions.
[0106] The problem solved by this disclosure is to maintain the encoding efficiency of next-generation video codecs and reduce their complexity when Petition 870260074236, dated 07 / 24 / 2026, p. 27 / 80 19 / 48 DIMD is used for luma and chroma.
[0107] At least one aspect of this disclosure has been conceived with the above in mind. SUMMARY
[0108] The following section presents a simplified summary of at least one modality in order to provide a basic understanding of some aspects of this disclosure. This summary is not an extensive overview of a modality. It is not intended to identify key or critical elements of a modality. The following summary presents only some aspects of at least one modality in a simplified form as a prelude to the more detailed description provided elsewhere in the document.
[0109] According to a first aspect of the present disclosure, a method is provided for decoding a block of samples from a video image, the method comprising determining an intraprediction mode to decode the block of samples according to an analysis of gradients of samples located in at least one defined model area around the block of samples by calculating a histogram of gradients, where each entry of the histogram of gradients corresponds to an angular intraprediction mode, filtering samples from at least one model area, said filtering using filter windows centered on midline sample positions of at least one model area; selecting at most two angular intraprediction modes by comparing amplitudes of angular intraprediction modes in the histogram of gradients; determining the intraprediction mode from at most two selected angular intraprediction modes;where an integer number of midline sample positions from at least one model area in which the filter windows are centered is less than a total integer number of midline sample positions from at least one model area.
[0110] According to a second aspect of this disclosure, it is Petition 870260074236, dated 07 / 24 / 2026, page 28 / 80 20 / 48 provided a method for encoding a block of samples from a video image, the method comprising determining an intraprediction mode to decode the block of samples according to an analysis of gradients of samples located in at least one defined model area around the block of samples by calculating a histogram of gradients, where each entry of the histogram of gradients corresponds to an angular intraprediction mode, filtering samples from at least one model area, said filtering using filter windows centered on midline sample positions of at least one model area; selecting at most two angular intraprediction modes by comparing amplitudes of angular intraprediction modes in the histogram of gradients; determining the intraprediction mode from at most two selected angular intraprediction modes;where an integer number of midline sample positions from at least one model area in which the filter windows are centered is less than a total integer number of midline sample positions from at least one model area.
[0111] In one embodiment, the gradient histogram for chroma components of the video image is calculated by filtering only chroma samples from at least one model area.
[0112] In one embodiment, at least one of the at least one model area comprises samples along a lower boundary of a neighboring Virtual Pipeline Data Unit above and along a right boundary of a neighboring Virtual Pipeline Data Unit to the left.
[0113] In one embodiment, at least one of the at least one model area comprises all luma samples on an edge of neighboring virtual Pipeline Data Units.
[0114] In one embodiment, the model area comprises a subset of luma samples on an edge of neighboring virtual Pipeline Data Units.
[0115] In one embodiment, a gradient histogram is calculated for each Petition 870260074236, dated 07 / 24 / 2026, page 29 / 80 21 / 48 chroma component of the video image.
[0116] In one embodiment, each gradient histogram for a chroma component of the video image is calculated from the sample luma and chroma model areas.
[0117] In one embodiment, the model's chroma samples are multiplied by a compensation coefficient that depends on a chroma subsampling defined by a video image format.
[0118] In one embodiment, the midline sample positions of at least one of at least one model area in which the filter windows are centered are determined to locate the filter windows at 1 of a quarter number (N4) of the midline sample positions of the model area.
[0119] In one embodiment, the midline sample positions of at least one of at least one model area in which the filter windows are centered are determined to avoid any overlap between the filter windows.
[0120] In one embodiment, the filtering windows have different sizes.
[0121] In one embodiment, the size of the filtering windows depends on the size of the sample block and the availability of the sample from the model area.
[0122] According to a third aspect of this disclosure, an apparatus is provided comprising means for carrying out one of the methods according to the first and / or second aspect of this disclosure.
[0123] According to a fourth aspect of this disclosure, a non-transient storage medium is provided that carries program code instructions to perform a method in accordance with the first and / or second aspect of this disclosure.
[0124] According to a fifth aspect of the present disclosure, an electronic device is provided, including: a processor; and a memory for storing instructions executable by the processor. The processor is configured to Petition 870260074236, dated 07 / 24 / 2026, p. 30 / 80 22 / 48 perform the method in accordance with the first and / or second aspect of this disclosure.
[0125] According to a sixth aspect of this disclosure, a computer program product is provided including instructions that, when executed by one or more processors, cause one or more processors to execute the method according to the first and / or second aspect of this disclosure.
[0126] The specific nature of at least one of the embodiments, as well as other objects, advantages, features and uses of said at least one of the embodiments, will become evident from the following description of examples taken together with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Reference will now be made, by way of example, to the attached drawings which show embodiments of the present disclosure, and in which: Figure 1 shows an example of a encoding tree unit according to HEVC; Figure 2 shows an example of partitioning coding units into prediction units according to HEVC; Figure 3 shows an example of a CTU division according to VVC; Figure 4 shows examples of splitting modes supported in multi-type tree partitioning according to VVC; Figure 5 shows a schematic block diagram of the steps of a method 100 for encoding a VP video image according to the previous technique; Figure 6 shows a schematic block diagram of the steps of a 200 method for decoding a VP video image according to the previous technique; Figure 7 schematically illustrates a block diagram of a 300 method for deriving a DIMD luma mode according to the prior art; Petition 870260074236, dated 07 / 24 / 2026, page 31 / 80 23 / 48 Figure 8 illustrates an L-shaped model area; Figure 9 illustrates an example of a HoG; Figure 10 schematically shows a method for merging multiple luma predictors according to the previous technique; Figure 11 schematically illustrates a block diagram of a 400 method for deriving a DIMD chroma mode according to the prior art; Figure 12 schematically shows examples of colocalized luma samples from a chroma block according to the previous technique; Figure 13 schematically shows other examples of colocalized lumen samples from a chroma block according to the previous technique; Figure 14 schematically illustrates a block diagram of a 500 method for deriving a DIMD mode according to a modality; Figure 15 schematically illustrates filter window positions (step 510) according to a mode; Figure 16 schematically illustrates filter window positions (step 510) according to a mode; Figure 17 schematically illustrates filter window positions (step 510) according to a mode; Figure 18 illustrates VPDU and an L-shaped model area around the actual VPDU according to a modality; Figure 19 illustrates the L-shaped model area when VPDUs are used according to a modality; Figure 20 illustrates the L-shaped model area when VPDUs are used according to a modality; Figure 21 schematically illustrates a block diagram of a 600 method for constructing a HoG for chroma DIMD derivation according to a modality; and Figure 22 illustrates a schematic block diagram of an example of Petition 870260074236, dated 07 / 24 / 2026, page 32 / 80 24 / 48 is a system in which various aspects and modalities are implemented.
[0128] Similar or identical elements are referenced with the same reference numbers. DETAILED DESCRIPTION
[0129] At least one embodiment is described more fully below with reference to the accompanying figures, in which examples of at least one embodiment are depicted. An embodiment may, however, be incorporated in many alternative forms and should not be interpreted as limited to the examples set forth in the present invention. Consequently, it should be understood that there is no intention to limit the embodiments to the particular forms disclosed. Rather, this disclosure is intended to cover all modifications, equivalents and alternatives that fall within the spirit and scope of this disclosure.
[0130] At least one aspect generally refers to the encoding and decoding of video images, another aspect generally refers to the transmission of a supplied or encoded bitstream, and yet another aspect refers to the reception / access to a decoded bitstream.
[0131] At least one of the modes is described for encoding / decoding a video image, but it extends to encoding / decoding video images (image sequences) because each video image is sequentially encoded / decoded as described below.
[0132] In addition, at least one of the modalities is not limited to MPEG standards, such as AVC (ISO / IEC 14496-10 Advanced video coding for general audiovisual services, ITU-T Recommendation H.264, https: / / www.itu.int / rec / T-REC-H.264-202108-P / en), EVC (ISO / IEC 23094-1 Essential video coding), HEVC (ISO / IEC 23008-2 High-efficiency video coding, ITU-T Recommendation H.265, https: / / www.itu.int / rec / T-REC-H.265202108-P / en), VVC (ISO / IEC 23090-3 Versatile video coding, ITU-T Recommendation H.266, https: / / www.itu.int / rec / T-REC-H.266-202008-'Jen, but it could be Petition 870260074236, dated 07 / 24 / 2026, page 33 / 80 25 / 48 applied to other standards and recommendations such as AV1 (AOMedia Video 1, http: / / aomedia.org / av1 / specification / ) for example. At least one embodiment may be applied to pre-existing or future-developed standards and recommendations, and extensions of any standards and recommendations. Unless otherwise indicated, or technically prevented, the aspects described in this disclosure may be used individually or in combination.
[0133] In general terms, the present disclosure relates to the decoding of a sample block of a video image in which a DIMD mode is derived for luma and chroma samples of a sample block to be predicted by filtering samples from at least one model area, said filtering using filter windows centered on midline sample positions of at least one model area filtering samples in at least one shaped model area. An integer number of midline sample positions of at least one model area in which the filter window is centered is less than a total integer number of midline sample positions of at least one model area.
[0134] This reduces the resource requirements (memory and computing power, complexity) for deriving the DIMD mode for luma and chroma samples of the block to be predicted compared with the resources required by the current DIMD mode derivation method.
[0135] The following embodiments are described using an L-shaped model area composed of reconstructed samples to the left, above, and above the left of the reconstructed area adjacent to a current block. But the present disclosure extends to model areas composed only of reconstructed samples to the left of the reconstructed area adjacent to a current block, or model areas composed only of reconstructed samples above the reconstructed area adjacent to a current block.
[0136] Figure 14 schematically illustrates a block diagram of a 500 method for deriving a DIMD mode according to a modality. Petition 870260074236, dated 07 / 24 / 2026, page 34 / 80 26 / 48
[0137] The 500 method can be applied to both the 100 method (encoding) and the 200 method (decoding) for intraprediction of a block.
[0138] In step 510, midline sample positions of an L-shaped model area T1 are determined. The integer number N1 of midline sample positions of the L-shaped model area T1 in which the filtering window is centered is less than the total number of midline sample positions of the L-shaped model area.
[0139] In step 520, a luma DIMD mode is derived from method 300 in which the filter windows are centered on each of the midline sample positions N1 of an L-shaped model area T1.
[0140] In a variant of step 520, the midline sample positions of a T2 L-shaped model area of co-located reconstructed luma samples are determined. The midline sample positions of a T3 L-shaped model area of available chroma samples are determined. The number N2, respectively N3, of midline sample positions of a T2 L-shaped model area, respectively T3, where the filter windows are centered, is less than the total number of midline sample positions of the T2 L-shaped model area, respectively T3. A chroma DIMD mode is derived from method 400 where the filter windows are centered at each of the N2 midline sample positions of the T2 L-shaped model area and at each of the N3 midline sample positions of the T3 L-shaped model area.
[0141] In one embodiment of step 510, the midline sample positions of at least one of the at least one L-shaped model area (T1, T2, T3) in which the filter windows are centered, are determined to locate the filter windows in 1 of a number N4 (N4 >= 2) of the midline sample positions of the L-shaped model area T1, T2 or T3.
[0142] Figure 15 schematically illustrates an example of window positions. Petition 870260074236, dated 07 / 24 / 2026, page 35 / 80 27 / 48 filtering points centered on sample positions along the midline (white circles) when N1=2. The filtering windows are represented by dashed lines.
[0143] In one embodiment of step 510, the midline sample positions of at least one of the at least one L-shaped model area (T1, T2, T3) in which the filter windows are centered are determined to avoid any overlap between the filter windows.
[0144] Figure 16 schematically illustrates an example of centered filter window positions at determined midline sample positions (white circles). The filter windows are represented by dashed lines.
[0145] In one embodiment of step 510, the midline sample positions of at least one of the at least one L-shaped model area (T1, T2, T3) in which the filter windows are centered are determined to allow the filter windows to have a common edge. These midline sample positions of an L-shaped model area can then be determined between the above left, above, and top locations of the L-shaped model area.
[0146] Figure 17 schematically illustrates an example of centered filter window positions at determined midline sample positions (white circles). The filter windows are represented by dashed lines.
[0147] In one embodiment of the 500 method, the filtering window size is larger than 3x3 to compensate for the reduced number of filtered samples (i.e., reduced number of HoG entries). For example, a 5x5 filter size is employed.
[0148] In one embodiment of the 500 method, the filtering windows have different sizes.
[0149] In one embodiment of the 500 method, the size of the filtering windows depends on the size of the block to be predicted and the availability of the sample area of the L-shaped model.
[0150] For example, depending on sample availability, the size of Petition 870260074236, dated 07 / 24 / 2026, page 36 / 80 The 28 / 48 filtering window is increased to the L-shaped model area in the larger block dimension (possibly when the block size exceeds a predetermined size, for example 8x8).
[0151] The previous methods for determining the midline sample positions of an L-shaped model area in which the filter windows are centered and the filter window size can be combined.
[0152] For example, a filtering window is sized 5x5 for samples in an L-shaped model area adjacent to the longest block size and the allowed midline sample positions are in specific positions or located in evenly spaced positions or located in positions where the filtering windows do not overlap.
[0153] As the number of midline sample positions is reduced (step 510) and may be less representative of the directions present in the neighborhood of the actual block to be predicted, in one embodiment, the derivation of the mixing weights (for luma) is modified so that, instead of calculating the weights wi and W2 by equation (1), the weights wi and W2 are calculated as follows: If the amplitude of the first angular intraprediction mode M1 is greater than twice the amplitude of the second angular intraprediction mode M2, then the second weight W2 equals 0 and, for example, w±= — and w3= 21 / 64.
[0154] In one variant, the third weight W3 is equal to 0 and = =—. - - · - - 64 if the amplitude of the first angular intraprediction mode M1 is less than or equal to twice the amplitude of the second angular intraprediction mode M2, then the second weight W2 is not equal to 0.
[0155] For example, w±= w2= w3= —. - - · - - 64
[0156] In one variant, the third weight is equal to 0 and for example, w= = — ew3= 32 / 64.
[0157] As discussed above in the introductory section, chroma samples in the L-shaped model area and reconstructed luma samples colocalized in Petition 870260074236, dated 07 / 24 / 2026, page 37 / 80 29 / 48 an L-shaped model is used to construct the HoG. This requires that luma samples be first encoded and decoded before chroma encoding with the DIMD chroma tool, adding latency from the 100 and 200 encoding / decoding methods. Furthermore, if the DIMD luma mode is not enabled to predict a luma block, the method for deriving the DIMD chroma mode must, in any case, calculate the HoG from co-located reconstructed luma samples to obtain a DIMD chroma predictor. According to the hardware architecture design, the computation of HoG from co-located reconstructed luma samples can be repeated and redundant between the derivation of the DIMD luma mode and the DIMD chroma mode.
[0158] In one embodiment of method 500, to derive a DIMD mode for a chroma component of the video image, the HoG is calculated only according to chroma samples filtered from the L-shaped model area.
[0159] This mode applies only to deriving DIMD chroma mode.
[0160] This method is advantageous because it avoids encoding and decoding luma samples before encoding chroma samples using DIMD chroma mode. This method is also advantageous because of its low complexity, low latency, and it requires less computing power and memory compared to the method that used colocalized reconstructed luma and chroma samples.
[0161] In VVC, the Virtual Pipeline Data Unit (VPDU) is a concept of non-overlapping sample units, typically sized 64x64 for luma and 32x32 for chroma. The goal is to ensure that the VPDU is fully processed before starting the processing of the next VPDU so that the memory footprint of the hardware implementation remains reasonable. The VPDU can have a strong impact on tool design.
[0162] Thus, when the VPDU is considered in the design of the chroma DIMD and Petition 870260074236, dated 07 / 24 / 2026, page 38 / 80 30 / 48 when constructing the HoG to derive the DIMD predictor uses not only chroma samples present in the available model areas, but also luma samples (as discussed above), these luma samples are retrieved from a current VPDUc so that the luma samples are fully processed and made available.
[0163] In one embodiment of method 500, at least one of the at least one L-shaped model area comprising luma samples (e.g., the width of at least 3 samples) along the lower VPDU boundary of the neighboring upper VPDU (VPDU1 in Figure 18) and along the right VPDU boundary of the neighboring left VPDU (VPDU2 in Figure 18) is employed.
[0164] This mode reduces the latency of method 500.
[0165] In one variant, at least one of the at least one L-shaped model area comprises all (i.e., 64x2 in the VVC projection or 64x3 with a width of 3) luma samples on an edge of the neighboring VPDUs VPDU1 and VPDU2 (Figure 19). All of these luma samples are used in the construction of the HoG instead of using neighboring colocalized luma samples.
[0166] In one variant, a subset of the luma samples on an edge of the neighboring VPDUs VPDU1 and VPDU2 (Figure 20) is used. These luma samples are used in the construction of the HoG instead of using neighboring colocalized luma samples.
[0167] In a variant, if a subset of luma samples belonging to a 3-sample-wide edge of the neighboring VPDU is used, this subset corresponds to colocalized luma samples with (part of) orthogonal projections of the chroma block onto the neighboring VPDUs, as shown in Figures 18, 19, and 20.
[0168] Only luma samples from the current VPDU VPDUc are used in the construction of the HoG to derive the chroma DIMD predictor. In the case where the neighboring VPDU is not present, the chroma samples present in L-shaped model areas are used alone to construct the HoG. Petition 870260074236, dated 07 / 24 / 2026, page 39 / 80 31 / 48
[0169] In one embodiment of the 500 method, a HoG is calculated for each chroma component of the video image.
[0170] This mode allows parallelization to calculate HoG by chroma component and therefore improves the latency of the 500 method.
[0171] Furthermore, this mode improves signal adaptation by separating the HoG construction for chroma components and, consequently, the final intraprediction mode that is allocated independently to each chroma component when the chroma DIMD is enabled.
[0172] This mode provides even more flexibility in the design, as the final intraprediction mode selected for each chroma component can be different. This mode does not require other encoding tools to act separately on chroma encoding blocks or components. The advantage is that no extra signaling (more than the DIMD chroma flag) is needed for the modes of both components.
[0173] Figure 21 schematically illustrates a block diagram of a 600 method for constructing a HoG for chroma DIMD derivation according to a modality.
[0174] In this embodiment, the chroma samples available in an L-shaped model area and the reconstructed colocalized luma samples available in an L-shaped model area are used in the construction of each HoG chroma component.
[0175] However, since in non-4:4:4 signals the chroma samples are smaller than the luma samples, a compensation coefficient is introduced in the HoG construction.
[0176] In step 610, a current filtering window W is considered.
[0177] In step 620, a horizontal Dx gradient and a vertical Dy gradient are computed by applying horizontal and vertical edge detection filters, such as horizontal and vertical 3x3 Sobel filters. Petition 870260074236, dated 07 / 24 / 2026, page 40 / 80 32 / 48
[0178] In step 635, a compensation coefficient is determined according to the chroma subsampling of the signal, as defined by a video image format.
[0179] For example, the compensation coefficient is equal to 2 for 4:2:0 signals, 1.5 for 4:2:2 signals and 1 for a 4:4:4 signal.
[0180] In one variant, this compensation coefficient is a multiple of 2 for 4:2:0 signals, 1.5 for 4:2:2 signals, and 1 for a 4:4:4 signal.
[0181] Alternatively, for 4:2:0 signals, one of two (uniformly spaced) luma samples from the L-shaped model area are used as centers of the filtering windows to construct the HoG (as in Figure 15) so that the number of processed midline luma samples equals the number of processed chroma samples (for each chroma component).
[0182] In a variant of this embodiment of Figure 21, only the chroma samples available in an L-shaped model area are used in the construction of each chroma component HoG. Filtering window positions are then added to compensate for the lack of samples. Typically, the third row or column of chroma samples is also used as a centered filtering window position to construct each chroma HoG.
[0183] Figure 22 shows a schematic block diagram illustrating an example of a 700 system in which various aspects and modalities are implemented.
[0184] The 700 system can be incorporated as one or more devices, including the various components described below. In various embodiments, the 700 system can be configured to implement one or more of the aspects described in this disclosure.
[0185] Examples of equipment that can form all or part of the 700 system include personal computers, laptops, smartphones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, connected vehicles and their systems. Petition 870260074236, dated 07 / 24 / 2026, page 41 / 80 33 / 48 associated processing, head-mounted display devices (HMDs, transparent glasses), projectors (beamers), “caves” (system including multiple displays), servers, video encoders, video decoders, post-processors processing output from a video decoder, pre-processors providing input to a video encoder, web servers, video servers (e.g., a streaming server, a video-on-demand server, or a web server), still or video camera, encoding or decoding chip, or any other communication devices. Elements of the 700 system, individually or in combination, may be incorporated into a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of the 700 system may be distributed across multiple ICs and / or discrete components.In various configurations, the 700 system can be communicatively coupled to other similar systems, or to other electronic devices, by means of, for example, a communications bus or by means of dedicated input and / or output ports.
[0186] The 700 system may include at least one 710 processor configured to execute instructions loaded into it to implement, for example, the various aspects described in this disclosure. The 710 processor may include embedded memory, input / output interface, and various other circuits known in the art. The 700 system may include at least one 720 memory (for example, a volatile memory device and / or a non-volatile memory device). The 700 system may include a 740 storage device, which may include non-volatile memory and / or volatile memory, including, but not limited to, electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash, magnetic disk drive, and / or disk drive. Petition 870260074236, dated 07 / 24 / 2026, page 42 / 80 34 / 48 optical. The 740 storage device may include an internal storage device, an attached storage device, and / or a network-accessible storage device, as non-limiting examples.
[0187] The 700 system may include a 730 encoder / decoder module configured, for example, to process data to provide encoded / decoded video image data, and the 730 encoder / decoder module may include its own processor and memory. The 730 encoder / decoder module may represent module(s) that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both encoding and decoding modules. Furthermore, the 730 encoder / decoder module may be implemented as a separate element of the 700 system or may be incorporated within the 710 processor as a hardware and software combination, as is known to those skilled in the art.
[0188] The program code to be loaded into processor 710 or encoder / decoder 730 to perform the various aspects described in this disclosure may be stored in storage device 740 and subsequently loaded into memory 720 for execution by processor 710. According to various embodiments, one or more of the processor 710, memory 720, storage device 740, and encoder / decoder module 730 may store one or more of several items during the execution of the processes described in this disclosure. Such stored items may include, but are not limited to, video image data, information data used to encode / decode video image data, a bitstream, matrices, variables, and intermediate or final results of the processing of equations, formulas, operations, and operational logic.
[0189] In several embodiments, the memory within the 710 processor and / or the 730 encoder / decoder module can be used to store instructions and Petition 870260074236, dated 07 / 24 / 2026, page 43 / 80 35 / 48 provide working memory for processing that can be performed during encoding or decoding.
[0190] In other embodiments, however, memory external to the processing device (for example, the processing device may be the processor 710 or the encoder / decoder module 730) may be used for one or more of these functions. The external memory may be the memory 720 and / or the storage device 740, for example, volatile dynamic memory and / or non-volatile flash memory. In several embodiments, non-volatile external flash memory may be used to store the operating system of a television. In at least one embodiment, fast external dynamic volatile memory, such as RAM, may be used as working memory for video encoding and decoding operations, such as for MPEG-2 part 2 (also known as ITU-T Recommendation H.262 and ISO / IEC 13818-2, also known as MPEG-2 Video), AVC, H EVC, EVC, VVC, AV1 etc.
[0191] Input to the 700 system elements may be provided by means of various input devices, as indicated in block 790. Such input devices include, but are not limited to, (i) an RF portion that can receive an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, (iv) an HDMI input terminal, (v) a bus, such as CAN (Controller Area Network), CAN FD (Controller Area Network Flexible Data Rate Bus), FlexRay (ISO 17458) or Ethernet bus (ISO / IEC 802-3) when this disclosure is implemented in the automotive domain.
[0192] In several embodiments, the 790 block input devices may have respective associated input processing elements, as known in the art. For example, the RF portion may be associated with elements necessary to (i) select a desired frequency (also referred to as signal selection), or limit the bandwidth of a signal to a frequency band), (ii) Petition 870260074236, dated 07 / 24 / 2026, p. 44 / 80 36 / 48 convert the selected signal down, (iii) limit the band again to a narrower band of frequencies to select (e.g.) a signal frequency band that may be termed a channel in certain modes, (iv) demodulate the down-converted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select the desired stream of data packets. The RF portion of various modes may include one or more elements to perform these functions, for example, frequency selectors, signal selectors, band limiters, channel selectors, filters, down-converters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs several of these functions, including, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband.
[0193] In one embodiment of a decoder, the RF portion and its associated input processing element can receive an RF signal transmitted over a wired medium (e.g., cable). Then, the RF portion can perform frequency selection by filtering, down-conversion, and filtering again to a desired frequency band.
[0194] Various modalities rearrange the order of the elements described above (and others), remove some of these elements and / or add other elements that perform similar or different functions.
[0195] Adding elements can include inserting elements between existing elements, such as inserting amplifiers and an analog-to-digital converter. In several embodiments, the RF portion may include an antenna.
[0196] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting the 700 system to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, for example, error correction, Petition 870260074236, dated 07 / 24 / 2026, page 45 / 80 37 / 48 Reed-Solomon interfaces can be implemented, for example, within a separate input processing IC or within the 710 processor as needed. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface ICs or within the 710 processor as needed. The demodulated, error-corrected, and demultiplexed stream can be provided to various processing elements, including, for example, the 710 processor and the 730 encoder / decoder operating in combination with the memory and storage elements to process the data stream as needed for presentation to an output device.
[0197] Several elements of the 700 system can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data between them using a suitable connection arrangement 790, for example, an internal bus as known in the art, including the I2C bus, wiring and printed circuit boards.
[0198] The 700 system may include the 750 communication interface which allows communication with other devices via the 751 communication channel. The 750 communication interface may include, but is not limited to, a transceiver configured to transmit and receive data via the 751 communication channel. The 750 communication interface may include, but is not limited to, a modem or network card and the 751 communication channel may be implemented, for example, within a wired and / or wireless medium.
[0199] Data can be transmitted to the 700 system in various modes using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal from these modes can be received by the 751 communications channel and the 750 communications interface, which are adapted for Wi-Fi communications. The 751 communications channel from these modes can typically be connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Petition 870260074236, dated 07 / 24 / 2026, page 46 / 80 38 / 48
[0200] Other modes may provide data transmitted to the 700 system using a decoder that provides the data via the HDMI connection of the 790 input block.
[0201] Still other modes can provide data transmitted to the 700 system using the RF connection of the 790 input block.
[0202] The transmitted data can be used as a form of signaling information used by the 700 system. The signaling information may comprise the B bitstream and / or information such as a number of pixels in a 7a video image and / or any encoding / decoding configuration parameters.
[0203] It should be appreciated that signaling can be carried out in several ways. For example, one or more syntax elements, flags, and so on can be used to signal information to a corresponding decoder in various modes.
[0204] System 700 can provide an output signal to various output devices, including a display 761, speakers 771, and other peripheral devices 781. Other peripheral devices 781 may include, in various embodiments, one or more standalone DVRs, a disc player, a stereo system, a lighting system, and other devices that provide a function based on the output of System 700.
[0205] In various embodiments, control signals can be communicated between the system 700 and the display 761, loudspeakers 771 or other peripheral devices 781 using signaling such as AV.Link (AudioNideo Link), CEC (Consumer Electronics Control) or other communication protocols that allow device-to-device control with or without user intervention.
[0206] Output devices can be communicatively coupled to the 700 system via dedicated connections through the respective 760, 770 and 780 interfaces. Petition 870260074236, dated 07 / 24 / 2026, page 47 / 80 39 / 48
[0207] Alternatively, the output devices can be connected to the system 700 using the communication channel 751 via the communication interface 750. The display 761 and the speakers 71 can be integrated into a single unit with the other components of the system 700 in an electronic device, such as, for example, a television.
[0208] In several embodiments, the 760 display interface may include a display driver, such as, for example, a time controller chip (T Con).
[0209] The display 761 and the speaker 771 can alternatively be separated from one or more of the other components, for example, if the RF portion of input 790 is part of a separate decoder. In various embodiments in which the display 761 and the speakers 771 can be external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports or COMP outputs.
[0210] In Figures 1 to 22, various methods are described in the present invention, and each of the methods includes one or more steps or actions to achieve the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0211] Some examples are described with respect to operational block diagrams and / or flowcharts. Each block represents a circuit element, module, or piece of code that includes one or more executable instructions to implement the specified logical function(s). It should also be noted that in other implementations, the function(s) shown in the blocks may occur out of the order indicated. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order, depending on the functionality involved.
[0212] The implementations and aspects described in the present invention can be implemented in, for example, a method or a process, an apparatus, a Petition 870260074236, dated 07 / 24 / 2026, pp. 48 / 80 40 / 48 computer program, a data stream, a bit stream, or a signal. Even if discussed only in the context of a single implementation form (e.g., discussed only as a method), the implementation of the discussed features can also be implemented in other forms (e.g., a device or computer program).
[0213] The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices.
[0214] Additionally, the methods can be implemented by instructions being executed by a processor, and such instructions (and / or data values produced by an implementation) can be stored in a computer-readable storage medium. A computer-readable storage medium can take the form of a computer-readable program product embedded in one or more computer-readable media and having computer-readable program code embedded therein that is executable by a computer. A computer-readable storage medium as used in the present invention can be considered a non-transient storage medium, given the inherent ability to store the information contained therein, as well as the inherent ability to provide retrieval of the information from it.A computer-readable storage medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. It should be appreciated that the following, while providing more specific examples of computer-readable storage media to which the present embodiments may be applied, is merely an illustrative and not exhaustive list, as is readily apparent to one with ordinary skill in the art: a portable computer floppy disk; a hard disk; a read-only memory (ROM); Petition 870260074236, dated 07 / 24 / 2026, page 49 / 80 41 / 48 a programmable erasable read-only memory (EPROM or Flash memory); a portable compact disk read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination thereof.
[0215] The instructions can form an application program tangibly embedded in a processor-readable medium.
[0216] Instructions can be found, for example, in hardware, firmware, software, or a combination thereof. Instructions can be found, for example, in an operating system, a separate application, or a combination of both. A processor can therefore be characterized as, for example, both a device configured to execute a process and a device that includes a processor-readable medium (such as a storage device) having instructions to execute a process. Additionally, a processor-readable medium can store, in addition to or instead of instructions, data values produced by an implementation.
[0217] A device can be implemented in, for example, appropriate hardware, software, and firmware. Examples of such devices include personal computers, laptops, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected headsets, head-mounted display devices (HMDs, transparent glasses), projectors (beamers), “caves” (systems including multiple monitors), servers, video encoders, video decoders, post-processors processing output from a video decoder, pre-processors providing input to a video encoder, web servers, decoders, and any other device for processing video images or other communication devices. As should be clear, the equipment can be mobile and even installed in a mobile vehicle.
[0218] Computer software can be implemented by the processor Petition 870260074236, dated 07 / 24 / 2026, pages 50 / 80 42 / 48 710 can be implemented either by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may also be implemented by one or more integrated circuits. 720 memory may be of any type appropriate to the technical environment and may be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. 710 processors may be of any type appropriate to the technical environment and may encompass one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on a multi-core architecture, as non-limiting examples.
[0219] As will be evident to anyone with ordinary skill in the art, implementations can produce a variety of formatted signals to carry information that can, for example, be stored or transmitted. The information may include, for example, instructions for performing a method or data produced by one of the described implementations. For example, a signal may be formatted to carry the bit stream of a described mode. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known.The signal can be stored on a processor-readable medium.
[0220] The terminology used in the present invention is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in the present invention, the singular forms “a”, “an”, “the” and “the” may be intended to include the plural forms as well, unless the context indicates otherwise. Petition 870260074236, dated 07 / 24 / 2026, pp. 51 / 80 43 / 48 clearly the opposite. It will also be understood that the terms “includes / comprises” and / or “including / comprises” when used in this descriptive report, may specify the presence of, for example, declared resources, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other resources, integers, steps, operations, elements, components and / or groups thereof. Furthermore, when an element is described as being “responsive” or “connected” or “associated with” another element, it may be directly responsive or connected or associated with the other element, or intervening elements may be present. In contrast, when an element is described as being “directly responsive” or “directly connected” or “directly associated with” another element, there are no intervening elements present.
[0221] It should be noted that the use of any of the symbols / terms “ / ”, “and / or” and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, may be intended to cover the selection of the first option listed (A) only, or the selection of the second option listed (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B and / or C” and “at least one of A, B and C”, such phrasing is intended to cover the selection of the first option listed (A) only, or the selection of the second option listed (B) only, or the selection of the third option listed (C) only, or the selection of the first and second options listed (A and B) only, or the selection of the first and third options listed (A and C) only, or the selection of the second and third options listed (A and (B and C) only), or the selection of all three options (A and B and C).This can be extended, as is clear to anyone with average skill in this and related techniques, to as many items as are listed.
[0222] Various numerical values may be used in this disclosure. The specific values may be for illustrative purposes and the aspects described are not limited to those specific values.
[0223] It will be understood that, although the terms first, second, etc. may Petition 870260074236, dated 07 / 24 / 2026, page 52 / 80 44 / 48 being used in the present invention to describe various elements, these elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be called a second element and, similarly, a second element may be called a first element without departing from the teachings of the present disclosure. No ordering is implied between a first element and a second element.
[0224] The reference to “an embodiment” or “an embodiment” or “an implementation” or “an implementation”, as well as other variations thereof, is frequently used to convey that a feature, structure, characteristic, and so forth (described in connection with the embodiment / implementation) is included in at least one embodiment / implementation. Thus, the appearances of the phrase “in an embodiment” or “in an embodiment” or “in an implementation” or “in an implementation”, as well as any other variations, that appear in various places throughout this disclosure are not necessarily all referring to the same embodiment.
[0225] Similarly, the reference in the present invention to “according to an embodiment / example / implementation” or “in an embodiment / example / implementation”, as well as other variations thereof, is frequently used to convey that a particular feature, structure or characteristic (described in connection with the embodiment / example / implementation) can be included in at least one embodiment / example / implementation. Thus, the appearances of the expression “according to an embodiment / example / implementation” or “in an embodiment / example / implementation” in various places in the present disclosure are not necessarily all referring to the same embodiment / example / implementation, nor are they necessarily separate or alternative embodiments / examples / implementations. Petition 870260074236, dated 07 / 24 / 2026, page 53 / 80 45 / 48 mutually exclusive of other modalities / examples / implementation.
[0226] The reference numbers appearing in the claims are for illustrative purposes only and shall not have a limiting effect on the scope of the claims. Although not explicitly stated, the present embodiments / examples and variants may be employed in any combination or subcombination.
[0227] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding device. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process.
[0228] Although some of the diagrams include arrows on communication paths to show a primary direction of communication, it should be understood that communication can occur in the opposite direction to the arrows shown.
[0229] Several implementations involve decoding. “Decoding,” as used in this application, may encompass all or part of the processes performed, for example, on a received video image (including possibly a received bitstream that encodes one or more video images) to produce a final output suitable for display or for further processing in the reconstructed video domain. In several embodiments, such processes include one or more of the processes normally performed by a decoder. In several embodiments, such processes also, or alternatively, include processes performed by a decoder of several implementations described in this disclosure, for example,
[0230] As further examples, in one modality “decoding” may refer only to dequantization, in another modality “decoding” may refer to entropy decoding, in yet another modality “decoding” may refer only to differential decoding, and in yet another modality “decoding” may Petition 870260074236, dated 07 / 24 / 2026, page 54 / 80 46 / 48 refers to combinations of dequantization, entropy decoding, and differential decoding. Whether the phrase “decoding process” refers specifically to a subset of operations or generally to the broader decoding process will become clear based on the context of the specific description and is believed to be well understood by those versed in the art.
[0231] Various implementations involve encoding. Analogous to the discussion above about “decoding,” “encoding” as used in this disclosure may encompass all or part of the processes performed, for example, on an input video image to produce an output bitstream. In various embodiments, such processes include one or more of the processes normally performed by an encoder. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application.
[0232] As further examples, in one embodiment “coding” may refer only to quantization, in another embodiment “coding” may refer only to entropy coding, in yet another embodiment “coding” may refer only to differential coding, and in yet another embodiment “coding” may refer to combinations of quantization, differential coding, and entropy coding. Whether the phrase “coding process” may refer specifically to a subset of operations or generally to the broader coding process will become clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art.
[0233] Additionally, this disclosure may refer to the “obtaining” of various information. Obtaining information may include one or more of the following: for example, estimating information, calculating information, predicting information or retrieving information from memory, processing information, moving information, copying information, deleting information, calculating information, determining information, predicting information or estimating information. Petition 870260074236, dated 07 / 24 / 2026, pp. 55 / 80 47 / 48 information.
[0234] Additionally, this request may refer to “receiving” various information. Receiving the information may include one or more of, for example, accessing the information, or receiving information from a communication network.
[0235] Furthermore, as used in the present invention, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals particular information, such as encoding parameters or encoded video image data. In this way, in one embodiment the same parameter can be used on both the encoder and decoder sides. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. On the other hand, if the decoder already has the particular parameter, as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functions, bit saving is achieved in several embodiments.It should be appreciated that signaling can be carried out in various ways. For example, one or more syntax elements, flags, and so on are used to signal information to a corresponding decoder in various modalities. While the foregoing relates to the verbal form of the word “signal,” the word “signal” can also be used in the present invention as a noun.
[0236] Several implementations have been described. However, it will be understood that various modifications can be made. For example, elements from different implementations can be combined, supplemented, modified, or removed to produce other implementations. Furthermore, anyone with average ability will understand that other structures and processes can be substituted for those disclosed, and the resulting implementations will perform at least Petition 870260074236, dated 07 / 24 / 2026, pp. 56 / 80 48 / 48 substantially the same function(s), at least substantially in the same manner(s), to achieve at least substantially the same result(s) as the disclosed implementations. Consequently, these and other implementations are contemplated by this request. Petition 870260074236, dated 07 / 24 / 2026, pp. 57 / 80
Claims
1 / 3 CLAIMS 1. A method for decoding a block of samples from a video image, the method CHARACTERIZED in that it comprises determining an intraprediction mode for decoding the block of samples according to an analysis of gradients of samples located in at least one model area (T, T1, T2, T3) defined around the block of samples by: - calculating (320, 430) a histogram of gradients, where each entry of the gradient histogram corresponds to an angular intraprediction mode, filtering samples from at least one model area, said filtering using filter windows (W) centered on midline sample positions of at least one model area; - selecting (330, 440) at most two angular intraprediction modes by comparing amplitudes of angular intraprediction modes in the gradient histogram; - determine (360, 450) the intraprediction mode from the maximum of two selected angular intraprediction modes;wherein at least one of the at least one model area comprises samples along a lower boundary of a neighboring virtual pipeline data unit (VPDU1) above and along a right boundary of a neighboring virtual pipeline data unit (VPDU2) to the left, a virtual pipeline data unit being a set of non-overlapping sample blocks that is completely processed before processing of a subsequent virtual pipeline data unit begins; and wherein an integer number of midline sample positions of the at least one model area in which the filtering windows are centered is less than a total integer number of midline sample positions of the at least one model area.
2. Method for encoding a block of samples from a video image, Petition 870260074236, dated 07 / 24 / 2026, page 58 / 80 2 / 3 method CHARACTERIZED by the fact that it comprises determining an intraprediction mode for decoding the block of samples according to an analysis of gradients of samples located in at least one model area (T, T1, T2, T3) defined around the block of samples by: - calculating (320, 430) a histogram of gradients, where each entry of the gradient histogram corresponds to an angular intraprediction mode, filtering samples from at least one model area, said filtering using filter windows (W) centered on midline sample positions of at least one model area; - select (330, 440) a maximum of two angular intraprediction modes by comparing amplitudes of angular intraprediction modes in the gradient histogram;- determine (360, 450) the intraprediction mode from the maximum of two selected angular intraprediction modes; wherein at least one of the at least one model area comprises samples along a lower boundary of a neighboring virtual pipeline data unit (VPDU1) above and along a right boundary of a neighboring virtual pipeline data unit (VPDU2) to the left, a virtual pipeline data unit being a set of non-overlapping sample blocks that is completely processed before processing of a subsequent virtual pipeline data unit begins; and wherein an integer number of midline sample positions of the at least one model area in which the filtering windows are centered is less than a total integer number of midline sample positions of the at least one model area.
3. Method, according to claim 1 or 2, CHARACTERIZED in that at least one of the model areas comprises all luma samples on an edge of the neighboring virtual Pipeline Data Units (VPDU1, Petition 870260074236, dated 07 / 24 / 2026, p. 59 / 80 3 / 3 VPDU2) above and to the left.
4. Method, according to claim 1 or 2, CHARACTERIZED in that at least one model area comprises a subset of luma samples on an edge of the neighboring virtual pipeline data units (VPDU1, VPDU2) above and to the left.
5. Method, according to claims 1 to 4, CHARACTERIZED in that the gradient histogram for chroma components of the video image is calculated by filtering only chroma samples from at least one model area.
6. A method, according to any one of claims 1 to 5, characterized in that the filtering windows have different sizes.
7. Method, according to any one of claims 1 to 6, CHARACTERIZED in that the size of the filtering windows depends on the size of the sample block and the availability of the sample from the model area.
8. Non-transient storage medium CHARACTERIZED in that it carries instructions which, when executed by a processor, cause the processor to execute a method as defined in any of claims 1 to 7.
9. Electronic device CHARACTERIZED in that it comprises: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the method as defined in any one of claims 1 to 7. Petition 870260074236, dated 07 / 24 / 2026, pp. 60 / 80