Block partitioning acceleration based on reinforcement learning

EP4802707A1Pending Publication Date: 2026-09-09INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2024789919
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-31
Filing Date
2024-10-17
Publication Date
2026-09-09

AI Technical Summary

Technical Problem

Existing video encoding and decoding methods face challenges in efficiently predicting partitioning modes for block-based video codecs, leading to prolonged encoding times due to exhaustive rate-distortion optimization searches.

Method used

The implementation of a reinforcement learning-based method that uses a neural network to predict partitioning modes for video blocks, allowing for accelerated encoding by directly inferring suitable splitting modes from neighbor block information and spatial representations.

Benefits of technology

This approach significantly reduces encoding time while maintaining compression efficiency, as it enables the prediction of optimal partitioning modes without exhaustive searches, thereby improving the overall performance of video codecs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024079327_08052025_PF_FP_ABST
    Figure EP2024079327_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and apparatus are provided for block partitioning acceleration based on reinforcement learning by a neural network. In one embodiment, a neural network receives inputs comprising at least one of neighbor features, parent features, coding unit information, and spatial features to determine block partitioning. In another embodiment, spatial features used by a neural network comprise a concatenation of two histograms of oriented gradients is used to determine block partitioning. In another embodiment, information is signaled from an encoder to a decoder for neural network determination of block partitioning.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]BLOCK PARTITIONING ACCELERATION BASED ON REINFORCEMENT LEARNING CROSS REFERENCE TO RELATED APPLICATION This application claims the benefit of European Serial No.23306885.7 filed October 31, 2023, which is incorporated by reference herein in its entirety. TECHNICAL FIELD At least one of the present embodiments generally relates to a method or an apparatus for video encoding or decoding, compression or decompression. BACKGROUND To achieve high compression efficiency, image and video coding schemes usually employ prediction, including motion vector prediction, and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter frame correlation, then the differences between the original image and the predicted image, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction. SUMMARY At least one of the present embodiments generally relates to a method or an apparatus for video encoding or decoding, and more particularly, to a method or an apparatus for prediction of partitioning modes with reinforcement learning (RL) in hybrid block-based video codecs. According to a first aspect, there is provided a method. The method comprises steps for performing a partitioning prediction of a video block using a neural network; partitioning the video block based on said partitioning prediction; and, encoding the partitioned video block. According to a second aspect, there is provided another method. The method comprises steps for performing a partitioning prediction of a video block using a neural network; partitioning the video block based on said partitioning prediction; and, decoding the partitioned video block. According to another aspect, there is provided an apparatus. The apparatus comprises a processor and a memory. The processor can be configured to operate on digital video data according to the aforementioned methods. According to another aspect, there is provided an apparatus. The apparatus comprises a processor and a memory. The processor can be configured to encode a block of a video or decode video data by executing any of the aforementioned methods. According to another general aspect of at least one embodiment, there is provided a device comprising an apparatus according to any of the decoding embodiments; and at least one of (i) an antenna configured to receive a signal, the signal including the video block, (ii) a band limiter configured to limit the received signal to a band of frequencies that includes the video block, or (iii) a display configured to display an output representative of the video block. According to another general aspect of at least one embodiment, there is provided a non-transitory computer readable medium containing data content generated according to any of the described encoding embodiments or variants. According to another general aspect of at least one embodiment, there is provided a signal comprising video data generated according to any of the described encoding embodiments or variants. According to another general aspect of at least one embodiment, video data or a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variants. According to another general aspect of at least one embodiment, there is provided a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out any of the described decoding embodiments or variants. These and other aspects, features and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments, which is to be read in connection with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 illustrates an example of partitioning of a luminance channel and two chrominance channels of an image. Figure 2 illustrates an example of a partitioning tree starting from a 64x64 root block in HEVC. Figure 3 illustrates an example rate distortion optimization for loop process to find splitting or prediction mode for coding units. Figure 4 illustrates an example of top / left neighboring blocks of a coding unit. Figure 5 illustrates neighboring blocks of a current coding unit in VVC, according to a fourth sub-embodiment of a first variant embodiment. Figure 6 illustrates neighboring blocks of a current coding unit in VVC, according to an alternate fourth sub-embodiment of a first variant embodiment Figure 7 illustrates. neighboring blocks of a current coding unit in VVC, according to a fifth sub-embodiment of a first variant embodiment Figure 8 illustrates variant neighborhood block considerations (a) considering middle top / left neighboring block only and (b) considering all the top and left neighboring blocks as neighborhood features. Figure 9 illustrates neighboring blocks of a current coding unit. Figure 10 illustrates downscaling / upscaling converting texture information of various sizes according to an embodiment. Figure 11 illustrates downscaling / upscaling converting texture information of various sizes according to another embodiment. Figure 12 illustrates a histogram of oriented gradients of each subdivision of a current coding unit and its causal borders. Figure 13 illustrates an example of training database collection during encoding. Figure 14 illustrates an example training process where an agent interacts with an environment and loops over episodes. Figure 15 illustrates an example of a flag controlling a model prediction inside an encoder. Figure 16 illustrates an example of a prediction process. Figure 17 illustrates one embodiment of a first method under the described aspects. Figure 18 illustrates one embodiment of a second method under the described aspects Figure 19 illustrates one embodiment of an apparatus under the described aspects. Figure 20 illustrates a standard, generic, video compression scheme. Figure 21 illustrates a standard, generic, video decompression scheme. Figure 22 illustrates a processor-based system for encoding / decoding under the general described aspects. DETAILED DESCRIPTION The embodiments described here are in the field of video compression and generally relate to video compression and video encoding and decoding more specifically to a method or an apparatus for prediction of partitioning modes with reinforcement learning (RL) in hybrid block-based video codecs. To achieve high compression efficiency, image and video coding schemes usually employ block-based prediction, including motion vector prediction, and transform to leverage spatial and temporal redundancy in the video content. Generally, intra or inter prediction is used to exploit the intra or inter frame correlation, then the differences between the original block and the predicted block, often denoted as prediction errors or prediction residuals, are transformed, quantized, and entropy coded. To reconstruct the video, the compressed data are decoded by inverse processes corresponding to the entropy coding, quantization, transform, and prediction. In the HEVC (High Efficiency Video Coding) video compression standard, motion compensated temporal prediction is employed to exploit the redundancy that exists between successive pictures of a video. To do, a motion vector is associated to each prediction unit (PU). Each CTU (Coding Tree Unit) is represented by a Coding Tree in the compressed domain. This is a quad-tree division of the CTU, where each leaf is called a Coding Unit (CU). Each CU is then given some Intra or Inter prediction parameters (Prediction Info). To do so, it is spatially partitioned into one or more Prediction Units (PUs), each PU being assigned some prediction information. The Intra or Inter coding mode is assigned on the CU level. The described embodiments are in the context of video compression. The embodiments specifically focus on prediction of partitioning modes with reinforcement learning (RL) in hybrid block-based video codecs such as H.265 / HEVC, H.266 / VVC, Enhanced Compression Model (ECM) or Essential Video Coding (EVC). In this description, the term “partitioning modes” refers to different ways the current block can be divided into sub-blocks. For instance, in H.266, there exist 6 partitioning modes : quadtree (QT) split divides the current block into four equal-sized non-overlapping sub- blocks, binary tree split divides the current block into two equal-sized non-overlapping sub-blocks horizontally (BTH) or vertically (BTV), ternary tree split divides the current block into three non-overlapping sub-blocks, one block being twice larger than the other two in horizontal (TTH) or vertical (TTV) direction, “no split” (NS) does not split the current block. In this description, for a given frame, the term “channel partitioning” refers to the final decomposition of this channel into blocks resulting from hierarchically dividing it into sub-blocks via partitioning modes selected during the encoding process. As illustration, Figure 1 depicts an example of partitioning the luminance channel and two chrominance channels of a given frame via H.266, version VTM-10.0. The encoding process aims at finding the optimal split mode for each Coding Block (CB) through Rate-Distortion-Optimization (RDO). By trying all the possibilities, the encoder chooses the split that leads to a minimum cost. For a given original CB and a given split mode, the cost is the weighted sum of the distortion between this original CB and its reconstruction arising from the encoding via this split mode, and the bitrate for encoding this original CB via this split mode. Consequently, this process substantially prolongs the total encoding time. However, by integrating neural networks (NNs) into block-based video codecs to forecast appropriate splitting modes for each CB, the partitioning process can be accelerated. For instance, the NN-based split prediction can work as follows. During the encoding step, for each luminance / chrominance CB, instead of letting the encoder find the appropriate split mode, a trained NN directly infers a suitable splitting mode to be applied to this CB from neighbor’s block information, spatial representation of this block, and other relevant features. The goal is to improve the compression efficiency, that is, to reduce the bitrate while maintaining the quality, or equivalently to improve the quality while maintaining the bitrate. In the described embodiments, block partitioning process is investigated in which splitting mode can be predicted via a reinforcement learning agent model that is independent of coding block size, prediction mode (such as intra / inter, for example), and channel component. In the realm of video codecs such as HEVC, VVC, and others, several techniques have been put forward to enhance the block partitioning process, employing deep learning and even deep reinforcement learning. These approaches aim to address the challenge of accelerating the splitting task to minimize encoder complexity. State-of-the-art methods have emerged to tackle this issue effectively and efficiently, focusing on strategies that expedite the splitting process while striving to maintain a high level of quality. These advancements in deep learning and reinforcement learning have paved the way for significant improvements in video codec performance and computational efficiency. Reinforcement learning based partitioning mode prediction model per depth In HEVC, a first approach tackles the challenge of coding block splitting using a deep reinforcement learning approach, the methodology involves training three Deep Q-learning Networks (DQN) models, one for each depth (only depths 0, 1 and 2) as shown in Figure 2 to anticipate the compression benefits associated with splitting a coding block. The output of the model is then utilized by the encoder to guide the decision-making process. The training process for the three Q-networks is carried out sequentially, starting from 16 ൈ 16 for the Q-networks of depth 2 and progressing to 64 ൈ 64 coding units (CUs) for the Q-network of depth 0. Each model takes as input the current CU block and its corresponding quantization parameter (QP) and uses the already trained Q-network of the next depth (not considered for depth 2) to calculate the accumulated reward regarding the action. Furthermore, the reward assignment for non-split action always results in a reward of zero, and it leads to a terminal state in the subsequent state. In HEVC, quad split is the only mode supported, as a result, each of the Q-networks only requires a single output, providing an estimation of the expected rate-distortion cost reduction for a split action. If the estimated value is positive, a split action is taken, otherwise, the current CU remains intact. In order to reduce the coding complexity in VVC standard, a second approach introduces a fast method for Coding Unit splitting based on Deep Reinforcement Learning. The method treats the 32 ൈ 32 CU only, where the CU splitting scenarios at a specific node represent the states, the selection of splitting modes represents the actions, the changes in Rate-Distortion cost serve as immediate rewards or penalties and the encoder acts as an agent that sequentially makes coding decisions. The model receives CU size of 32 ൈ 32 and its corresponding QP as inputs and outputs a Q-value (accumulated reward if the action is selected) for each of the six possible splits (NS, QT, BTH, BTV, TTH, TTV), then the split of the maximum Q-value is selected for the current input. The mentioned methods indeed demonstrate their capability to accelerate encoding and can enhance the partitioning process. However, a notable constraint is that for each block size, a new model is required to handle it. In other words, these approaches heavily rely on the specific block size being considered. Hierarchical deep learning approach for partitioning mode prediction The problem of speeding up the partitioning is also tackled in VP9, in which a third approach proposes a hierarchical fully convolutional network (H-FCN), and the partition tree is represented using four matrices ( ^^^, ^^^, ^^ଶ, ^^ଷ^ that correspond to the four depths of the VP9 partition tree. Each matrix element represents the type of merge of a group of four blocks at its corresponding location and level within the block 64x64. “Fully merging four blocks” means that there is no split, no merge corresponds to quad split, horizontal merge and vertical merge correspond to horizontal split and vertical split respectively of the four blocks of interest. For example, ^^^represents the merges of four 4 ൈ 4 blocks, ^^^represents merges of non-overlapping groups of four 8 ൈ 8 blocks, ^^ଶmerges of non- overlapping groups of four 16 ൈ 16 blocks and ^^ଷrepresents merges of non- overlapping groups of four 32 ൈ 32 blocks. The H-FCN has a main and four outputs branches that stem from the principal branch, the model takes the 64 ൈ 64 block as input and each branch outputs the merging choice for each of the four depths. A fourth approach proposes a multi-stage exit CNN model with an early exit mechanism to accelerate the encoding process in intra mode VVC configuration. Inthis approach the process of partitioning a 128 ൈ 128 Coding Tree Unit (CTU) into64 ൈ 64 Coding Units (CUs) can be identified as Stage 1. Similarly, the subsequentsubdivision of these 64 ൈ 64 CUs into 32 ൈ 32 CUs can be considered as Stage 2, andso forth. In the default setting of the intra-mode Versatile Video Coding (VVC), all128 ൈ 128 CTUs are mandated to undergo splitting into 64 ൈ 64 CUs, resulting in the support for quad-tree mode exclusively in Stage 1. As for Stage 2, both non-splitting and quad-tree modes are accommodated. Subsequent stages provide the possibility of up to six modes, including non-splitting, quad-tree, horizontal binary-tree, vertical binary-tree, horizontal ternary-tree, and vertical ternary-tree. It is ensured that the minimum width or height of CUs remains at 4 for all these modes. In addition to the computational complexity, these approaches suffer from the problem of consistency and prediction can be contradictory due to the limitations of capturing global or long-range contextual information and this can impact the understanding of the relationship between sub-blocks within the input CU and, they are size independent approaches. The encoding process is a combinatorial problem as shown in Figure 3. In VVC, determining the optimal CU partitioning involves an exhaustive rate-distortion optimization (RDO) search, where the RD cost of all potential CUs is evaluated, and the combination with the lowest RD cost is selected. For each tested mode, a cost ^^ is calculated by adding the distortion ^^ between the original and the decoded pixels and the necessary bitrate ^^ and balance the two terms by ^^ as shown in equation (1). min ^^ ൌ ^^ ^ ^^ ^^ (1)Partitioning process is the head of the combinatorial tree where it recursively loops over the CUs to determine the optimal split, which is time consuming. By inserting NNs, this process can be accelerated. As discussed in the previous sections, this problem already exists and is treated by researchers via different approaches. However, all the methods in the literature are size dependent, and each of the proposed models takes inputs of a different size. In this case, CU with similar textures and different sizes cannot be processed via a unified robust neural network. Alternatively, in these embodiments, the solution can cover all CU sizes by making a representation of all the CUs considering features that are considered by the encoder when encoding a given CU. Hereafter, a generic reinforcement learning, size independent method is presented in which all CUs are treated by the proposed framework. The key to this method is to create a representation of all the CUs. Representation of the CUs based on features sets Since a neural network takes an input of fixed size and these embodiments aim at using a single neural network for CUs of any sizes, in these embodiments, the representation of a CU fed into a neural network must be independent of the CU size. Thus, the representation of a CU may be projected into a dimension where the representation of all CUs may be vector ^^ based on several feature sets. In the subsequent sections, the case where the vector ^^ is composed of the four feature sets {“Neighbor features”, “parent features”, “CU information”, “spatial features”} is detailed. ^^ ൌ^Neighbor features, parent features, CU information, spatial features^Yet, in another variant embodiment, the vector ^^ may include only a subset of the four feature sets {“Neighbor features”, “parent features”, “CU information”, “spatial features”}. For instance, the vector ^^ may be composed of “Neighbor features” and “Parent features” exclusively. As another example, the vector ^^ may be composed of “parent features”, “CU information”, and “spatial features” exclusively. In another variant embodiment, the vector ^^ may include, among other feature sets, the four feature sets {“Neighbor features”, “parent features”, “CU information”, “spatial features”}. For instance, let us consider that, in the hybrid block-based video codec of interest, “no split” belongs to the set of possible splits of a given CU. Let us also consider that this hybrid block-based video codec includes a NN taking the vector ^^ of a given CU to infer the most probable splits of this CU, excluding “no split”. For a given CU, this NN is run after testing “no split” and before testing any other split. In this case, the vector ^^ may include {“Neighbor features”, “parent features”, “CU information”, “spatial features”} and features characterizing the intra / inter prediction mode selected to predict the current CU when testing “no split”. Another example can be derived from this case. The vector ^^ may include {“Neighbor features”, “parent features”, “CU information”, “spatial features”}, features characterizing the intra / inter prediction mode selected to predict the current CU when testing “no split”, and a feature representing the number of bits needed to write to the bitstream the quantized transform coefficients of the current CU when testing “no split”. Neighborhood features (NF) for partitioning structure understanding During the encoding of a specific Coding Unit (CU), the encoder may take advantage of the neighboring blocks. These neighboring blocks may correspond to previously processed CUs. Features of the Top and Left neighboring blocks In a first variant embodiment, Top and Left neighboring blocks may be considered. For instance, Figure 4 depicts an example of Top and Left neighboring blocks of the current ^^ ൈ ^^ CU. These neighboring blocks may have already undergone their own individual encoding procedures. Thus, for each of the Top and Left neighboring blocks, valuable data such as rate-distortion cost or quadtree depth (number of quadtree splits needed to move from the root block (CTU) to this block in the partitioning tree) may be used. For instance, the neighboring features included in the vector ^^ may be the rate-distortion cost ^^ ^^ ^^ ^^^୭୮of the Top neighboring block, the rate-distortion cost ^^ ^^ ^^ ^^^^^^of the Left neighboring block, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮of the Top neighboring block, and the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^of theLeft neighboring block. In this case, ^^ ^^ ൌ^ ^^ ^^ ^^ ^^^୭୮, ^^ ^^ ^^ ^^^^^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^^.Normalization of the rate-distortion cost per pixel and by the “no split” cost In a first sub-embodiment of the first variant embodiment, the rate-distortion cost of each neighboring block may be normalized per pixel and by the “no split” cost of this neighboring block. For example, following Figure 4, the rate-distortion cost of the Top neighboring block may be as follows. ^^ ^^ ^^ ^^^^^௧^୭୮ൌ^^౦^ೈ ^.^ೈ ^.^^^௧ ಿೄ(2) ^^౦ where ^^ ^^ ^^ ^^்^^is the rate-distortion cost of the Top neighboring block of the current CU, W / 2 and W / 2 are the width and the height respectively of the Top neighboring block, ^^ ^^ ^^ ^^^ே୭ௌ୮ is the “no split” cost of the Top neighboring block of the current CU. Normalization of the rate-distortion cost per pixel In a second sub-embodiment of the first variant embodiment, the rate-distortion cost of each neighboring block may be normalized per pixel exclusively. For example, following Figure 4, the rate-distortion cost of the Top neighboring block may be as follows. ^^ ^^ ^^ ^^ ൌ^^^௧^^౦^୭୮ ^ೈ ^.^ೈ ^(3) Normalization of the “no split” cost In a third sub-embodiment of the first variant embodiment, the rate-distortion cost of each neighboring block may be normalized by the “no split” cost of this neighboring block exclusively. For example, following Figure 4, the rate-distortion cost of the Top neighboring block may be as follows. ^^ ^^ ^^ ^^^୭୮ൌ^^^௧^^౦^^^௧ ಿೄ^^౦(4) Rate-distortion cost a leaf of the current partitioning tree In a fourth sub-embodiment of the first variant embodiment, given the current state of the partitioning of the current frame when running the RDO on the current CU, the Top neighboring block may correspond to the “leaf” CU containing the pixel above the pixel at the top-left of the current CU. Then, the rate-distortion cost of the Top neighboring block may correspond to its “no split” cost. The Left neighboring block may correspond to the “leaf” CU containing the pixel on the left side of the pixel at the top-left of the current CU. Then, the rate-distortion cost of the Left neighboring block may correspond to its “no split” cost. Note that, at the current state of the partitioning of the current frame, as the Top / Left neighboring block corresponds to a leaf of the partitioning tree, its “no split” cost is always its minimum cost among the tested splits. An example of the fourth sub-embodiment of the first variant embodiment is depicted in Figure 5. Note that, in this fourth sub-embodiment of the first variant embodiment, any alternative criterion may locate the Top / Left neighboring block of the current CU. For instance, the Top neighboring block may correspond to the “leaf” CU containing the pixel above the pixel at the center of the first row of the current CU. The Left neighboring block may correspond to the “leaf” CU containing the pixel on the left side of the pixel at the center of the first column of the current CU. An example of this alternative fourth sub-embodiment is presented in Figure 6. Rate-distortion cost of a neighboring block, in the case where the current CU and this neighboring block result from the same split In a fifth sub-embodiment of the first variant embodiment, if the current CU and its above area result from the same split, as shown in Figure 7 (a), the Top neighboring block may correspond to the CU complying with the conditions (i) and (ii). (i) It is the child of index ^^ resulting from this split, where ^^ ∈^0, ^^ െ 1^, the current CU being the child of index ^^ resulting from this split. (ii) It contains the pixel above the pixel at the top-left of the current CU. If the current CU and its left area result from the same split, as illustrated in Figure 7 (b), the Left neighboring block may correspond to the CU complying with the conditions (i) and (iii). (iii) It contains the pixel on the left side of the pixel at the top-left of the current CU. Note that, in this fifth sub-embodiment of the first variant embodiment, (ii) and (iii) may be adapted in any manner. For instance, (ii) may become “It contains the pixel above the pixel at the center of the first row of the current CU.” Also, (iii) may become “It contains the pixel on the left side of the pixel at the center of the first column of the current CU”. In this fifth sub-embodiment of the first variant embodiment, the rate-distortion cost of the Top neighboring block may correspond to the cost of its selected split. The rate-distortion cost of the Left neighboring block may correspond to the cost of its selected split. Adding the MTT depth of each neighboring block to the neighboring features In a sixth sub-embodiment of the first variant embodiment, the MTT depth of the Top neighboring block and the MTT depth of the Left neighboring block may be part of the neighboring features for the current CU. For instance, the neighboring features included in the vector ^^ may be ^^ ^^ ^^ ^^^୭୮, ^^ ^^ ^^ ^^^^^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^, the MTT depth ^^ ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮of the Top neighboring block, and the MTT depth ^^ ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^ of the Left neighboring block. Then, ^^ ^^ ൌ^^^ ^^ ^^ ^^^୭୮, ^^ ^^ ^^ ^^^^^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^, ^^ ^^ ^^ ^^ ^^ ^^ ^^ℎ^୭୮, ^^ ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^^^. Features of multiple Top neighboring blocks and multiple Left neighboring blocks In a second variant embodiment, multiple Top neighboring blocks and multiple Left neighboring blocks may be considered. For instance, Figure 8 presents an example involving multiple Top and Left neighboring blocks of the current ^^ ൈ ^^ CU. These neighboring blocks may have already undergone their own individual encoding procedures. Thus, for each of these neighboring blocks, valuable data such as rate- distortion cost or quadtree depth may be used. In Figure 8 (a), the neighboring features included in the vector ^^ may be the rate-distortion cost ^^ ^^ ^^ ^^ெ்^^^ௗௗ^^of the middle Top neighboring block, the rate-distortion cost ^^ ^^ ^^ ^^^ெ^^^ௗ௧ௗ^^of the middle Left neighboring block, the quadtree depth ^ௗ^^of the middle Top neighboring block, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎெ^ௗௗ^^ ^^^௧of the middle Left neighboring block. In this case, ^^ ^^ ൌ^^^ ^^ ^^ ^^ெ^ௗௗ^^ ெ^ௗௗ^^ ெ^்^^ , ^^ ^^ ^^ ^^^^^௧ , ^^ ^^ ^^ ^^ ^^ ^^ℎ ௗௗ^^ ெ^ௗௗ^^்^^ , ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ ^.of index 1 may correspond to the “leaf” CU containing the pixel above the pixel at the top-left of the current CU. Then, the rate-distortion cost of the Top neighboring block of index 1 may correspond to its “no split” cost. The Top neighboring block of index 4 may correspond to the “leaf” CU containing the pixel above the pixel at the top-right of the current CU. Then, the rate- distortion cost of the Top neighboring block of index 4 may correspond to its “no split” cost. The Left neighboring block of index 1 may correspond to the “leaf” CU containing the pixel on the left side of the pixel at the top-left of the current CU. Then, the rate- distortion cost of the Left neighboring block of index 1 may correspond to its “no split” cost. The Left neighboring block of index 4 may correspond to the “leaf” CU containing the pixel on the left side of the pixel at the bottom-left of the current CU. Then, the rate- distortion cost of the Left neighboring block of index 4 may correspond to its “no split” cost. In this case, ^^ ^^ ൌ^^^ ^^ ^^ ^^்^^^, ^^ ^^ ^^ ^^்^^ସ, ^^ ^^ ^^ ^^^^^௧^, ^^ ^^ ^^ ^^^^^௧ସ, ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^ସ, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ସ^Considering features of the entire row and column Top / Left neighboring blocks of the current CU In a first sub-embodiment of the second variant embodiment, neighboring features included in the vector ^^ may be all the Top blocks and Left blocks, as shown in Figure 8 (b), including the rate-distortion cost ^^ ^^ ^^ ^^்^^^of the Top neighboring block of index 1, the rate-distortion cost ^^ ^^ ^^ ^^்^^ଶof the Top neighboring block of index 2, the rate-distortion cost ^^ ^^ ^^ ^^்^^ଷof the Top neighboring block of index 3, the rate- distortion cost ^^ ^^ ^^ ^^்^^ସof the Top neighboring block of index 4, the rate-distortion cost ^^ ^^ ^^ ^^^^^௧^of the Left neighboring block of index 1, the rate-distortion cost ^^ ^^ ^^ ^^^^^௧ଶof the Left neighboring block of index 2, the rate-distortion cost ^^ ^^ ^^ ^^^^^௧ଷof the Left neighboring block of index 3, the rate-distortion cost ^^ ^^ ^^ ^^^^^௧ସof the Left neighboring block of index 4, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^^of the Top neighboring block of index 1, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^ଶof the Top neighboring block of index 2, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^ଷof the Top neighboring block of index 3, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^ସof the Top neighboring block of index 4, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧^of the Left neighboring block of index 1, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ଶof the Left neighboring block of index 2, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ଷof the Left neighboring block of index 3, and the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ସof the Left neighboring block of index 4. In this case, ^^ ^^ ൌ ^ ^^ ^^ ^^ ^^்^^^, ^^ ^^ ^^ ^^்^^ଶ, ^^ ^^ ^^ ^^்^^ଷ, ^^ ^^ ^^ ^^்^^ସ, ^^ ^^ ^^ ^^^^^௧^, ^^ ^^ ^^ ^^^^^௧ଶ, ^^ ^^ ^^ ^^^^^௧ଷ, ^^ ^^ ^^ ^^^^^௧ସ, ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^ ^^௧ସFeatures of neighboring blocks, beyond Top / Left neighboring blocks In previous sections, it is assumed that, as in H.265, H.266, and ECM, the coding order of the CTUs and that of the CUs goes from left-to-right and top-to-bottom. In a hybrid block-based video codec with change of coding order, depending on the coding parametrization, the current CU may be coded from either left-to-right or right- to-left and either top-to-bottom or bottom-to-top. In this case, in a third variant embodiment, different neighboring blocks may be considered. For instance, Top neighboring block, Left neighboring block, Right Neighboring block, and Bottom neighboring block may be considered. Figure 9 shows an example of Top neighboring block, Left neighboring block, Right neighboring block, and Bottom neighboring block of the current ^^ ൈ ^^ CU. For each of these neighboring blocks, valuable data such as rate-distortion cost or quadtree depth may be used. For instance, following Figure 9, the neighboring features included in the vector ^^ may be the rate-distortion cost ^^ ^^ ^^ ^^்^^of the Top neighboring block, the rate-distortion cost ^^ ^^ ^^ ^^^^^௧of the Left neighboring block, the rate-distortion cost ^^ ^^ ^^ ^^ோ^^^௧of the Right neighboring block, the rate-distortion cost ^^ ^^ ^^ ^^^^௧௧^^of the Bottom neighboring, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^of the Top neighboring block, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧of the Left neighboring block, the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎோ^^^௧of the Right neighboring, and the quadtree depth ^^ ^^ ^^ ^^ ^^ ^^ℎ^^௧௧^^of the Bottom neighboring block. In this case, ^^ ^^ ൌ^^^ ^^ ^^ ^^்^^, ^^ ^^ ^^ ^^^^^௧, ^^ ^^ ^^ ^^ோ^^^௧, ^^ ^^ ^^ ^^^^௧௧^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧ , ^^ ^^ ^^ ^^ ^^ ^^ℎோ^^^௧ , ^^ ^^ ^^ ^^ ^^ ^^ℎ^^௧௧^^^ Among these four neighboring blocks, if a neighboring block does not exist as it is out of the bounds of the current frame or it has not been encoded yet, its associated data may be set to default values. For instance, for the current CU, if its Right neighboring block does not exist, ^^ ^^ ^^ ^^ோ^^^௧may be set to the maximum value of a 64- bit unsigned integer, i.e.2^ସെ 1. ^^ ^^ ^^ ^^ ^^ ^^ℎோ^^^௧may be set to -1. As another example, for the current CU, if its Right neighboring block does not exist, ^^ ^^ ^^ ^^ோ^^^௧and ^^ ^^ ^^ ^^ ^^ ^^ℎோ^^^௧may be set to -2. Any of the sub-embodiments of the first variant embodiment can be straightforwardly adapted to the third variant embodiment. In particular, the fourth sub- embodiment of the first variant embodiment can be turned into the sub-embodiment in the following section. Rate-distortion cost of a neighboring block being a leaf of the current partitioning tree In a sub-embodiment of the third variant embodiment, the example in Figure 9 is reused. Given the current state of the partitioning of the current frame when running the RDO on the current CU, the Top neighboring block may correspond to the “leaf” CU containing the pixel above the pixel at the top-left of the current CU. Then, the rate- distortion cost of the Top neighboring block may correspond to its “no split” cost. The Left neighboring block may correspond to the “leaf” CU containing the pixel on the left side the pixel at the top-left of the current CU. Then, the rate-distortion cost of the Left neighboring block may correspond to its “no split” cost. The Right neighboring block may correspond to the “leaf” CU containing the pixel on the right side of the pixel at the top-right of the current CU. Then, the rate-distortion cost of the Right neighboring block of may correspond to its “no split” cost. The Bottom neighboring block may correspond to the “leaf” CU containing the pixel at the bottom of the pixel at the bottom- left of the current CU. Then, the rate-distortion cost of the Bottom neighboring block may correspond to its “no split” cost. Note that, in this sub-embodiment, any alternative criterion may locate the Top / Left / Right / Bottom neighboring block of the current CU. Rate-distortion as two feature values instead of the cost of Top / Left neighboring blocks In a fourth variant embodiment, based on equation (1), cost values of the neighboring blocks may be considered as two feature values including rate ^^ and distortion ^^. In this embodiment, any of the sub-embodiments of the first and second variants may be straightforwardly adapted. As an example, the case illustrated in Figure 8 (a) may be adapted as follows. The vector ^^ may contain the rate ^^்^^of the middle Top neighboring block, the distortion ^^்^^of the middle Top neighboring block, the rate ^^^^^௧of the middle Left neighboring block, the distortion ^^^^^௧of the middleLeft neighboring block, in this case, ^^ ^^ ൌ^ ^^்^^, ^^்^^, ^^^^^௧ , ^^^^^௧, ^^ ^^ ^^ ^^ ^^ ^^ℎ்^^, ^^ ^^ ^^ ^^ ^^ ^^ℎ^^^௧^.Note that any of the sub-embodiments can be straightforwardly adapted to this sub- embodiment. Parent Costs (PC) During the application of Rate-Distortion Optimization (RDO) on a Coding Unit (CU), the process generates sub-CUs. This implies that each CU owns a parent CU. To maintain a comprehensive understanding of the overall partitioning structure, it may become crucial to incorporate the costs of split modes evaluated on the parent CU as informative data for the current child CU. For instance, in H.266, the PC feature sets may be ^^ ^^ ൌ^^^ ^^ ^^ ^^ேௌ, ^^ ^^ ^^ ^^ொ், ^^ ^^ ^^ ^^^்ு, ^^ ^^ ^^ ^^^்^, ^^ ^^ ^^ ^^்்ு, ^^ ^^ ^^ ^^்்^^. Each time a split mode is applied to the parent, its associated cost normalized by the no split cost of the parent may be filled in the list of parent costs for the next state (CU) and assigning an arbitrary value to the remaining costs that have not been tested yet. Block Information of the current CU (BI) The block information may include, amongst others, the block size, QP value, channel component. This allows for distinguishing between CUs of varying sizes, channels, prediction type and the frame type. ^^ ^^ ൌ ^ ^^, ^^, ^^ ^^, ^^ℎ ^^ ^^ ^^ ^^ ^^, ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^, ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^^ For instance, in H.266, the characteristics of BI may be ^ ^^, ^^ ∈ ^64,32, 16, 8, 4^ ^ ^^ ^^ ∈ ℕ ^^ ^^ ^^ ∈ ^0 െ 51^ 0^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^^ ^^ℎ ^^ ^^ ^^ ^^ ^^ ൌ ^1 ^^ℎ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^2 ^^ℎ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^^^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ^^ ൌ ^0 ^^ ^^ ^^ ^^ ^^1^^ ^^ ^^ ^^ ^^^^ ^^ ^^ ^^^^ ^^ ^^ ^^^^ ^^ ^^ ^^current CU belonging to an intra slice, BI includes the channel component whereas, for the current CU belonging to an inter slice, BI does not include the channel component. In another variant embodiment, for the current CU, BI always includes the channel component. Spatial Features representation of all CUs (SF) Downscaling / upscaling converting texture information of various sizes into fixed-size input A target image size, e.g.12 ൈ 12, may be chosen. A CU of any size, along with its causal borders, may be downscaled / upscaled to this target image size. A following section on training database collection describes an example of database collection where, for each CU, four pixel-wide causal borders are extracted with the actual CU. Yet, note that a different width of causal borders may be used, not affecting the principle being explained. Also, note that the causal borders of a CU may have a different position with respect to this CU, not affecting the principle being explained. For instance, the target image size may be 12 ൈ 12 and the width of causal borders may be 4. In this case, the CUs with both height and width strictly larger than 8, along with their causal borders, may downscaled to 12 ൈ 12. In this case, the CUs of size 8 ൈ 8, along with their causal borders, may be neither downscaled nor upscaled. In this case, the CUs with both height and width strictly smaller than 8, along with their causal borders, may be upscaled to 12 ൈ 12, see Figure 10. In this case, the CUs with one dimension ^^^∈^height, width^strictly larger than 8, another dimension ^^^∈ ^height, width^ strictly smaller than 8, ^^^് ^^^, along with their causal borders, may be downscaled along ^^^and upscaled along ^^^to reach 12 ൈ 12. In this case, the CUs with one dimension ^^^∈ ^height, width^ strictly larger than 8, another dimension ^^^∈ ^height, width^ equal to 8, ^^^് ^^^, along with their causal borders, may be downscaled along ^^^exclusively to reach 12 ൈ 12. In this case, the CUs with one dimension ^^^∈^height, width^strictly smaller than 8, another dimension ^^^∈^height, width^equal to 8, ^^^് ^^^, along with their causal borders, may be upscaled along ^^^exclusively to reach 12 ൈ 12. For example, the target image size may be 18 ൈ 18 and the width of causalborders may be 2. In this case, the CUs with both height and width strictly larger than16, along with their causal borders, may downscaled to 18 ൈ 18. In this case, the CUsof size 16 ൈ 16, along with their causal borders, may be neither downscaled nor upscaled. In this case, the CUs with both height and width strictly smaller than 16, along with their causal borders, may be upscaled to 18 ൈ 18. In this case, the CUs with one dimension ^^^∈ ^height, width^ strictly larger than 16, another dimension ^^^∈ ^height, width^ strictly smaller than 16, ^^^് ^^^, along with their causal borders, may be downscaled along ^^^and upscaled along ^^^to reach 18 ൈ 18. In this case, the CUs with one dimension ^^^∈^height, width^strictly larger than 16, another dimension ^^^∈ ^height, width^ equal to 16, ^^^് ^^^, along with their causal borders, may be downscaled along ^^^exclusively to reach 18 ൈ 18. In this case, the CUs with one dimension ^^^∈ ^height, width^ strictly smaller than 16, another dimension ^^^∈ ^height, width^ equal to 16, ^^^് ^^^, along with their causal borders, may be upscaled along ^^^exclusively to reach 18 ൈ 18. In a variant embodiment, to downscale / upscale a given CU, along with its causal borders, the nearest-neighbor method may be used. The nearest-neighbor method assigns the value of the nearest pixel to the new position after resizing the input image, which makes it negligible in terms of computational complexity. In another variant embodiment, to downscale / upscale a given CU, along with its causal borders, a linear interpolation may be used. Downscaled / upscaled image directly used as spatial features representation In a variant embodiment, the image resulting from the optional downscaling / upscaling of a given CU with its causal borders may be directly used as spatial features representation. For instance, in Figure 10, this may mean that the pixels in the 12 ൈ 12 image are put into the spatial feature vector ^^ௌி. In another variant embodiment, in addition to the resulting from the optional downscaling / upscaling of a given CU, the spatial features vector ^^ௌிmay include any other spatial feature. Downscaled / upscaled image undergoing a processing before being put into the spatial features vector In a variant embodiment, the image resulting from the optional downscaling / upscaling of a given CU with its causal borders may be processed before being put into the spatial features vector. For instance, a convolutional neural network may be fed with the image resulting from the optional downscaling / upscaling of a given CU with its causal borders, and the neural network output may be put into the spatial features vector, see Figure 11. Traditional feature extraction converting texture information of various sizes into fixed-size input For a CU of any size, instead of using the downscaling / upscaling to convert texture information into fixed-size input, traditional features may be extracted from this CU, along with its causal borders, to construct a spatial features vector of fixed size. For instance, Gray-Level Co-occurrence Matrix (GLCM) and Histogram of Oriented Gradients (HOG) may be picked as traditional features. Gray-Level Co-occurrence Matrix (GLCM) features as part of the spatial features (SF) For a given component ^^ (either luminance, blue chrominance, or red chrominance) of a given CU, the GLCM ^^ may have size 2^ൈ 2^, ^^ denoting the internal bit depth of the codec of interest. For instance, ^^ ൌ 10. As another example,^^ ൌ 8. ^^ may be defined at a given distance range ∈ ^^୰ୟ୬^^ and a given direction ∈^^^୧୰^ୡ^୧୭୬. For instance, ^^୰ୟ୬^^ ൌ ^1, 2, 3^. As another example, ^^୰ୟ୬^^ ൌ ^1, 2, 4^. Forinstance, ^^^୧୰^ୡ^୧୭୬∈ ^0,గସ ,గଶ ,ଷగସ ^. As another example, ^^^୧୰^ୡ^୧୭୬∈ ^0,గଶ^. For a given distance range value and a given direction value, ^^^ ^^, ^^^ may be number of times the pair occurs at the defined spatial relationship in ^^. ^^^ ^^, ^^^ ൌ ∑^௫ୀି^^∑ு௬ୀି^^^1, if ^^^ ^^, ^^^ ൌ ^^ and ^^^ ^^ ^ ^^ ^^, ^^ ^ ^^ ^^^ ൌ ^^0, otherwise (6)the spatial offset between the current pair of pixels. For instance, if the distance range value is 1 and the direction value isగସ, ^^ ^^ ൌ 1 and ^^ ^^ ൌ 1. As another example, if the distance range value is 1 and the direction value isଷగସ , ^^ ^^ ൌ െ1 and ^^ ^^ ൌ 1. Once ^^ is may be derived from ^^ to form a vector ^^^^ ^^ ^^ ^^. The spatial features vector ^^ௌிmay include ^^^^ ^^ ^^ ^^. Contrast, energy, homogeneity, correlation, and dissimilarity derived from the GLCM In a variant embodiment, for a given component of a given CU, for a given distance range value and a given direction value, the contrast, energy, homogeneity, correlation, and dissimilarity may be derived from the computed ^^. For instance, the formula of each of them may be as follows. contrast ൌ ∑^,^ ^ ^^ െ ^^^ଶ ∙ ^^^ ^^, ^^^ (7)^^^ ^^^ଶ (8)(9)(10)^^^ (11)^^௫and ^^௬are the means of respectively of ^^. ^^௫and ^^௬are their respective standard deviations. ^^^^ ^^ ^^ ^^ൌ ^contrast, energy, homogenity, correlation, dissimilarity^. Subset of {contrast, energy, homogeneity, correlation, dissimilarity} In another variant embodiment, for a given component of a given CU, for a given distance range value and a given direction value, ^^^^ ^^ ^^ ^^may include any subset of ^contrast, energy, homogenity, correlation, dissimilarity^. For instance, ^^^^ ^^ ^^ ^^ൌ ^contrast, energy, homogenity, dissimilarity^. As another example, ^^^^ ^^ ^^ ^^ൌ ^contrast, homogenity, dissimilarity^. {contrast, energy, homogeneity, correlation, dissimilarity} along with any additional statistical measures In another variant embodiment, for a given component of a given CU, for a given distance range value and a given direction value, ^^^^ ^^ ^^ ^^may include, amongst other statistical measures,^contrast, energy, homogenity, correlation, dissimilarity^. Concatenation of the statistical measures associated to each CU component In a variant embodiment, for each component compID belonging to {Y, ^^ୠ, ^^୰} of a given CU, for a given distance range value and a given direction value, the GLCM ^^ୡ୭୫୮୍ୈmay be computed, and a set of statistical measures may be derived from ^^ୡ୭୫୮୍ୈ. For instance, if the set of statistical measures is^contrast, energy, homogeneity, correlation, dissimilarity^, for a given CU, for a given distance range value and a given direction value, ^^^^ ^^ ^^ ^^ൌ^contrastଢ଼, energyଢ଼, homogenityଢ଼, correlationଢ଼, dissimilarityଢ଼,contrast^^ ^^, energyେ^, homogenityେ^, correlationେ^, dissimilarityେ^,contrast ^^ ^^ , energyେ౨ , homogenityେ౨ , correlationେ౨ , dissimilarityେ౨^.As another example, if the set of statistical measures is ^contrast, correlation^, for a given CU, for a given distance range value and a given direction value, ^^^^ ^^ ^^ ^^ൌ ^contrastଢ଼, correlationଢ଼, contrast^^ ^^, correlationେ^, contrast^^ ^^, correlationେ౨^. Multiple values of distance range and direction In a variant embodiment, for a given component of a given CU, for each pair ^^^of a given distance range value and a given direction value, the GLCM ^^^ೖmay be computed, and a set of statistical measures may be derived from ^^^ೖ. For instance, if the set of statistical measures is ^contrast, energy, homogeneity, correlation, dissimilarity^, for a given component of a given CU, using pair ^^^ൌ ^1,గସ^ and pair ^^^ൌ ^1,ଷగସ ^, ^^^బand ^^^భmay be computed. ,contrast^భ , energy^భ , homogenity^భ , correlation^భ , dissimilarity^భ^.As another example, if the set of statistical measures is^contrast, correlation^, for a given component of a given CU, using pair ^^గ ଷగ^ ൌ ^1,ସ^ and pair ^^^ൌ ^1,ସ^, ^^^బ^^ ൌ ^^ ^^ ^^ ^^^contrast^బ, correlation^బ, contrast^భ, correlation^భ^. Different definition of the statistical measures In a variant embodiment, the already defined statistical measures may take on different expressions. For instance, homogenity ൌ∑ ெ^^,^^^,^ ^ା^^ି^^మ. Any combination of embodiments Different embodiments in the preceding sections can be straightforwardly combined. Histogram of Oriented Gradients (HOG) as part of the spatial features (SF) For a given CU, the HOG of this CU may be calculated and put into a vector ^^^^ ^^ ^^. The spatial features vector ^^ௌிmay include ^^^^ ^^ ^^. HOG of a given CU and its causal borders In a variant embodiment, for a given CU, the HOG of this CU and its causal borders may be calculated and put into ^^^^ ^^ ^^. HOG of a given CU and HOG of its causal borders In a variant embodiment, for a given CU, the HOG of this CU and the HOG of its causal borders may be calculated, and the concatenation of these two HOGs may be put into ^^^^ ^^ ^^. HOG of each subdivision of a given CU and its causal borders In a variant embodiment, a given CU and its causal borders may be divided into portions. Any division into portions may apply. Then, for each of these portions, a HOG may be calculated. Finally, the concatenation of the resulting HOGs may be put into ^^^^ ^^ ^^. For instance, let us say the causal border, a.k.a “Top border”, comprising 4 rows of pixels above the current ^^ ൈ ^^ CU and the causal border, a.k.a “Left border” comprising 4 columns of pixels on the left side of the current ^^ ൈ ^^ CU are considered. Let us also say that the current CU and its causal borders are divided into portions as: {“Top border”, “Left border”, 4 non-overlapping sub-blocks of the current CU with same size}, see Figure 12. Then, a HOG, denoted Top HOG features, may be computed on “Top border”. A HOG, denoted Left HOG features, may be computed on “Left border”. A HOG, denoted HOG features i, may be computed on the non-overlapping sub-blockof index ^^, ^^ ∈ ^0,3^. ^^ ^^ ^^ ^^ ൌ ^Top HOG features,Left HOG features, HOG features 0, HOG features 1, HOG features 2, HOG features 3^.HOG computation Any of the different definitions of HOG may apply here. As an example, all the bins of the HOG may be initialized to 0. Then, for each pixel in a set of pixels (e.g. all pixels; e.g. one pixel out of two) belonging to the region (e.g. a causal border of the current CU; e.g. a sub-block in the current CU) on which HOG is computed, a magnitude ^^ and an angle ^^ may be computed. ^^ ൌ^^^௫ଶ ^ ^^௬ଶ (12)^^ (13) where ^^௫, ^^௬may be of the neighboring pixels at the right and left of the current between the neighboring pixels above and below the current pixel. ^^௫ ൌ | ^^^ ^^, ^^ ^ 1^ െ ^^^ ^^, ^^ െ 1^| (14)ൌ^^^ ^^ െ 1, ^^^ െ ^^^ ^^ ^ 1, (15)^^, ^^ denoting the region ^^ of interest. Then, the magnitude ^^ and the angle ^^ may be used to find the index ^^ of the HOG bin to be incremented. ^^^^ൌ ^^. Finally, after completing the loop over pixels, the resulting HOG may be normalized. As another example, the magnitude may be defined differently. The previous example may be re-used, except that ^^ ൌ | ^^௫| ^ ห ^^௬ห. As another example, ^^௫, ^^௬may be defined differently. The previous example may be re-used, except that ^^௫ൌ | ^^^ ^^, ^^^ െ ^^^ ^^, ^^ െ 1^| and ^^௬ൌ | ^^^ ^^ െ 1, ^^^ െ ^^^ ^^, ^^^|. Note that any variant of the current embodiment may be straightforwardly combined with any of the other embodiments. Spatial features (SF) combining GLCM and HOGs In a variant embodiment, for a given CU and its causal borders, GLCM and HOGs may be put into the associated spatial features vector. For instance, if the spatial features vector of a given CU and its causal borders comprises GLCM and HOGs exclusively, ^^^^ ^^ൌ ^ ^^^^ ^^ ^^ ^^, ^^^^ ^^ ^^^. The training process of the framework Training database collection To form a training database, numerous CUs may be collected via the encodings of many sequences of different resolutions (e.g.4K, HD, SD) using several QPs (e.g. 22, 27, 32, 37). Figure 13 depicts an example of training database collection. When starting the encoding of a CU, first, it may be verified whether its height and width belong to a set of allowed sizes. To form the state vector, the coding block may be accompanied by a ^ ^^ ^ 4, ^^ ^ 4^ patch, encompassing the block itself and four pixel-wide causal borders. Additionally, the associated cost of the current tested split mode may be also extracted for the reward signal later. ^^ ^^, ^^ ^^ and ^^ ^^ features may also be extracted to form the total vector ^^ representing the current state. ^^ ൌ^^^ ^^, ^^ ^^, ^^ ^^, block pixels^. At this stage the block pixels may not be treated yet to calculate the spatial features to form the vector ^^^^ ^^. Reward function incorporating cost and temporal complexity for agent-based action prediction The reward signal may be considered as the accumulated penalties of the cost of the current action ( ^^ ^^ ^^ ^^ୟୡ^୧୭୬^, which is the split mode decision, plus atemporal complexity term ^^^^ times ^^ to balance the two terms.reward^action, CU^ ൌ െ ^^ ^^ ^^ ^^ୟୡ^୧୭୬ െ ^^. ^^^^ (17) Theorical RDO iteration as temporal complexity In a variant embodiment, the temporal complexity may be considered as the theorical number of times that the RDO process is done starting from a given CU. Total number of tested intra / inter modes as temporal complexity In a variant embodiment, when encoding a CU, the number of modes to be tested in different encoding components can be determined beforehand. For instance, these components may include the transformed mode, intra mode, and inter mode.The total number of tested modes may be considered as temporal complexity.reward^action, CU^ ൌ െ ^^ ^^ ^^ ^^ୟୡ^୧୭୬ െ ^^. ^^ே (18)^^ேdenoting the number of tested modes during the encoding process of the CU. Reinforcement learning algorithm adaptation to train neural networks In reinforcement learning, the agent may learn a Q-value function ^^^ ^^^ to make decisions in an environment by interacting with it over multiple iterations. The environment may be defined with its states, actions, and reward signal. In this description, the state may be the feature vector ^^ of the current CU, actions may be the decision of the split mode, and the reward may be the combination of the cost and the temporal complexity as defined in equations (17-18). Figure 14 shows an example of training process where the agent interacts with the environment and loops over the episodes. The episode may represent the total partitioning tree of a CU parent 64 ൈ 64, the agent may take the parent as first state ^^^, and at this stage, pixels of the current CU can be treated to calculate the spatial feature vector and form the entire feature vector ^^. The next state may depend on the current action. Flag characterizing the neural network partitioning mode prediction for inference step The "ActivatePrediction" flag may allow the model to make predictions on the current CU within the RDO loop. If the classical encoder heuristics are activated, the encoder may prepare a list of split modes to be tested based on the heuristics. If the model is activated (ActivatePrediction=1) and the prediction action (output of the model) matches “currTestMode” (current split mode being tested on the current CU), the encoder may test the specific split mode on the CU. On the other hand, if the prediction action does not match “currTestMode”, the mode may be skipped, then the encoder may move on to the next configuration. Figure 15 provides an example of flag controlling the model prediction. Figure 16 gives an example of Prediction Process (PP). If the classical encoder heuristics are deactivated, if the model is activated (ActivatePrediction=1), the RDO process may test all the split modes that must not be skipped according to the agent. One embodiment of a method 1700 under the general aspects described here is shown in Figure 17. The method commences at start block 1701 and control proceeds to block 1710 for performing a partitioning prediction of a video block using a neural network. Control proceeds from block 1710 to block 1720 for partitioning the video block based on said partitioning prediction. Control proceeds from block 1720 to block 1730 for encoding the partitioned video block. One embodiment of a method 1800 under the general aspects described here is shown in Figure 18. The method commences at start block 1801 and control proceeds to block 1810 for performing a partitioning prediction of a video block using a neural network. Control proceeds from block 1810 to block 1820 for partitioning the video block based on said partitioning prediction. Control proceeds from block 1820 to block 1830 for decoding the partitioned video block. Figure 19 shows one embodiment of an apparatus 1900 for encoding, decoding, compressing or decompressing, or filtering of video data using the aforementioned methods. The apparatus comprises Processor 1910 and can be interconnected to a memory 1920 through at least one port. Both Processor 1910 and memory 1920 can also have one or more additional interconnections to external connections. Processor 1910 is also configured to either insert or receive information in a bitstream and, either compressing, encoding, or decoding using any of the described aspects. The embodiments described here include a variety of aspects, including tools, features, embodiments, models, approaches, etc. Many of these aspects are described with specificity and, at least to show the individual characteristics, are often described in a manner that may sound limiting. However, this is for purposes of clarity in description, and does not limit the application or scope of those aspects. Indeed, all of the different aspects can be combined and interchanged to provide further aspects. Moreover, the aspects can be combined and interchanged with aspects described in earlier filings as well. The aspects described and contemplated in this application can be implemented in many different forms. Figures 20, 21, and 22 provide some embodiments, but other embodiments are contemplated and the discussion of Figures 20, 21, and 22 does not limit the breadth of the implementations. At least one of the aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a bitstream generated or encoded. These and other aspects can be implemented as a method, an apparatus, a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or a computer readable storage medium having stored thereon a bitstream generated according to any of the methods described. In the present application, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “pixel” and “sample” may be used interchangeably, the terms “image,” “picture” and “frame” may be used interchangeably. Usually, but not necessarily, the term “reconstructed” is used at the encoder side while “decoded” is used at the decoder side. Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Various methods and other aspects described in this application can be used to modify modules, for example, the intra prediction, entropy coding, and / or decoding modules (160, 260, 145, 230), of a video encoder 100 and decoder 200 as shown in Figure 20 and Figure 21. Moreover, the present aspects are not limited to VVC or HEVC, and can be applied, for example, to other standards and recommendations, whether pre-existing or future-developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise, or technically precluded, the aspects described in this application can be used individually or in combination. Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values. Figure 20 illustrates an encoder 100. Variations of this encoder 100 are contemplated, but the encoder 100 is described below for purposes of clarity without describing all expected variations. Before being encoded, the video sequence may go through pre-encoding processing (101), for example, applying a color transform to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing a remapping of the input picture components in order to get a signal distribution more resilient to compression (for instance using a histogram equalization of one of the color components). Metadata can be associated with the pre-processing and attached to the bitstream. In the encoder 100, a picture is encoded by the encoder elements as described below. The picture to be encoded is partitioned (102) and processed in units of, for example, CUs. Each unit is encoded using, for example, either an intra or inter mode. When a unit is encoded in an intra mode, it performs intra prediction (160). In an inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which one of the intra mode or inter mode to use for encoding the unit, and indicates the intra / inter decision by, for example, a prediction mode flag. Prediction residuals are calculated, for example, by subtracting (110) the predicted block from the original image block. The prediction residuals are then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the non-transformed residual signal. The encoder can bypass both transform and quantization, i.e., the residual is coded directly without the application of the transform or quantization processes. The encoder decodes an encoded block to provide a reference for further predictions. The quantized transform coefficients are de-quantized (140) and inverse transformed (150) to decode prediction residuals. Combining (155) the decoded prediction residuals and the predicted block, an image block is reconstructed. In-loop filters (165) are applied to the reconstructed picture to perform, for example, deblocking / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored at a reference picture buffer (180). Figure 21 illustrates a block diagram of a video decoder 200. In the decoder 200, a bitstream is decoded by the decoder elements as described below. Video decoder 200 generally performs a decoding pass reciprocal to the encoding pass as described in Figure 20. The encoder 100 also generally performs video decoding as part of encoding video data. In particular, the input of the decoder includes a video bitstream, which can be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. The picture partition information indicates how the picture is partitioned. The decoder may therefore divide (235) the picture according to the decoded picture partitioning information. The transform coefficients are de-quantized (240) and inverse transformed (250) to decode the prediction residuals. Combining (255) the decoded prediction residuals and the predicted block, an image block is reconstructed. The predicted block can be obtained (270) from intra prediction (260) or motion- compensated prediction (i.e., inter prediction) (275). In-loop filters (265) are applied to the reconstructed image. The filtered image is stored at a reference picture buffer (280). The decoded picture can further go through post-decoding processing (285), for example, an inverse color transform (e.g. conversion from YcbCr 4:2:0 to RGB 4:4:4) or an inverse remapping performing the inverse of the remapping process performed in the pre-encoding processing (101). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream. Figure 22 illustrates a block diagram of an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 1000, singly or in combination, can be embodied in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 1000 is communicatively coupled to one or more other systems, or other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 1000 is configured to implement one or more of the aspects described in this document. The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this document. Processor 1010 can include embedded memory, input output interface, and various other circuitries as known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device, and / or a non- volatile memory device). System 1000 includes a storage device 1040, which can include non-volatile memory and / or volatile memory, including, but not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 can include an internal storage device, an attached storage device (including detachable and non-detachable storage devices), and / or a network accessible storage device, as non-limiting examples. System 1000 includes an encoder / decoder module 1030 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 1030 can include its own processor and memory. The encoder / decoder module 1030 represents module(s) that can be included in a device to perform the encoding and / or decoding functions. As is known, a device can include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 1030 can be implemented as a separate element of system 1000 or can be incorporated within processor 1010 as a combination of hardware and software as known to those skilled in the art. Program code to be loaded onto processor 1010 or encoder / decoder 1030 to perform the various aspects described in this document can be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. In accordance with various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 can store one or more of various items during the performance of the processes described in this document. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic. In some embodiments, memory inside of the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of, for example, a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also referred to as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard being developed by JVET, the Joint Video Experts Team). The input to the elements of system 1000 can be provided through various input devices as indicated in block 1130. Such input devices include, but are not limited to, (i) a radio frequency (RF) portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Component (COMP) input terminal (or a set of COMP input terminals), (iii) a Universal Serial Bus (USB) input terminal, and / or (iv) a High Definition Multimedia Interface (HDMI) input terminal. Other examples, not shown in Figure 22, include composite video. In various embodiments, the input devices of block 1130 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) downconverting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the downconverted and band- limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs various of these functions, including, for example, downconverting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements can include inserting elements in between existing elements, such as, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna. Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 1000 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing IC or within processor 1010 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface Ics or within processor 1010 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010, and encoder / decoder 1030 operating in combination with the memory and storage elements to process the datastream as necessary for presentation on an output device. Various elements of system 1000 can be provided within an integrated housing, Within the integrated housing, the various elements can be interconnected and transmit data therebetween using suitable connection arrangement, for example, an internal bus as known in the art, including the Inter-IC (I2C) bus, wiring, and printed circuit boards. The system 1000 includes communication interface 1050 that enables communication with other devices via communication channel 1060. The communication interface 1050 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 1060. The communication interface 1050 can include, but is not limited to, a modem or network card and the communication channel 1060 can be implemented, for example, within a wired and / or a wireless medium. Data is streamed, or otherwise provided, to the system 1000, in various embodiments, using a wireless network such as a Wi-Fi network, for example IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi- Fi signal of these embodiments is received over the communications channel 1060 and the communications interface 1050 which are adapted for Wi-Fi communications. The communications channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 1000 using a set-top box that delivers the data over the HDMI connection of the input block 1130. Still other embodiments provide streamed data to the system 1000 using the RF connection of the input block 1130. As indicated above, various embodiments provide data in a non- streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, for example a cellular network or a Bluetooth network. The system 1000 can provide an output signal to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes one or more of, for example, a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be for a television, a tablet, a laptop, a cell phone (mobile phone), or another device. The display 1100 can also be integrated with other components (for example, as in a smart phone), or separate (for example, an external monitor for a laptop). The other peripheral devices 1120 include, in various examples of embodiments, one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disk player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide a function based on the output of the system 1000. For example, a disk player performs the function of playing the output of the system 1000. In various embodiments, control signals are communicated between the system 1000 and the display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communications protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to system 1000 using the communications channel 1060 via the communications interface 1050. The display 1100 and speakers 1110 can be integrated in a single unit with the other components of system 1000 in an electronic device such as, for example, a television. In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip. The display 1100 and speaker 1110 can alternatively be separate from one or more of the other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs. The embodiments can be carried out by computer software implemented by the processor 1010 or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type appropriate to the technical environment and can be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory, as non-limiting examples. The processor 1010 can be of any type appropriate to the technical environment, and can encompass one or more of microprocessors, general purpose computers, special purpose computers, and processors based on a multi-core architecture, as non- limiting examples. Various implementations involve decoding. “Decoding”, as used in this application, can encompass all or part of the processes performed, for example, on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such processes also, or alternatively, include processes performed by a decoder of various implementations described in this application. As further examples, in one embodiment “decoding” refers only to entropy decoding, in another embodiment “decoding” refers only to differential decoding, and in another embodiment “decoding” refers to a combination of entropy decoding and differential decoding. Whether the phrase “decoding process” is intended to refer specifically to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art. Various implementations involve encoding. In an analogous way to the above discussion about “decoding”, “encoding” as used in this application can encompass all or part of the processes performed, for example, on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more of the processes typically performed by an encoder, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also, or alternatively, include processes performed by an encoder of various implementations described in this application. As further examples, in one embodiment “encoding” refers only to entropy encoding, in another embodiment “encoding” refers only to differential encoding, and in another embodiment “encoding” refers to a combination of differential encoding and entropy encoding. Whether the phrase “encoding process” is intended to refer specifically to a subset of operations or generally to the broader encoding process will be clear based on the context of the specific descriptions and is believed to be well understood by those skilled in the art. Note that the syntax elements as used herein are descriptive terms. As such, they do not preclude the use of other syntax element names. When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of a corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of a corresponding method / process. Various embodiments may refer to parametric models or rate distortion optimization. In particular, during the encoding process, the balance or trade-off between the rate and distortion is usually considered, often given the constraints of computational complexity. It can be measured through a Rate Distortion Optimization (RDO) metric, or through Least Mean Square (LMS), Mean of Absolute Errors (MAE), or other such measurements. Rate distortion optimization is usually formulated as minimizing a rate distortion function, which is a weighted sum of the rate and of the distortion. There are different approaches to solve the rate distortion optimization problem. For example, the approaches may be based on an extensive testing of all encoding options, including all considered modes or coding parameters values, with a complete evaluation of their coding cost and related distortion of the reconstructed signal after coding and decoding. Faster approaches may also be used, to save encoding complexity, in particular with computation of an approximated distortion based on the prediction or the prediction residual signal, not the reconstructed one. Mix of these two approaches can also be used, such as by using an approximated distortion for only some of the possible encoding options, and a complete distortion for other encoding options. Other approaches only evaluate a subset of the possible encoding options. More generally, many approaches employ any of a variety of techniques to perform the optimization, but the optimization is not necessarily a complete evaluation of both the coding cost and related distortion. The implementations and aspects described herein can be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed can also be implemented in other forms (for example, an apparatus or program). An apparatus can be implemented in, for example, appropriate hardware, software, and firmware. The methods can be implemented in, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users. Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment. Additionally, this application may refer to “determining” various pieces of information. Determining the information can include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory. Further, this application may refer to “accessing” various pieces of information. Accessing the information can include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information. Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information can include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations such as, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information. It is to be appreciated that the use of any of the following “ / ”, “and / or”, and “at least one of”, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed. Also, as used herein, the word “signal” refers to, among other things, indicating something to a corresponding decoder. For example, in certain embodiments the encoder signals a particular one of a plurality of transforms, coding modes or flags. In this way, in an embodiment the same transform, parameter, or mode is used at both the encoder side and the decoder side. Thus, for example, an encoder can transmit (explicit signaling) a particular parameter to the decoder so that the decoder can use the same particular parameter. Conversely, if the decoder already has the particular parameter as well as others, then signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding transmission of any actual functions, a bit savings is realized in various embodiments. It is to be appreciated that signaling can be accomplished in a variety of ways. For example, one or more syntax elements, flags, and so forth are used to signal information to a corresponding decoder in various embodiments. While the preceding relates to the verb form of the word “signal”, the word “signal” can also be used herein as a noun. As will be evident to one of ordinary skill in the art, implementations can produce a variety of signals formatted to carry information that can be, for example, stored or transmitted. The information can include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal can be formatted to carry the bitstream of a described embodiment. Such a signal can be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium. The preceding sections describe a number of embodiments, across various claim categories and types. Features of these embodiments can be provided alone or in any combination. Further, embodiments can include one or more of the following features, devices, or aspects, alone or in any combination, across various claim categories and types: At least one embodiment comprises determining a partition prediction using a neural network in an encoding and decoder process. At least one embodiment comprises encoding or decoding according to the above embodiments to implement a partitioning determination of a video block. At least one embodiment comprises any of the above embodiments wherein a neural network takes as input a vector comprising at least one of neighbor features, parent features, coding unit information, and spatial features. At least one embodiment comprises any of the above embodiments with neighbor features comprising at least one of a rate distortion cost of a neighboring block, and a quadtree depth of a neighboring block. At least one embodiment comprises any of the above embodiments with parent features comprising at least one of a cost of a split mode normalized by a no split cost of a parent block. At least one embodiment comprises any of the above embodiments with coding unit information comprising at least one of video block size, quantization parameter, and channel component. At least one embodiment comprises any of the above embodiments with spatial features comprising at least one of a resized video block, a processed resized video block, a gray-level co-occurrence matrix, information derived from a gray-level co- occurrence matrix, contrast, energy, homogeneity, correlation, dissimilarity, a histogram of oriented gradients, and a histogram of oriented gradients of causal borders of the video block. At least one embodiment comprises any of the above embodiments wherein a concatenation of histogram of gradients is input to a neural network. At least one embodiment comprises any of the above embodiments wherein information is signaled from an encoder to a decoder for neural network determination of partition of a video block. At least one embodiment comprises any encoding or decoding operation based on the above operations. At least one embodiment comprises performing encoding or decoding with the aforementioned methods on a sub-block. At least one embodiment comprises a bitstream or signal that includes one or more of the described syntax elements, or variations thereof. At least one embodiment comprises a bitstream or signal that includes syntax conveying information generated according to any of the embodiments described. At least one embodiment comprises creating and / or transmitting and / or receiving and / or decoding according to any of the embodiments described. At least one embodiment comprises a method, process, apparatus, medium storing instructions, medium storing data, or signal according to any of the embodiments described. At least one embodiment comprises inserting in the signaling syntax elements that enable the decoder to determine decoding information in a manner corresponding to that used by an encoder. At least one embodiment comprises creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements, or variations thereof. At least one embodiment comprises a TV, set-top box, cell phone, tablet, or other electronic device that performs transform method(s) according to any of the embodiments described. At least one embodiment comprises a TV, set-top box, cell phone, tablet, or other electronic device that performs transform method(s) determination according to any of the embodiments described, and that displays (e.g., using a monitor, screen, or other type of display) a resulting image. At least one embodiment comprises a TV, set-top box, cell phone, tablet, or other electronic device that selects, bandlimits, or tunes (e.g., using a tuner) a channel to receive a signal including an encoded image, and performs transform method(s) according to any of the embodiments described. At least one embodiment comprises a TV, set-top box, cell phone, tablet, or other electronic device that receives (e.g., using an antenna) a signal over the air that includes an encoded image, and performs transform method(s).

Claims

CLAIMS 1. A method, comprising: performing a partitioning prediction of a video block using a neural network; partitioning the video block based on said partitioning prediction; and, encoding the partitioned video block.

2. An apparatus, comprising: a memory, and a processor, configured to: perform a partitioning prediction of a video block using a neural network; partition the video block based on said partitioning prediction; and, encode the partitioned video block.

3. A method, comprising: performing a partitioning prediction of a video block using a neural network; partitioning the video block based on said partitioning prediction; and, decoding the partitioned video block.

4. An apparatus, comprising: a memory, and a processor, configured to: perform a partitioning prediction of a video block using a neural network; partition the video block based on said partitioning prediction; and, decode the partitioned video block.

5. The method of Claim 1 or 3, or the apparatus of Claim 2 or 4, wherein said neural network receives inputs comprising at least one of neighbor features, parent features, coding unit information, and spatial features.

6. The method or the apparatus of Claim 5, wherein said neighbor features comprise at least one of a rate distortion cost of a neighboring block, and a quadtree depth of a neighboring block.

7. The method or the apparatus of Claim 5, wherein said parent features comprise at least one of a cost of a split mode normalized by a no split cost of a parent block.

8. The method or the apparatus of Claim 5, wherein said coding unit information comprises at least one of a current video block size, quantization parameter, and channel component, one prediction type and one frame type.

9. The method or the apparatus of Claim 5, wherein said spatial features comprise at least one of a resized video block, a processed resized video block, a gray-level co-occurrence matrix, information derived from a gray-level co-occurrence matrix, contrast, energy, homogeneity, correlation, dissimilarity, a histogram of oriented gradients, and a histogram of oriented gradients of causal borders of the video block.

10. The method or the apparatus Claim 9, wherein said spatial features comprise a concatenation of two histograms of oriented gradients is used.

11. The method of any one of Claims 1, 3, 5-10, or the apparatus of any one of Claims 2, 4, or 5-10, wherein information is explicitly signaled in a bitstream used for said neural network.

12. A device comprising: an apparatus according to Claim 2; and at least one of (i) an antenna configured to receive a signal, the signal including the video block, (ii) a band limiter configured to limit the received signal to a band of frequencies that includes the video block, and (iii) a display configured to display an output representative of the video block.

13. A non-transitory computer readable medium containing data content generated according to the method of any one of claims 1, 3, or 5 through 11, or by the apparatus of any one of claims 2, 4, or 5 through 11, for playback using a processor.

14. A signal comprising video data generated according to the method of any one of claims 1, or 3, or 5 through 11, or by the apparatus of any one of claims 2, or 4, or 5 through 11, for playback using a processor.

15. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of claims 1, or 3 or 5 through 11.