Decoding of segmentation information using signaling

The hierarchical structure with cascading segmentation information processing layers addresses inefficiencies in existing image and video coding by enabling efficient decoding and adaptation, enhancing computational efficiency and scalability.

JP7839251B2Active Publication Date: 2026-04-01HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing image and video coding technologies, including hybrid codecs and machine learning applications, face challenges in optimizing encoding and decoding efficiency, particularly in handling feature maps across devices and adapting to desired parameters and content.

Method used

A method and apparatus for decoding data using a hierarchical structure with cascading segmentation information processing layers, enabling efficient decoding by processing segmentation information in multiple layers with upsampling and parallelization, and utilizing a trainable convolutional upsampling layer for improved efficiency and flexibility.

Benefits of technology

This approach enhances decoding efficiency by reducing computational complexity and improving processing time while allowing adaptation to various segments, making it suitable for scalable and parallelizable implementations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007839251000039
    Figure 0007839251000039
  • Figure 0007839251000040
    Figure 0007839251000040
  • Figure 0007839251000041
    Figure 0007839251000041
Patent Text Reader

Abstract

To provide a method and a device for decoding data.SOLUTION: Two or more sets of segmentation information elements are obtained from a bitstream. The two or more sets of segmentation information elements are then input to two or more segmentation information processing layers of a plurality of cascaded layers, respectively. At each of the two or more segmentation information processing layers, a respective set of segmentation information is processed. Decoding data for picture or video processing is obtained on the basis of segmentation information processed by the plurality of cascaded layers. Therefore, data can be decoded from the bitstream in an efficient method in a hierarchical structure.SELECTED DRAWING: Figure 9
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of this disclosure generally relate to the field of decoding data from a bitstream for image or video processing using multiple processing layers. In particular, some embodiments relate to methods and apparatus for such decoding. [Background technology]

[0002] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, the signal is typically encoded block by block by predicting blocks and further coding only the difference between the original block and its prediction. In particular, such coding can include transformation, quantization, and bitstream generation, and usually includes some form of entropy coding. Typically, the three components of the hybrid coding method—transformation, quantization, and entropy coding—are optimized separately. Modern video compression standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use the transformed representation to code the residual signal after prediction.

[0003] In recent years, machine learning has been applied to image and video coding. Generally, machine learning can be applied to image and video coding in a variety of different ways. For example, some end-to-end optimized image or video coding schemes have been discussed. Furthermore, machine learning has been used to determine or optimize certain parts of end-to-end coding, such as the selection or compression of predictive parameters. These applications share the commonality of generating some feature map data that is transmitted between the encoder and decoder. Efficient bitstream construction can significantly contribute to reducing the number of bits used to encode the image / video source signal.

[0004] A neural network typically contains two or more layers. Feature maps are the outputs of these layers. In a neural network that is split between devices, for example, between an encoder and a decoder, between a device and the cloud, or between different devices, the feature maps at the output of the split location (e.g., the first device) are compressed and sent to the remaining layers of the neural network (e.g., the second device).

[0005] It may be desirable to further improve encoding and decoding using a trained network architecture. [Overview of the project] [Means for solving the problem]

[0006] Some embodiments of this disclosure provide methods and apparatus for decoding pictures in an efficient manner and enabling some scalability to adapt to desired parameters and content.

[0007] The above and other objectives are achieved by the subject matter of the independent claims. Further implementations are evident from the dependent claims, specification, and drawings.

[0008] According to one embodiment, a method is provided for decoding data from a bitstream for picture or video processing, the method comprising: obtaining two or more sets of segmentation information elements from a bitstream; inputting each of the two or more sets of segmentation information elements into two or more segmentation information processing layers of a plurality of cascading layers; and processing each of the two or more segmentation information processing layers, wherein the step of obtaining the decoded data for picture or video processing is based on the segmentation information processed by the plurality of cascading layers.

[0009] Such a method can provide improved efficiency by enabling the decoding of data in various segments that can be constructed layer-based in a hierarchical structure. The provision of segments can take into account the characteristics of the data to be decoded.

[0010] For example, the step of obtaining a set of segmentation information elements is based on segmentation information processed by at least one segmentation information processing layer among multiple cascaded layers.

[0011] In some exemplary embodiments, the step of inputting a set of segmentation information elements is based on processed segmentation information output by at least one of a plurality of cascading layers.

[0012] Cascaded segmentation information processing enables efficient analysis of segmentation information.

[0013] For example, the segmentation information processed in two or more segmentation information processing layers may have different resolutions.

[0014] In some embodiments and examples, the step of processing segmentation information in two or more segmentation information processing layers includes the step of upsampling.

[0015] The hierarchical structure of segmentation information provides a small amount of side information to be inserted into the bitstream, thus improving efficiency and / or processing time.

[0016] In particular, upsampling the segmentation information includes nearest neighbor upsampling. Nearest neighbor upsampling has a low computational complexity and can be easily implemented. Moreover, it is efficient especially for logical representations such as flags. For example, upsampling the segmentation information includes transposed convolution. By performing upsampling, the upsampling quality can be improved. Further, such a convolutional upsampling layer may be provided as trainable or configurable in a decoder, and as a result, the convolutional kernel may be controlled by instructions parsed from a bitstream or derived by other means.

[0017] In an exemplary implementation, for each segmentation information processing layer j of a plurality of N segmentation information processing layers among a plurality of cascade layers, the input step, when j = 1, is to input initial segmentation information from a bitstream, and otherwise, to input the segmentation information processed by the (j - 1)th segmentation information processing layer, and the step of outputting the processed segmentation information.

[0018] For example, the step of processing the segmentation information input by each layer j < N of a plurality of N segmentation information processing layers is to parse segmentation information elements from a bitstream and associate the parsed segmentation information elements with the segmentation information output by the preceding layer, where the position of the parsed segmentation information elements within the associated segmentation information is determined based on the segmentation information output by the preceding layer, and further includes the step. In particular, the amount of segmentation information elements parsed from the bitstream is determined based on the segmentation information output by the preceding layer. For example, the parsed segmentation information elements are represented by a set of binary flags.

[0019] Such a hierarchical structure may be parallelizable, can be easily executed on a GPU / NPU, and provides a process that can utilize parallelism. A fully trainable method of transferring gradients enables its use in an end-to-end trainable video coding solution.

[0020] In some exemplary embodiments and examples, the step of obtaining decoded data for picture or video processing includes determining at least one of an intra-picture or inter-picture prediction mode, a picture reference index, single-reference or multi-reference prediction (including bi-prediction), present or absent prediction residual information, quantization step size, motion information prediction type, length of a motion vector, motion vector resolution, motion vector prediction index, size of a motion vector difference, motion vector difference resolution, motion compensation filter, in-loop filter parameters, post-filter parameters, based on the segmentation information. The decoding of the present disclosure is very generally applicable to any type of data related to picture or video coding.

[0021] The method of the above embodiments or examples further includes obtaining a set of feature map elements from a bitstream and inputting the set of feature map elements into a feature map processing layer among a plurality of layers based on the segmentation information processed by a segmentation information processing layer, and obtaining decoded data for picture or video processing based on the feature maps processed by a plurality of cascade layers.

[0022] In particular, at least one of the plurality of cascade layers is a segmentation information processing layer and a feature map processing layer. In other embodiments, each of the plurality of layers is either a segmentation information processing layer or a feature map processing layer.

[0023] The separated layer functions provide a clean design and functional separation. However, the present disclosure can also function when a layer implements both functions.

[0024] In one embodiment, a computer program product stored in a non-temporary medium is provided, and when this computer program product is executed on one or more processors, it performs the method according to any of the above examples and embodiments.

[0025] According to one embodiment, a device for decoding an image or video is provided, which includes a processing circuit configured to perform the method according to any of the above examples and embodiments.

[0026] According to one embodiment, a device is provided for decoding data from a bitstream for picture or video processing, the device including: an acquisition unit configured to acquire two or more sets of segmentation information elements from a bitstream; an input unit configured to input each of the two or more sets of segmentation information elements to two or more segmentation information processing layers of a plurality of cascade layers, respectively; a processing unit configured to process each of the two or more segmentation information processing layers for each of the two or more segmentation information processing layers; and a decoded data acquisition unit configured to acquire the decoded data for picture or video processing based on the segmentation information processed in the plurality of cascade layers.

[0027] Any of the above-described devices can be realized on an integrated chip. The present invention can be implemented in hardware (HW) and / or software (SW). Furthermore, the HW-based implementation can be combined with the SW-based implementation.

[0028] Please note that this disclosure is not limited to any particular framework. Furthermore, this disclosure is not limited to image or video compression, but may also apply to object detection, image generation, and recognition systems.

[0029] For clarity, any one of the embodiments described above may be combined with one or more of the other embodiments described above to create a new embodiment within the scope of this disclosure.

[0030] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. [Brief explanation of the drawing]

[0031] [Figure 1] This is a schematic diagram showing the channels processed by the layers of a neural network. [Figure 2] This is a schematic diagram illustrating an autoencoder-type neural network. [Figure 3A] This is a schematic diagram showing an exemplary network architecture for the encoder and decoder sides, including a highly pre-modeled version. [Figure 3B] This is a schematic diagram showing a typical network architecture for the encoder side, including a highly pre-modeled model. [Figure 3C] This is a schematic diagram showing a typical network architecture for the decoder side, including a highly pre-modeled model. [Figure 4] This is a schematic diagram showing an exemplary network architecture for the encoder and decoder sides, including a highly pre-modeled version. [Figure 5A] A block diagram showing a neural network-based end-to-end video compression framework. [Figure 5B] This block diagram shows some illustrative details of the application of neural networks for motion field compression. [Figure 5C] This block diagram shows some illustrative details of the application of neural networks for motion compensation. [Figure 6] This is a schematic diagram of the layers of the U-net. [Figure 7A] This is a block diagram illustrating an example of a hybrid encoder. [Figure 7B]This is a block diagram illustrating an exemplary hybrid decoder. [Figure 8] This flowchart shows an exemplary method for encoding data for picture / video processing, such as coding. [Figure 9] This block diagram shows the structure of a network that transfers information from layers of different resolutions within a bitstream. [Figure 10A] This is a schematic diagram showing maximum pooling. [Figure 10B] This is a schematic diagram illustrating mean pooling. [Figure 11] This is a schematic diagram illustrating the processing of feature maps and segmentation information by an exemplary encoder. [Figure 12] This block diagram shows the generalized processing of motion information feature maps by the encoder and decoder. [Figure 13] This block diagram shows the network structure for transferring information from layers of different resolutions within a bitstream to process motion vector-related information. [Figure 14] This block diagram shows an exemplary cost calculation unit with high cost tensor resolution. [Figure 15] This block diagram shows an exemplary cost calculation unit with low cost tensor resolution. [Figure 16] This is a block diagram illustrating the functional structure of signal selection logic. [Figure 17] This block diagram illustrates the functional structure of signal selection logic having a cost calculation unit that provides multiple coding options. [Figure 18] This block diagram shows the structure of a network that transfers information from layers of different resolutions within a bitstream, which has convolutional downsampling and upsampling layers. [Figure 19] This is a block diagram showing a structure that transfers information from layers of different resolutions within a bitstream that has additional layers. [Figure 20]This block diagram shows a structure for transferring information from layers of different resolutions within a bitstream, with layers that allow for downsampling or upsampling filter selection. [Figure 21] This block diagram shows the structure of a network that transfers information from layers of different resolutions within a bitstream, each having a layer that enables convolutional filter selection. [Figure 22] This block diagram illustrates the functional structure of a network-based RDO decision unit for selecting a coding mode. [Figure 23] This is a block diagram of an exemplary cost calculation unit that may be used in a network-based RDO decision unit for selecting a coding mode. [Figure 24] This is a block diagram of an exemplary cost calculation unit that may be used in a network-based RDO decision unit for selecting a coding mode that supports multiple options. [Figure 25] This is a schematic diagram showing possible block divisions or filter shapes. [Figure 26] This is a schematic diagram illustrating the derivation of segmentation information. [Figure 27] This is a schematic diagram illustrating how segmentation information is processed on the decoder side. [Figure 28] This is a block diagram illustrating exemplary signal supply logic for reconfiguring dense optical flow. [Figure 29] This is a block diagram illustrating exemplary signal supply logic for reconfiguring dense optical flow. [Figure 30] This is a block diagram of a convolutional filter set. [Figure 31] This is a block diagram showing an upsampling filter set. [Figure 32A] This is a schematic diagram showing the upsampling process on the decoder side using the nearest copy method. [Figure 32B] This is a schematic diagram showing the upsampling process on the decoder side using convolution. [Figure 33] This is a flowchart illustrating an exemplary method for decoding data such as feature map information used in the decoding of pictures or videos. [Figure 34] This is a flowchart illustrating an exemplary method for encoding data such as segmentation information used in picture or video encoding. [Figure 35] This is a block diagram showing an example of a video coding system configured to implement embodiments of the present invention. [Figure 36] This is a block diagram showing another example of a video coding system configured to implement embodiments of the present invention. [Figure 37] This is a block diagram showing an example of an encoding or decoding device. [Figure 38] This is a block diagram showing other examples of encoding or decoding devices. [Modes for carrying out the invention]

[0032] The following description refers to the accompanying drawings, which form part of this disclosure and, as an example, illustrate specific embodiments of the present invention or specific ways in which embodiments of the present invention may be used. It is understood that embodiments of the present invention may be used in other ways, and may include structural or logical modifications not depicted in the drawings. Therefore, the following detailed description should not be construed as restrictive, and the scope of the invention is defined by the appended claims.

[0033] For example, disclosure relating to a described method may also apply to a corresponding device or system configured to perform that method, and vice versa. For example, if one or more specific method steps are described, a corresponding device may include one or more units, e.g., functional units (e.g., one unit performing one or more steps, or multiple units each performing one or more of the steps), even if such one or more units are not explicitly described or shown in the drawings. On the other hand, for example, if a particular device is described based on one or more units, e.g., functional units, a corresponding method may include one step performing the function of one or more units (e.g., one step performing the function of one or more units, or multiple steps each performing one or more of the functions of multiple units), even if such one or more steps are not explicitly described or shown in the drawings. Furthermore, it is understood that the various exemplary embodiments and / or features of the aspects described herein can be combined with each other unless otherwise specified.

[0034] Some embodiments aim to improve the quality of encoded and decoded picture or video data and / or reduce the amount of data required to represent encoded picture or video data. Some embodiments provide efficient selection of information signaled from encoder to decoder. Below is an overview of some of the technical terms and frameworks used in embodiments of this disclosure that may be employed.

[0035] Artificial neural networks Artificial neural networks (ANNs), or connectionist systems, are computing systems that draw loose inspiration from the biological neural networks that make up animal brains. Such systems "learn" to perform tasks by considering examples and are generally not programmed with task-specific rules. For example, in image recognition, they may learn to identify images containing cats by analyzing exemplary images that have been manually labeled as "cat" or "not a cat," and using the results to identify cats in other images. They do this without prior knowledge of cats, for example, that cats have fur, tails, whiskers, and cat-like faces. Instead, they automatically generate discriminative characteristics from the examples they process.

[0036] ANNs are based on a collection of connected units or nodes called artificial neurons, which roughly model neurons in the brain of living organisms. Each connection can transmit signals to other neurons, similar to synapses in the brain of living organisms. The artificial neuron that receives the signal can then process the signal and signal to the neurons connected to it.

[0037] In ANN implementations, the "signals" in a connection are real numbers, and the output of each neuron is calculated by some nonlinear function of the sum of its inputs. These connections are called edges. Neurons and edges typically have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal in the connection. Neurons may have thresholds such that they only transmit a signal if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. The signal travels from the first layer (input layer) to the last layer (output layer), sometimes traversing the layers multiple times.

[0038] The initial goal of the ANN approach was to solve problems in the same way the human brain solves them. Over time, attention shifted to performing specific tasks, leading to deviations from biology. ANNs have been used for a wide range of tasks, including computer vision, speech recognition, machine translation, social network filtering, board games and video games, medical diagnosis, and even activities traditionally considered uniquely human, such as drawing.

[0039] The name "Convolutional Neural Network" (CNN) indicates that this network employs a mathematical operation called convolution. Convolution is a special type of linear operation. A convolutional network is a neural network that uses convolution instead of general matrix multiplication in at least one of its layers.

[0040] Figure 1 schematically illustrates the general concept of processing by neural networks such as CNNs. A convolutional neural network consists of an input layer, an output layer, and several hidden layers. The input layer is the layer to which the input (such as a portion of an image as shown in Figure 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that convolve by multiplication or other dot products. The result of the layers is one or more feature maps (f.map in Figure 1), sometimes called channels. Subsampling may exist that involves some or all of the layers. As a result, the feature maps may become smaller, as shown in Figure 1. The activation function in a CNN is usually a ReLU (Normalized Linear Unit) layer, followed by additional convolutions such as pooling layers, fully connected layers, and normalization layers, the inputs and outputs of which are masked by the activation function and the final convolution and are therefore called hidden layers. The layers are colloquially called convolutions, but this is merely a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This is important for indices in a matrix in that it affects how weights are determined at specific index points.

[0041] When programming a CNN for image processing, as shown in Figure 1, the input is a tensor with shape (number of images) × (image width) × (image height) × (image depth). Then, after passing through the convolutional layer, the image is abstracted into a feature map with shape (number of images) × (feature map width) × (feature map height) × (feature map channels). The convolutional layer in the neural network must have the following attributes: a convolution kernel (hyperparameter) defined by width and height; the number of input and output channels (hyperparameters); and the depth of the convolutional filter (input channels) must be equal to the number of channels (depth) in the input feature map.

[0042] In the past, conventional multilayer perceptron (MLP) models have been used for image recognition. However, due to the fully connected nodes, MLP models suffered from high dimensionality and did not handle high-resolution images well. A 1000x1000 pixel image with RGB color channels has 3 million weights, which is too high to be processed efficiently and conveniently on a large scale with fully connected nodes. Furthermore, such network architectures do not take into account the spatial structure of the data, treating distant input pixels as if they were close to each other. This ignores the locality of reference in image data, both computationally and semantically. Therefore, for purposes such as image recognition where spatially local input patterns are dominant, the fully connected nature of neurons is redundant.

[0043] Convolutional neural networks (CNNs) are a biologically inspired variation of multilayer perceptrons specifically designed to emulate the behavior of the visual cortex. These models mitigate the challenges posed by MLP architectures by leveraging the strong spatially local correlations present in natural images. Convolutional layers are the core building blocks of CNNs. The layer parameters consist of a set of learnable filters (kernels, as described above), which have small receptive fields but extend across the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, calculating the dot product between the filter's entry and the input to generate a two-dimensional activation map of that filter. As a result, the network learns which filters are activated when it detects a particular type of feature at a given spatial location in the input.

[0044] Stacking the activation maps of all filters along the depth dimension forms the total output volume of the convolutional layer. Therefore, all entries in the output volume can also be interpreted as the outputs of neurons that look at a small region in the input and share parameters with neurons in the same activation map. A feature map, or activation map, is the output activation of a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a mapping that corresponds to the activation of different parts of an image, and also called a feature map because it is a mapping of where a particular type of feature is found in the image. High activation means that a particular feature has been found.

[0045] Another important concept in CNNs is pooling, which is a form of nonlinear downsampling. There are several nonlinear functions for implementing pooling, the most common of which is maximum pooling. Maximum pooling divides the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.

[0046] Intuitively, the precise location of a feature is less important than its approximate location relative to other features. This is the concept behind the use of pooling in convolutional neural networks. Pooling layers help progressively reduce the spatial size of representations, thereby reducing the number of parameters in the network, the memory footprint, and the computational complexity, and thus also controlling overfitting. In CNN architectures, it is common to periodically insert pooling layers between consecutive convolutional layers. Pooling operations provide another form of transformation invariance.

[0047] A pooling layer operates independently for each depth slice of the input, spatially resizing it. The most common form is a pooling layer with a 2x2 filter, applied with a stride of 2 for each depth slice of the input, 2 along both width and height, discarding 75% of the activations. In this case, all maximization operations span four numbers. The depth dimension remains invariant. In addition to maximal pooling, the pooling unit can also use other functions such as mean pooling or L2 norm pooling. Mean pooling has been commonly used in the past but has become less preferred recently compared to maximal pooling, which often performs better in practice. There is a recent trend to use smaller filters or discard the pooling layer altogether for aggressive reduction of the size of the representation. "Region of Interest" pooling (also known as ROI pooling) is a variation of maximal pooling where the output size is fixed and the input rectangle is a parameter. Pooling is a key component of convolutional neural networks for object detection based on the fast R-CNN architecture.

[0048] The above ReLU stands for Normalized Linear Unit and applies a non-saturated activation function. It effectively removes negative values ​​from the activation map by setting negative values ​​to 0. It increases the nonlinear properties of the decision function and the entire network without affecting the receptive field of the convolutional layer. Other functions, such as the saturated hyperbolic tangent and sigmoid functions, are also used to increase nonlinearity. ReLU is often preferred over other functions because it trains neural networks several times faster without significantly impairing generalized accuracy.

[0049] After several convolutional and maximum pooling layers, high-level inference in the neural network is performed via fully connected layers. The neurons in the fully connected layers are connected to all activations of the previous layer, as seen in typical (non-convolutional) artificial neural networks. Thus, their activations can be computed as affine transformations, followed by matrix multiplication and then bias offsets (vector addition of learned or fixed bias terms).

[0050] The "loss layer" (which includes the calculation of the loss function) specifies how training penalizes the deviation between the predicted (output) label and the true label, and is usually the final layer of a neural network. Various loss functions can be used that are suitable for different tasks. Softmax loss is used to predict a single class from K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values ​​in [0,1]. Euclidean loss is used to regress to real-valued labels.

[0051] In summary, Figure 1 illustrates the data flow in a typical convolutional neural network. First, the input image passes through a convolutional layer and is abstracted into a feature map containing multiple channels corresponding to the number of filters in the set of learnable filters in this layer. The feature map is then subsampled, for example, using a pooling layer, which reduces the dimensionality of each channel in the feature map. Next, the data reaches another convolutional layer, which may have a different number of output channels. As mentioned above, the number of input and output channels are hyperparameters of the layer. To establish network connectivity, these parameters need to be synchronized between two connected layers so that the number of input channels in the current layer is equal to the number of output channels in the previous layer. For a first layer processing input data, e.g., an image, the number of input channels is typically equal to the number of channels in the data representation, e.g., three channels for an RGB or YUV representation of an image or video, or one channel for a grayscale image or video representation.

[0052] Autoencoders and unsupervised learning An autoencoder is a type of artificial neural network used to learn efficient data coding in an unsupervised manner. A schematic diagram is shown in Figure 2. The purpose of an autoencoder is to learn a representation (encode) of a dataset, typically for dimensionality reduction, by training the network to ignore signal "noise". Along with the reduction side, the reconstruction side is learned, and the autoencoder attempts to generate a representation from the reduced encoding that is as close as possible to its original input, and thus its name. In its simplest form, given one hidden layer, the encoder stage of the autoencoder takes input x and maps it to h. h = σ(Wx + b).

[0053] This image h is commonly referred to as the code, latent variable, or latent representation. Here, σ is an element-wise activation function, such as a sigmoid function or normalized linear unit. W is the weight matrix, and b is the bias vector. The weights and biases are typically initialized randomly and then iteratively updated during training via backpropagation. The decoder stage of the autoencoder then maps h to a reconstructed x' of the same shape as x: x'=σ'(W'h'+b') Here, the decoder's σ', W', and b' may be independent of the corresponding σ, W, and b of the encoder.

[0054] Variational autoencoder models make strong assumptions about the distribution of latent variables. They use a variational approach to latent representation learning, resulting in an additional loss component and specific estimators for the training algorithm, called stochastic gradient variational Bayes (SGVB) estimators. The data is used in directed graphical model p θ (x|h) is generated, and the encoder is the posterior distribution p θ Approximation q for (x|h) Φ Assuming that (h|x) is being learned, Φ and θ represent the parameters of the encoder (recognition model) and decoder (generative model), respectively. The probability distribution of the latent vector of a VAE typically closely matches the probability distribution of the training data much better than a standard autoencoder. The objective of a VAE has the following form:

number

[0055] In the formula, D KL This represents the Kullback-Leibler divergence. The prior distribution for the latent variables is a central isotropic multivariate Gaussian distribution p θ (h) is typically set to N(0,I). Generally, the shapes of the variational and likelihood distributions are chosen so that they are factored Gaussian distributions: q Φ (h|x)=N(ρ(x),ω 2 (x)I) pΦ (x|h) = N(μ(h), σ 2 (h)Ι) where ρ(x) and ω 2 (x) is the encoder output, and μ(h) and σ 2 (h) is the decoder output.

[0056] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, enable researchers' interest in applying neural network-based techniques to the tasks of image and video compression. For example, end-to-end optimized image compression using a network based on variational autoencoders has been proposed.

[0057] Correspondingly, data compression is regarded as a fundamental and well-studied problem in engineering and is generally formulated with the aim of designing codes for a given discrete data ensemble with minimum entropy. This solution relies heavily on knowledge of the probabilistic structure of the data, and thus the problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as vectors of image pixel intensities) must be quantized into a finite set of discrete values, which results in errors.

[0058] In this context, a trade-off must be made between two competing costs: the entropy (rate) of the discretized representation, known as the irreversible compression problem, and the error (distortion) resulting from quantization. Different compression applications, such as data storage or transmission over a channel with limited capacity, require different rate-distortion trade-offs.

[0059] Simultaneous optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional spaces is unwieldy. For this reason, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-value representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This method is called transformative coding due to the central role of the transformation.

[0060] For example, JPEG uses the discrete cosine transform for blocks of pixels, while JPEG2000 uses multiscale orthogonal wavelet decomposition. Typically, the three components of the transform coding method—transform, quantization, and entropy coding—are optimized separately (often by manual parameter tuning). Modern video compression standards such as HEVC, VVC, and EVC also use transformed representations to code residual signals after prediction. Several transforms are used for this purpose, including the discrete cosine transform and discrete sine transform (DCT, DST), as well as the low-frequency inseparable manual optimization transform (LFNST).

[0061] Variational image compression The Variable Automatic Encoder (VAE) framework can be considered a nonlinear transformation coding model. The transformation process can be broadly divided into four parts, which are illustrated in Figure 3A, which shows the VAE framework.

[0062] The transformation process can be divided into four main parts, and Figure 3A illustrates the VAE framework. In Figure 3A, the encoder 101 maps the input image x to a latent representation (represented by y) via the function y=f(x). This latent representation may hereafter be referred to as part of the “latent space” or “points in the latent space”. The function f() is a transformation function that converts the input signal x to a more compressible representation y. The quantizer 102 uses Q, which represents the quantizer function, to convert the latent representation y

number

number

number

[0063] The latent space can be understood as a compressed representation of data where similar data points are close to each other within the latent space. The latent space is useful for learning data features and finding simpler representations of the data for analysis. Quantized latent representation T

number

number

number

number

number

number

[0064] In Figure 3A, component AE105 is an arithmetic coding module and represents a quantized latent representation.

number

number

number

number

[0065] Arithmetic decoding (AD) 106 is the process of reversing the binary conversion, where the binary numbers are converted back to sample values. Arithmetic decoding is provided by the arithmetic decoding module 106.

[0066] Please note that this disclosure is not limited to this particular framework. Furthermore, this disclosure is not limited to image or video compression, but can also be applied to object detection, image generation, and recognition systems.

[0067] In Figure 3A, there are two interconnected subnetworks. In this context, a subnetwork is a logical division between parts of an entire network. For example, in Figure 3A, modules 101, 102, 104, 105, and 106 are referred to as the "encoder / decoder" subnetwork. The "encoder / decoder" subnetwork is responsible for encoding (generating) and decoding (analyzing) the first bitstream, "bitstream 1". The second network in Figure 3A, including modules 103, 108, 109, 110, and 107, is referred to as the "hyperencoder / decoder" subnetwork. The second subnetwork is responsible for generating the second bitstream, "bitstream 2". The two subnetworks have different purposes.

[0068] The first subnetwork is, • Transformation of input image x to its latent representation y (this makes it easier to compress x), • Latent expression y is a quantized latent expression

number

number

number

[0069] The purpose of the second subnetwork is to obtain the statistical properties of the samples in "bitstream 1" (e.g., mean, variance, and correlation between samples in bitstream 1) so that the compression of bitstream 1 by the first subnetwork becomes more efficient. The second subnetwork generates a second bitstream, "bitstream 2," which contains this information (e.g., mean, variance, and correlation between samples in bitstream 1).

[0070] The second network is a quantized latent representation

number

number

number

number

number

number

number

number

number

number

number

number

number

[0071] Figure 3A illustrates an example of a VAE (Variational Autoencoder), the details of which may differ in different implementations. For example, in certain implementations, additional components may exist to more efficiently obtain the statistical properties of the samples in bitstream 1. In one such implementation, there may be a context modeler that targets the extraction of cross-correlation information of bitstream 1. The statistical information provided by the second subnetwork may be used by the AE (Arithmetic Encoder) 105 and AD (Arithmetic Decoder) 106 components.

[0072] Figure 3A shows an encoder and decoder in a single diagram. As will be apparent to those skilled in the art, encoders and decoders may, and very often, be embedded in different devices.

[0073] Figure 3B shows the encoder, and Figure 3C shows the decoder components of the VAE framework separated. As input, the encoder receives a picture, according to some embodiments. The input picture may include one or more channels, such as a color channel or other types of channels, for example, a depth channel or a motion information channel. The outputs of the encoder (as shown in Figure 3B) are bitstream 1 and bitstream 2. Bitstream 1 is the output of the encoder's first subnetwork, and bitstream 2 is the output of the encoder's second subnetwork.

[0074] Similarly, in Figure 3C, two bitstreams, namely bitstream 1 and bitstream 2, are received as input, and the resulting image is reconstructed (decoded).

number

[0075] Specifically, as shown in Figure 3B, the encoder includes an encoder 121 that converts input x into a signal y, which is then provided to a quantizer 322. The quantizer 122 provides information to the arithmetic coding module 125 and the hyperencoder 123. The hyperencoder 123 provides the bitstream 2 already described above to the hyperdecoder 147, which then provides this information to the arithmetic coding module 105 (125).

[0076] The output of the arithmetic coding module is bitstream 1. Bitstreams 1 and 2 are the outputs of the coded signal, which are then provided (transmitted) to the decoding process. Unit 101 (121) is referred to as the “encoder,” but the entire subnetwork described in Figure 3B can also be referred to as the “encoder.” The coding process generally refers to a unit (module) that converts an input to an coded (e.g., compressed) output. As seen in Figure 3B, unit 121 can be considered the core of the entire subnetwork, as it performs the conversion of input x to a compressed version, y. Compression in encoder 121 may be achieved, for example, by applying a neural network, or any processing network generally having one or more layers. In such a network, compression may be performed by cascading processes that include downsampling, which reduces the size and / or number of channels in the input. Thus, the encoder may be referred to, for example, a neural network (NN) based encoder.

[0077] The remaining parts in the diagram (quantization unit, hyperencoder, hyperdecoder, arithmetic encoder / decoder) are all responsible for improving the efficiency of the encoding process or converting the compressed output y into a sequence of bits (bitstream). Quantization may be provided to further compress the output of the NN encoder 121 by lossy compression. The AE125, in combination with the hyperencoder 123 and hyperdecoder 127 used to constitute the AE125, may perform binarization, which can further compress the quantized signal by lossless compression. Thus, the entire subnetwork in Figure 3B can also be referred to as the "encoder".

[0078] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of a signal before converting it to binary digits (bits). For example, in the VAE framework, the encoder, which is a nonlinear transformation, maps the input image x to y, where y has a smaller width and height than x. Because y has a smaller width and height, and therefore a smaller size, the dimensionality (size) of the signal is reduced, and thus it is easier to compress the signal y. Generally, it should be noted that encoders do not necessarily have to reduce the size in both (or generally all) dimensions. Rather, some exemplary implementations may provide encoders that reduce the size in only one dimension (or generally a subset of dimensions).

[0079] J. Balle, L. Valero Laparra, and EP Simoncelli (2015). In "Density Modeling of Images Using a Generalized Normalization Transformation," In:arXiv e-prints, presented at the 4th Int.Conf. for Learning Representations, 2016 (hereinafter referred to as "Balle"), the authors proposed a framework for end-to-end optimization of image compression models based on nonlinear transformations. The authors optimize for mean squared error (MSE) but use a more flexible transformation constructed from a cascade of linear convolution and nonlinearity. In particular, the authors use a generalized partition-normalization (GDN) coupled nonlinearity, conceived from a model of neurons in the biological visual system and proven effective for Gaussian image density. Following this cascaded transformation, a uniform scalar quantization (i.e., each element is rounded to the nearest integer) is performed, effectively implementing parametric vector quantization in the original image space. The compressed image is reconstructed from these quantized values ​​using an approximate parametric nonlinear inverse transform.

[0080] An example of the VAE framework is shown in Figure 4, which utilizes six downsampling layers marked 401-406. The network architecture includes a highly prior model. (g) a ,g s ) shows the image autoencoder architecture, and on the right (h a ,h s ) corresponds to an autoencoder that implements ultra-priority. The factored prior model is analyzed and composited g a and g s The same architecture is used. Q represents quantization, and AE and AD represent arithmetic encoder and arithmetic decoder, respectively. The encoder takes g from the input image x. a This process is applied, resulting in a response y (latent representation) with a spatially varying standard deviation. Encoding g a This includes multiple convolutional layers with subsampling and Generalized Decomposition Normalization (GDN) as the activation function.

[0081] The response is h a It is supplied to summarize the distribution of the standard deviation at z. Then z is quantized, compressed and transmitted as side information. The encoder then processes the quantized vector

number

number

number

number

number

number

number

[0082] Layers that include downsampling are indicated by a downward arrow in the layer description. The layer description "Conv Nx5x5 / 2↓" means that the layer is a convolutional layer with N channels and the convolutional kernel is 5x5 in size. As mentioned above, 2↓ means that 2x downsampling is performed in this layer. 2x downsampling results in one of the dimensions of the input signal being reduced by half in the output. In Figure 4, 2↓ indicates that both the width and height of the input image are reduced by half. Since there are six downsampling layers, if the width and height of the input image 4¹⁴ (also indicated by x) are given by w and h, the output signal z^4¹⁴ will have a width and height equal to w / 64 and h / 64, respectively. The modules indicated by AE and AD are arithmetic encoders and arithmetic decoders, which are described with reference to Figures 3A to 3C. Arithmetic encoders and decoders are specific implementations of entropy coding. AE and AD can be replaced by other means of entropy coding. In information theory, entropy coding is a lossless data compression method used to convert the values ​​of symbols into a binary representation, which is a reversible process. The "Q" in the figure corresponds to the quantization operation described above in relation to Figure 4, which will be further explained in the "Quantization" section. Furthermore, the quantization operation and corresponding quantization unit as part of component 413 or 415 are not necessarily present and / or can be replaced by other units.

[0083] Figure 4 also shows a decoder including upsampling layers 407-412. An additional layer 420 is provided between the upsampling layers 411 and 410, implemented as a convolutional layer, but in an input processing order that does not provide upsampling to the received input. A corresponding convolutional layer 430 is also shown for the decoder. Such a layer can be provided in the NN to perform operations on inputs that do not change the size of the input but change certain characteristics. However, such a layer is not required.

[0084] Seen in the processing order of bitstream 2 passing through the decoder, the upsampling layers are executed in the reverse order, i.e., from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide upsampling with an upsampling ratio of 2, indicated by ↑. Of course, it is not necessarily required that all upsampling layers have the same upsampling ratio, and other upsampling ratios such as 3, 4, 8 may be used. Layers 407-412 are implemented as convolutional layers (conv). Specifically, since they may be intended to provide an operation on the input that is the inverse of the encoder operation, the upsampling layers may apply a deconvolution operation to the received input such that their size increases by a coefficient corresponding to the upsampling ratio. However, this disclosure is not limited in general to deconvolution, and upsampling may be performed in any other way, such as by bilinear interpolation between two nearest neighbor samples or by nearest neighbor sample copying.

[0085] In the first subnetwork, after some convolutional layers (401-403), generalized division normalization (GDN) is applied on the encoder side, and inverse GDN (IGDN) is applied on the decoder side. In the second subnetwork, the activation function applied is ReLU. Note that this disclosure is not limited to this implementation, and generally, other activation functions may be used instead of GDN or ReLU.

[0086] End-to-end image or video compression DNN-based image compression methods can leverage large-scale end-to-end training and highly nonlinear transformations not available in conventional approaches. However, directly applying these techniques to build an end-to-end learning system for video compression is not trivial. Firstly, learning how to generate and compress motion information tailored for video compression remains an unresolved problem. Video compression methods rely heavily on motion information to reduce temporal redundancy in video sequences.

[0087] A direct solution is to represent motion information using learning-based optical flow. However, current learning-based optical flow approaches aim to generate the most accurate flow field possible. Accurate optical flow is often not optimal for a particular video task. In addition, the amount of data in optical flow is significantly larger compared to motion information in conventional compression systems, and directly applying existing compression techniques to compress optical flow values ​​would significantly increase the number of bits required to store motion information. Secondly, it is unclear how to construct a DNN-based video compression system by minimizing rate distortion-based objectives for both residual and motion information. Rate distortion optimization (RDO) aims to achieve higher quality (i.e., less distortion) of the reconstructed frame given the number of bits (or bitrate) for compression. RDO is critical to video compression performance. To leverage the end-to-end training capability for learning-based compression systems, RDO strategies need to be optimized for the entire system.

[0088] Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, Zhiyong Gao; “DVC: ​​An End-to-end Deep Video Compression Framework”. In the Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015, the authors proposed an end-to-end deep video compression (DVC) model that learns motion estimation, motion compression, and residual coding together.

[0089] Such an encoder is shown in Figure 5A. In particular, Figure 5A shows the overall structure of an end-to-end trainable video compression framework. To compress motion information, a CNN is specified to transform the optical flow into a corresponding representation suitable for better compression. Specifically, an autoencoder-type network is used to compress the optical flow. The motion vector (MV) compression network is shown in Figure 5B. The network architecture is somewhat similar to ga / gs in Figure 4. In particular, the optical flow is fed into a series of convolutional operations and nonlinear transformations, including GDN and IGDN. The number of output channels for convolution (deconvolution) is 128, except for the last deconvolution layer which is equal to 2. Given an optical flow of size M × N × 2, the MV encoder generates a motion representation of size M / 16 × N / 16 × 128. The motion representation is then quantized, entropy coded, and sent to a bitstream. The MV decoder receives the quantized representation and uses the MV encoder to reconstruct the motion information.

[0090] Figure 5C shows the structure of the motion compensation unit. Here, the previously reconstructed frame x t-1Using the reconstructed motion information, the warping unit generates warped frames (usually with the help of interpolation filters such as bilinear interpolation filters). A separate CNN with three inputs then generates the predicted picture. The architecture of the motion-compensated CNN is also shown in Figure 5C.

[0091] The residual information between the original frame and the predicted frame is encoded by a residual encoder network. A highly nonlinear neural network is used to transform the residuals into their corresponding latent representations. Compared to the discrete cosine transform in conventional video compression systems, this approach can better utilize the power of the nonlinear transform and achieve higher compression efficiency.

[0092] From the above overview, it can be seen that CNN-based architectures can be applied to both image and video compression, taking into account different parts of the video framework, including motion estimation, motion compensation, and residual coding. Entropy coding is a common method used for data compression, widely adopted in the industry, and is also applicable to feature map compression for either human perception or computer vision tasks.

[0093] Machine video coding Machine video coding (VCM) is another direction in computer science that is gaining popularity today. The main concept behind this approach is to transmit coded representations of image or video information intended for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to traditional image and video coding aimed at human perception, the quality characteristic is not the reconstructed quality, but rather the performance on computer vision tasks, such as object detection accuracy.

[0094] Recent research has proposed a new deployment paradigm called collaborative intelligence, in which deep models are partitioned between mobile and cloud. Extensive experimentation under various hardware configurations and wireless connectivity modes has revealed that the optimal operating point in terms of energy consumption and / or computational latency typically involves partitioning the model at a deep point within the network. It has been found that today's common solutions, where the model resides entirely in the cloud or entirely on mobile, are (if any) rarely optimal. The concept of collaborative intelligence has also been extended to model training. In this case, data flows from cloud to mobile during backpropagation in training, from mobile to cloud during the forward pass in training, and in both inference and inference directions.

[0095] Lossy compression of deep feature data has been studied based on HEVC intracoding in the context of recent deep models for object detection. The degradation of detection performance with increasing compression levels has been noted, and compression-extension training has been proposed to minimize this loss by generating models more robust to quantization noise in feature values. However, this remains a suboptimal solution because the codecs employed are highly complex and optimized for natural scene compression rather than deep feature compression.

[0096] The problem of deep feature compression for collaborative intelligence has been addressed by approaches for object detection tasks using the common YOLOv2 network, for the purpose of studying the trade-off between compression efficiency and recognition accuracy. Here, the term deep feature is synonymous with feature map. The word "deep" derives from the concept of collaborative intelligence where some hidden (deep) layer output feature map is captured and transferred to the cloud for inference. This appears to be more efficient than sending compressed natural image data to the cloud and performing object detection using the reconstructed image.

[0097] Efficient feature map compression is useful for compressing and reconstructing images and videos for both human perception and machine vision. The shortcomings of modern autoencoder-based approaches to compression are also relevant to machine vision tasks.

[0098] Artificial neural networks with skip connections Residual neural networks (ResNets) are a type of artificial neural network (ANN) built on a construct known from pyramidal cells in the cerebral cortex. Residual neural networks do this by utilizing skip connections or shortcuts to skip certain layers. Typical ResNet models are implemented with two or three skips, including nonlinearity (ReLU) and batch normalization between them. Additional weight matrices may be used to learn the skip weights, and these models are known as HighwayNets. Models with multiple parallel skips are called DenseNets. In the context of residual neural networks, non-residual networks can be described as plain networks.

[0099] One motivation for skipping layers is to avoid the vanishing gradient problem by reusing activations from previous layers until adjacent layers learn their weights. During training, weights are adapted to mute upstream layers and amplify previously skipped layers. In the simplest case, only weights are adapted for connections in adjacent layers, with no explicit weights for upstream layers. This works best when a single nonlinear layer is stepped over, or when all intermediate layers are linear. Otherwise, an explicit weight matrix should be learned for skipped connections (HighwayNet should be used).

[0100] Skipping effectively simplifies the network by using fewer layers in the initial training phase. This speeds up learning by reducing the effects of vanishing gradients as fewer layers propagate. The network then gradually reconstructs the skipped layers as it learns the feature space. Towards the end of training, when all layers are expanded, it remains closer to the manifold and therefore learns faster. A neural network without residuals explores a larger feature space. This makes it more vulnerable to perturbations that move it away from the manifold and requires extra training data to reconstruct.

[0101] As shown in Figure 6, longer skip connections were introduced to the U-Net. The U-Net architecture is derived from the so-called "fully convolutional network" first proposed by Long and Shelhamer. The main concept is to supplement a normal collating network with continuous layers in which pooling operations are replaced by upsampling operators. Thus, these layers increase the resolution of the output. Furthermore, the continuous convolutional layers can then learn to assemble a precise output based on this information.

[0102] One significant variation in U-Net is the presence of numerous feature channels in the upsampling portion, which allows the network to propagate contextual information to higher resolution layers. As a result, the dilation path is more or less symmetric to the collision path, leading to a U-shaped architecture. The network uses only the effective portion of each convolution without fully connected layers. Missing context is extrapolated by mirroring the input image to predict pixels within the image's boundary regions. This tiling strategy is crucial for applying the network to large images, where resolution is otherwise limited by GPU memory.

[0103] Introducing skip connections allows for better capture of features at different spatial resolutions, which is successfully applied to computer vision tasks such as object detection and segmentation. However, implying such skip connections for image or video compression is not a simple task, as information from the encoding side must be transferred in the communication channel, and direct connections between layers require the transfer of a considerable amount of data.

[0104] Conventional hybrid video encoding and decoding Neural network frameworks can also be employed in combination with, or within, conventional hybrid coding and decoding, as will be illustrated later. A very brief overview of exemplary hybrid coding and decoding is given below.

[0105] Figure 7A shows a schematic block diagram of an exemplary video encoder 20 configured to implement the technology of the present application. In the example of Figure 7A, the video encoder 20 includes an input 201 (or input interface 201), a residual calculation unit 204, a transformation processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transformation processing unit 212, a reconstruction unit 214, a loop filter unit 220, a decoding picture buffer (DPB) 230, a mode selection unit 260, an entropy coding unit 270, and an output 272 (or output interface 272). The mode selection unit 260 may include an inter-prediction unit 244, an intra-prediction unit 254, and a splitting unit 262. The inter-prediction unit 244 may include a motion estimation unit and a motion compensation unit (not shown). The video encoder 20 shown in Figure 7A may also be referred to as a hybrid video encoder or a video encoder with a hybrid video codec.

[0106] The encoder 20 may be configured to receive, for example, a picture 17 (or picture data 17) via input 201, for example, a picture of a sequence of pictures that form a video or video sequence. The received picture or picture data may be a pre-processed picture 19 (or pre-processed picture data 19). For simplicity, the following description will refer to picture 17. Picture 17 may also be referred to as the current picture or the picture being coded (in particular, in video coding to distinguish the current picture from other pictures, for example, previously coded and / or coded pictures of the same video sequence, i.e., the video sequence that also includes the current picture).

[0107] A (digital) picture is, or can be considered as, a two-dimensional array or matrix of samples having intensity values. Samples in an array are sometimes referred to as pixels (a shortened form of picture element) or pels. The number of samples in the horizontal and vertical (or axis) directions of an array or picture defines the size and / or resolution of the picture. For color representation, typically three color components are employed; i.e., a picture can represent or contain three sample arrays. In the RGB format or color space, a picture contains corresponding red, green, and blue sample arrays. However, in video coding, each pixel is typically represented in a luminance and chrominance format or color space, e.g., YCbCr, where each pixel contains a luminance component represented by Y (sometimes L is used instead) and two chrominance components represented by Cb and Cr. The luminance (or short luma) component Y represents brightness or gray level intensity (for example, in a grayscale picture), and the two chrominance (or short chroma) components Cb and Cr represent chromaticity or color information components. Therefore, a picture in YCbCr format contains a luminance sample array of luminance sample values ​​(Y) and two chrominance sample arrays of chrominance values ​​(Cb and Cr). A picture in RGB format can be converted to or from YCbCr format, and vice versa; this process is also known as color conversion or transformation. If a picture is monochrome, it may contain only a luminance sample array. Thus, a picture could be, for example, an array of luma samples in a monochrome format, or an array of luma samples and two corresponding arrays of chroma samples in 4:2:0, 4:2:2, and 4:4:4 color formats.

[0108] Embodiments of the video encoder 20 may include a picture splitting unit (not shown in Figure 7A) configured to divide a picture 17 into multiple (typically non-overlapping) picture blocks 203. These blocks may also be referred to as root blocks, macroblocks (H.264 / AVC), coding tree blocks (CTB), or coding tree units (CTU) (H.265 / HEVC and VVC). The picture splitting unit may be configured to use the same block size and corresponding grid defining the block size for all pictures in a video sequence, or to change the block size between pictures or subsets or groups of pictures, dividing each picture into a corresponding block. The abbreviation AVC stands for Advanced Video Coding.

[0109] In a further embodiment, the video encoder may be configured to directly receive blocks 203 of picture 17, for example, one, more, or all of the blocks that make up picture 17. Picture blocks 203 may also be referred to as the current picture block or the coded picture block.

[0110] Similar to picture 17, picture block 203 is, or can be considered as, a two-dimensional array or matrix of samples having intensity values ​​(sample values), although it has fewer dimensions than picture 17. In other words, block 203 may contain, depending on the applied color format, for example, one sample array (e.g., a lumen array for monochrome picture 17, or a lumen or chromen array for color pictures), three sample arrays (e.g., a lumen and two chromen arrays for color picture 17), or any other number and / or type of arrays. The number of samples in the horizontal and vertical (or axis) directions of block 203 defines the size of block 203. Thus, a block may be, for example, an M×N (M columns × N rows) array of samples, or an M×N array of conversion coefficients.

[0111] An embodiment of the video encoder 20, as shown in Figure 7A, may be configured to encode the picture 17 block by block, for example, encoding and prediction being performed for each block 203.

[0112] The embodiment of the video encoder 20 shown in Figure 7A may be further configured to divide and / or encode a picture using slices (also referred to as video slices), the picture may be divided into one or more slices (typically non-overlapping) or encoded using them, each slice may contain one or more blocks (e.g., CTUs).

[0113] Embodiments of the video encoder 20, as shown in Figure 7A, may be further configured to divide and / or encode a picture using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), wherein a picture may be divided into one or more tile groups (typically non-overlapping) or encoded using them, each tile group may, for example, contain one or more blocks (e.g., CTUs) or one or more tiles, each tile may, for example, be rectangular in shape and contain one or more blocks (e.g., CTUs), e.g., a complete block or a partial block.

[0114] Figure 7B shows an example of a video decoder 30 configured to implement the technology of the present application. The video decoder 30 is configured to receive encoded picture data 21 (e.g., encoded bitstream 21), encoded by, for example, the encoder 20, in order to obtain a decoded picture 331. The encoded picture data or bitstream includes information for decoding the encoded picture data, for example, data representing picture blocks of an encoded video slice (and / or tile group or tile), and associated syntax elements.

[0115] The entropy decoding unit 304 is configured to analyze the bitstream 21 (or generally the encoded picture data 21) and perform entropy decoding on the encoded picture data 21, for example, to obtain one or all of the following: for example, quantization coefficients 309 and / or decoding coding parameters (not shown in Figure 3), for example, inter-prediction parameters (e.g., reference picture index and motion vector), intra-prediction parameters (e.g., intra-prediction mode or index), transformation parameters, quantization parameters, loop filter parameters, and / or other syntax elements. The entropy decoding unit 304 may be configured to apply a decoding algorithm or scheme corresponding to an encoding scheme such as the one described with respect to the entropy coding unit 270 of the encoder 20. The entropy decoding unit 304 may be further configured to provide the inter-prediction parameters, intra-prediction parameters, and / or other syntax elements to the mode application unit 360 and other parameters to other units of the decoder 30. The video decoder 30 may receive syntax elements at the video slice level and / or video block level. In addition to slices and their respective syntax elements, tile groups and / or tiles and their respective syntax elements may be received and / or used as alternatives.

[0116] The reconstruction unit 314 (e.g., an adder or summer 314) may be configured to add the reconstructed residual block 313 to the predicted block 365 by, for example, adding the sample value of the reconstructed residual block 313 to the sample value of the predicted block 365, thereby obtaining the reconstructed block 315 in the sample region.

[0117] An embodiment of the video decoder 30, as shown in Figure 7B, may be configured to divide and / or decode a picture using slices (also referred to as video slices), the picture may be divided into one or more slices (typically non-overlapping) or decoded using them, each slice may contain one or more blocks (e.g., CTUs).

[0118] Embodiments of the video decoder 30, as shown in Figure 7B, may be configured to divide and / or decode a picture using tile groups (also referred to as video tile groups) and / or tiles (also referred to as video tiles), wherein the picture may be divided into one or more tile groups (typically non-overlapping) or decoded using them, each tile group may, for example, contain one or more blocks (e.g., CTUs) or one or more tiles, each tile may, for example, be rectangular in shape and may contain one or more blocks (e.g., CTUs), e.g., a complete block or a partial block.

[0119] Other variations of the video decoder 30 may be used to decode the encoded picture data 21. For example, the decoder 30 can generate an output video stream without using the loop filtering unit 320. For example, a non-transformation-based decoder 30 can directly dequantize the residual signal for some blocks or frames without the inverse transformation processing unit 312. In another implementation, the video decoder 30 may have an inverse quantization unit 310 and an inverse transformation processing unit 312 combined into a single unit.

[0120] It should be understood that the encoder 20 and decoder 30 may further process the results of the current step and then output them to the next step. For example, after interpolation filtering, motion vector derivation, or loop filtering, further operations such as clipping or shifting may be performed on the results of interpolation filtering, motion vector derivation, or loop filtering.

[0121] Improving coding efficiency As mentioned above, image and video compression methods based on the variational autoencoder approach suffer from the lack of spatial adaptation and object segmentation targeting to capture real-world object boundaries. Therefore, content adaptability is limited. Furthermore, for certain types of video information, such as motion or residual information, sparse representation and coding are desirable to keep signaling overhead at a reasonable level.

[0122] Accordingly, some embodiments of this disclosure introduce segmentation information coding and feature map coding from different spatial resolution layers of the autoencoder to enable content adaptability and sparse signal representation and transmission.

[0123] In some exemplary implementations, connections are introduced between the encoder layers and decoder layers, excluding the low-resolution layers (latent space), which are transmitted in the bitstream. In some exemplary implementations, to conserve bandwidth, only a portion of the feature maps from different resolution layers are provided in the bitstream. For example, signal selection and signal supply logic is introduced to select, transmit, and use portions of the feature maps from different resolution layers. On the receiver side, tensor combinational logic is introduced to combine the output from the previous resolution layer with the information received from the bitstream corresponding to the current resolution layer.

[0124] Some detailed embodiments and examples relating to the encoder and decoder sides are provided below.

[0125] Encoding method and device According to one embodiment, a method is provided for encoding data for picture or video processing into a bitstream. Such a method includes the step of processing the data, which includes generating feature maps in a plurality of cascaded layers, each feature map including its own resolution, and at least two of the generated feature maps having different resolutions from one another.

[0126] In other words, the resolutions of two or more cascaded layers can be different from each other. Here, when we refer to the resolution of a layer, we mean the resolution of the feature map processed by that layer. In exemplary implementations, it is the resolution of the feature map output by the layer. A feature map containing a resolution means that at least a portion of the feature map has that resolution. In some implementations, the entire feature map may have the same resolution. The resolution of a feature map can be given, for example, by multiple feature map elements in the feature map. However, it may be more specifically defined by the number of feature map elements in one or more dimensions (e.g., x, y, or alternatively or additionally, the number of channels may be considered).

[0127] The term "layer" here refers to a processing layer. It does not necessarily have to be a trainable or trained layer with parameters (weights), like the layers of some neural networks mentioned above. Rather, a layer can represent a specific processing of a layer input to obtain a layer output. In some embodiments, a layer may be trained or trainable. Training here refers to machine learning or deep learning.

[0128] When referring to a cascaded layer, it means that the layers have a predetermined sequence, and the input to the first layer (in the given order) is processed sequentially by the first layer, and then by subsequent layers, in the given order. In other words, the output of layer j is the input to layer j+1, where j is an integer from 1 to the total number of cascaded layers. In certain non-limiting examples, layer j+1 contains (or has) the same resolution as or lower than layer j for all possible j values. In other words, the resolution of a layer does not increase with the sequence of cascading (processing order) (e.g., on the encoder side). However, it should be noted that this disclosure is not limited to such specific cascaded layers. In some embodiments, the layers of cascaded processing may also include layers that increase the resolution. In any case, there may be layers that do not change the resolution.

[0129] A low-resolution feature map may mean, for example, fewer feature elements per feature map. A high-resolution feature map may mean, for example, more feature elements per feature map.

[0130] The method further includes the step of generating a bitstream which involves selecting layers from among several layers that are different from the layer that generates the lowest resolution feature map, and inserting information related to the selected layer into the bitstream.

[0131] In other words, in addition to outputting the results of processing by all layers in the cascade to a bitstream (or instead), information is provided to another (selected) layer. There may be one or more selected layers. The information related to the selected layer can be any kind of information, such as the layer's output or segmentation information (discussed later) of part of the layer, or other information related to the feature maps processed by the layer and / or the processing performed by the layer. In other words, in some examples, the information may be the elements of the feature map and / or the positions of elements within the feature map (in the layer).

[0132] The input to the cascading process is data for picture or video processing. Such data may be related to predictive coding, such as inter-prediction or intra-prediction. It may be motion vectors, or other parameters of prediction such as prediction mode or reference picture or direction, or other parts of coding separate from prediction, such as transformation, filtering, entropy coding, or quantization. Bitstream generation may involve any transformation (binaryization) of values ​​to bits, including fixed codewords, variable-length codes, or arithmetic coding.

[0133] Here, a picture can be a still picture or a video picture. A picture refers to one or more samples, such as those captured by a camera or generated by computer graphics, for example. A picture may contain samples representing luminance levels in grayscale, or may have multiple channels, including one or more of the following: luminance channels, chrominance channels, depth channels, or other channels. The picture or video encoding may be either a hybrid coding (such as HEVC or VVC) or an autoencoder as described above.

[0134] Figure 8 is a flowchart illustrating the method described above. Therefore, this method includes a step 810 for processing input data. From the processed data, a portion is selected in a selection step 820 and included in a bitstream in a generation step 830. Not all data generated in the processing steps needs to be included in the bitstream.

[0135] According to an exemplary implementation, the processing further includes downsampling by one or more of the cascaded layers. An exemplary network 900 that implements (performs during operation) such processing is shown in Figure 9.

[0136] In particular, Figure 9 shows the input data for image or video processing 901 entering the network 900. The input data for image or video processing can be any type of data used for such processing, such as direct image (picture) or video samples, prediction modes, motion vectors, etc., as already mentioned above. The processing in Figure 9 applied to input 901 is carried out by multiple processing layers 911-913, each processing layer reducing the resolution of each motion vector array. That is, the cascaded layers 911-913 are downsampling layers. Note that when a layer is referred to as a downsampling layer, it performs downsampling. There may be embodiments in which the downsampling layers 911-913 perform downsampling as their sole task, and there may be embodiments in which the downsampling layers 911-913 do not perform downsampling as their sole task. Rather, the downsampling layers can generally also perform other types of processing.

[0137] As shown in Figure 9, the downsampling layers 911-913 also have additional select outputs connected to the signal selection logic 920, separate from the processed data inputs and outputs. Note that the term "logic" here refers to any circuit that implements a function (in this case, signal selection). The signal selection logic 920 selects the information contained in the bitstream 930 from the select outputs of any of the layers. In the example in Figure 9, each layer 911-913 downsamples the layer input. However, layers that do not apply downsampling may be added between the downsampling layers. For example, these layers can process the input by filtering or other operations.

[0138] In the example shown in Figure 9, the signal selection logic 920 selects information contained in the bitstream from the outputs of layers 911 to 913. The goal of this selection may be to select information relevant to reconstructing an image or video from multiple feature maps output by different layers. In other words, the downsampling layer and signal selection logic may be implemented as part of an encoder (picture or video encoder). For example, the encoder may be encoder 101 shown in Figure 3A, encoder 121 in Figure 3B, MV encoder net (part of end-to-end compression in Figure 5A), MV encoder in Figure 5B, or part of an encoder according to Figure 7A (e.g., loop filtering 220 or part of mode selection unit 260 or prediction units 244, 254).

[0139] Figure 9 further includes the decoder-side portion (which may be called the bulging path) including the signal supply logic 940 and the upsampling layers 951 to 953. The encoder-side input is the bitstream 930. Output 911 is, for example, the reconfigured input 901. The decoder side is described below with greater delay.

[0140] Downsampling can be performed, for example, through max pooling, average pooling, or any other operation that results in downsampling. Another example of such an operation involves convolution. Figure 10A shows an example of max pooling. In this example, each of the four elements (adjacent 2x2 squares) in array 1010 is grouped and used to determine one element in array 1020. Arrays 1020 and 1010 may correspond to feature maps in some embodiments of this disclosure. However, arrays may also correspond to parts of feature maps in these embodiments. Fields (elements) in arrays 1020 and 1010 may correspond to elements of feature maps. In this figure, feature map 1020 is determined by downsampling feature map 1010. The numbers in the fields of arrays 1010 and 1020 are illustrative. Instead of numbers, fields may include, for example, motion vectors. In the max pooling example shown in Figure 10A, the top-left four fields of array 1010 are grouped and the maximum value among them is selected. This group of values ​​determines the top-left field of array 1020 by assigning its maximum value to this field. In other words, the largest of the four top-left values ​​in array 1010 is inserted into the top-left field of array 1020.

[0141] Alternatively, some implementations may use min pooling. In min pooling, instead of selecting the field with the maximum value, the field with the minimum value is selected. However, these downsampling techniques are merely examples, and various downsampling strategies may be used in different embodiments. Some implementations may use different downsampling techniques in different layers, in different regions within the feature map, and / or for different types of input data.

[0142] In some implementations, downsampling is performed using mean pooling. In mean pooling, the average of a group of feature map elements is calculated and associated with the corresponding field in the feature map of the downsampled feature map.

[0143] An example of average pooling is shown in Figure 10B. In this example, the top-left feature map element of feature map 1050 is averaged, and the top-left element of feature map 1060 takes this average value. The same applies to the top-right, bottom-right, and bottom-left groups in Figure 10B.

[0144] In another embodiment, a convolution operation is used for downsampling in part or all of the layers. In the convolution, a filter kernel is applied to a group or block of elements in the input feature map. The kernel itself may be an array of elements the same size as the block of input elements, and each element of the kernel stores a weight for the filter operation. In downsampling, a sum is calculated of the elements from the input block, each weighted by the corresponding value taken from the kernel. If the weights of all elements in the kernel are fixed, such a convolution may correspond to the filter operation described above. For example, a convolution with a kernel having the same fixed weights and a stride of the kernel size corresponds to an average pooling operation. However, the stride of the convolution used in this embodiment may differ from the kernel size, and the weights may differ. In one example, the kernel weights may be such that some features in the input feature map are emphasized or distinguished from each other. Furthermore, the kernel weights may be learnable or pre-trained.

[0145] According to one embodiment, information related to a selected layer includes elements 1120 of the feature map of that layer. For example, the information can convey feature map information. Generally, the feature map may include any features related to the motion picture.

[0146] Figure 11 shows an exemplary implementation where the feature map 1110 is a dense optical flow of motion vectors having width W and height H. The motion segmentation net 1140 includes three downsampling layers (corresponding, for example, to the downsampling layers 911-913 in Figure 9) and a signal selection circuit (logic) 1100 (corresponding, for example, to the signal selection logic 920). Figure 11 shows an example of the outputs (L1-L3) of different layers in the erosion path on the right.

[0147] In this example, the output of each layer (L1-L3) is a feature map with progressively lower resolution. The input to L1 is a dense optical flow 1110. In this example, one element of the feature map output from L1 is determined from 16 (4x4) elements of the dense optical flow 1110. Each square in the L1 output (bottom right of Figure 11) corresponds to a motion vector obtained by downsampling (downspl4) from the 16 motion vectors of the dense optical flow. Such downsampling could be, for example, average pooling or another operation as described above. In this exemplary implementation, only a portion of the feature map L1 of that layer is included in the information 1120. A layer L1 is selected, and the portion corresponding to the four motion vectors (feature map elements) associated with the selected layer is signaled within the selected information 1120.

[0148] The output L1 of the first layer is then input to the second layer (downspl2). The feature map elements of the output L2 of the second layer are determined from the four elements of L1. However, in other examples, each element of a feature map with a lower resolution may also be determined by a group consisting of any other number of elements from a feature map with a next higher resolution. For example, the number of elements in the group that determines one element in the next layer may be any power of 2. In this example, the output L2 feature map corresponds to three motion vectors, which are also included in the selected information 1120, and therefore the second layer is also a selected layer. The third layer (downspl2) downsamples the output L2 from the second layer by 2 in each of the two dimensions. Thus, one feature map element of the output L3 of the third layer is obtained based on the four elements of L2. In feature map L3, the elements are not signaled, i.e., the third layer is not a selected layer in this example.

[0149] The signal selection module 1100 of the motion segmentation network 1140 selects the aforementioned motion vectors (elements of the feature maps from the outputs of the first and second layers) and provides them to the bitstream 1150. This provision may be simple binarization and may include, but is not required to include, entropy coding.

[0150] The element groups may be arranged in a square shape, as in the example in Figure 11. However, the groups may be arranged in any other shape, such as a rectangle, and the longer side of the rectangle may be oriented horizontally or vertically. These shapes are merely examples. Any shape can be used in the implementation. This shape can also be signaled within bitstream 1150. Signaling can be implemented by a map of flags indicating which feature elements belong to a shape and which do not. Alternatively, signaling can be done using a more abstract description of the shape.

[0151] In this exemplary implementation, feature map elements are grouped such that every element belongs to exactly one group of elements that determine one element of the feature map in the next layer. In other words, feature map element groups do not overlap, and only one group contributes to feature map elements in higher (later in the cascaded processing order) layers. However, it is conceivable that elements in one layer may contribute to two or more elements in the next layer. In other words, in processing 810, when a new layer output, e.g., layer output L2, is generated based on layer output L1 which has a higher resolution, a filtering operation may be used.

[0152] In this embodiment, selection 820 (for example, by signal selection 1100) selects elements to be included in the bitstream from a plurality of output feature maps (L1-L3). The selection may be implemented to minimize the amount of data required to signal the selected data while keeping the amount of information related to decoding as large as possible. For example, rate distortion optimization or other optimizations may be employed.

[0153] The above example illustrates processing with three layers. In general, the method is not limited to this. Any number of processing layers (one or more) can be employed. In other words, according to a more generalized example, the method includes obtaining data to be encoded. This may be a dense flow of motion vectors, as described above. However, the disclosure is not limited thereto, and other data may be processed instead of, or in addition to, motion vectors, such as prediction modes, prediction directions, filtering parameters, or even spatial picture information (samples) or depth information.

[0154] Processing of the data to be encoded 810 in this example involves processing by each layer j of multiple N cascaded layers. Processing by the j-th layer is: If -j=1, the data to be encoded is taken as the layer input; otherwise, the feature map processed by the (j-1)th layer is taken as the layer input (i.e., if the j-th layer is the layer currently being processed, then j-1 is the preceding layer), - Processing the acquired layer input, and processing includes downsampling. - This includes outputting a downsampled feature map.

[0155] In this example, j=1 is the highest-resolution layer among the N processing layers. Note that the input to this layer may be a dense optical flow (which can generally be considered a feature map). Therefore, in some specific embodiments, the j=1 layer may be an input layer. However, this is not always the case, as the N processing layers may be preceded by some preprocessing layers. Typically, it is characteristic of encoders that earlier processing layers have higher resolution than slower processing layers (shrinkage path). This can be reversed on the decoder side. Some of the processing layers may not change the resolution or may not further improve it, but the disclosure may still be applicable.

[0156] In the example described above, bitstream 1150 carries selected information 1120, which can be, for example, a motion vector or any other feature. In other words, bitstream 1150 carries feature map elements from at least one layer of the processing network (encoder-side processing network) that is not the output layer. In the embodiment of Figure 11, only a portion of the selected feature map is transmitted in the bitstream. This portion has one or more feature elements. Rules for determination may be defined to allow the decoder to determine which portion of the feature map is transmitted. In some embodiments, segmentation information may be transmitted in bitstream 1150 to constitute which portion of the feature map is transmitted. Such exemplary embodiments are described below. However, it should be noted that the embodiments described above are illustrative and, in general, such additional signaling is not necessary, as there may be rules for deriving information and relying on other known or signaled parameters.

[0157] In an exemplary embodiment relating to segmentation information, the information relating to the selected layer includes information 1130 indicating (in addition to or instead of the selected information 1120) from which layer and / or from which portion of the feature map of that layer the elements of the feature map of that layer were selected.

[0158] In the example shown in Figure 11, segmentation information is indicated by binary flags. For example, on the right, each low-resolution feature map or feature map portion is assigned either 0 or 1. For example, L3 is assigned zero (0) because it is not selected and no motion vectors (feature elements) are signaled for L3. Feature map L2 has four parts. Layer processing L2 is the selected layer. Three of the four feature map elements (motion vectors) are signaled, and correspondingly the flag is set to 1. The remaining one part of feature map L2 does not contain any motion vectors, and therefore the corresponding motion vectors are indicated by the L1 feature map, so the flag is set to 0. Since L1 is the first layer, it is implicit that the remaining motion vectors are provided in this layer. Note here that the binary flag takes a first value (e.g., 1) when the corresponding feature map portion is partial selection information, and a second value (e.g., 0) when the corresponding feature map portion is not partial selection information. Since this is a binary flag, it can only take one of these two values.

[0159] Such segmentation information can be provided in a bitstream. The left side of Figure 11 shows the processing of the segmentation information 1130. Note that the segmentation information 1130 can also be processed by layers of the motion segmentation net 1140. It may be processed in the same layer as the feature map or in a different layer. The segmentation information 1130 can also be interpreted as follows: One superpixel in the layer with the lowest resolution covers a 16x16 cell of the feature map obtained by downsampling the downspl4 of the dense optical flow 1110. The flag assigned to the superpixel covering the 16x16 cell is set to 0, which means that feature map elements (multiple) (in this case, motion vectors) are not shown for this layer (the layer is not selected). Thus, feature map elements may be shown in an area corresponding to a 16x16 cell in the next layer, represented by four equally sized superpixels, each covering an 8x8 feature element cell. Each of these four superpixels is associated with a flag. For superpixels associated with a flag having a value of 1, the feature map element (motion vector) is signaled. For superpixels with a flag set to 0, the motion vector is not signaled. Unsignaled motion vectors are signaled for layers that have superpixels covering 4x4 element cells.

[0160] More generally, a method for encoding data for picture / video decoding may further include selecting (segmentation) information for insertion into a bitstream. This information relates to a first region (superpixel) within a feature map processed by a layer j > 1. The first region corresponds to a region of a feature map or initial data that is encoded at a layer smaller than j and contains a plurality of elements. The method further includes excluding from the selection a region corresponding to the first region, from a selection in a feature map processed by a layer k (where k is an integer greater than or equal to 1 and k < j). Correspondence of regions between different layers, as used herein, means that the corresponding regions (superpixels) spatially cover the same feature elements (initial data elements) within the feature map (initial data) being encoded. In the example of FIG. 11, the initial data to be segmented is L1 data. However, the association may be made with reference to the dense optical flow 1110.

[0161] In the specific configuration shown in Figure 11, it is ensured that each feature element of the initial feature map (e.g., L1) is covered by only one superpixel in only one of the N layers. This configuration offers the advantage of efficiently coding the feature map as well as the segmentation information. The cascaded layer processing framework corresponds to the neural network processing framework and can thus be used to segment data and provide data for various segments with different resolutions. In particular, the advantage of downsampling in some of the layers may include reducing the amount of data required to signal the representation of the initial feature map. Specifically, in the example of signaling motion vectors, downsampling allows a group of similar motion vectors to be signaled by a single common motion vector. However, the prediction error resulting from grouping motion vectors needs to be small to achieve good interpretation. This may mean that different levels of grouping motion vectors for different areas of a picture may be optimal for achieving the desired prediction quality while simultaneously requiring a small amount of data to signal the motion vectors. This can be achieved by using multiple layers with different resolutions.

[0162] In one embodiment, where the feature map elements are motion vectors, the length and direction of the motion vectors may be averaged for downsampling purposes, and the averaged motion vectors are associated with the corresponding feature map elements of the downsampled feature map. In typical averaging, all elements in a group of elements corresponding to one element in the downsampled feature map have the same weight. This corresponds to applying a filter with equal weights to a group or block of elements to compute the downsampled feature map elements. However, in other implementations, such a filter may have different weights for different elements in the layer input. In other implementations, instead of computing the average of a group or block of elements in downsampling, the median of each group of elements may be computed.

[0163] In the example in Figure 11, the downsampling filter operation uses a filter with a 2x2 square shape for the input elements and calculates a filtered output that maps to one element in the downsampled feature map, depending on the selected filter operation. The filter operation uses a stride of 2, which is equal to the edge length or the square filter. This means that between the two filtering operations, the filter is moved by a stride equal to the size of the filter. As a result, in downsampling, the downsampled elements are calculated from non-overlapping blocks in the layer to which the downsampling filter is applied.

[0164] However, in some further possible embodiments, the stride may differ from the filter edge length. For example, the stride may be less than the filter edge length. As a result, the filter blocks used to determine the elements in the downsampled layer may overlap, meaning that one element from the downsampled feature map contributes to the calculation of two or more elements in the downsampled feature map.

[0165] Generally, the data associated with the selected layer includes indications of the location of feature map elements within the feature map of the selected layer. Here, similar to the concept in Figure 11 (feature maps L1-L3), the feature map of the selected layer refers to the output from the selected layer, i.e., the feature map processed by the selected layer.

[0166] For example, the locations of selected and unselected feature map elements are indicated by a number of binary flags based on the flag locations in the bitstream. In the above description with reference to Figure 11, the binary flags are included in the bitstream 1150 as segmentation information 1130. To enable the decoder to parse and correctly interpret the segmentation information, an assignment should be defined between the flags and the areas in the feature maps processed by the layers and / or layers. This can be done by defining the order in which the flags, known to both the encoder and the decoder, are binaryized.

[0167] The above example is provided for data used when encoding a picture / video that is a motion vector. However, this disclosure is not limited to such embodiments. In one embodiment, the data to be encoded includes image information and / or predicted residual information and / or predicted information, where image information means sample values ​​of the original image (or image to be coded). The sample values ​​may be samples of one or more colors or other channels.

[0168] Information about the selected layer does not necessarily have to be motion vectors or superpixel motion vectors. In addition or alternative to this, in some embodiments, the information includes prediction information. Prediction information may include a reference index and / or a prediction mode. For example, the reference index may indicate which particular picture from the reference picture set should be used for interpretation. The index may be relative to the current image in which the predicted current block is located. The prediction mode may indicate, for example, whether to use one or more reference frames and / or a combination of different predictions, such as combined intra-interpretation.

[0169] Nevertheless, efficient motion vector field coding and reconstruction can be achieved when the data to be encoded is a motion vector field. A corresponding general block scheme of a device capable of performing such encoding and decoding of a motion field is shown in Figure 12. On the encoding side, motion information is obtained using some motion estimation or optical flow estimation module (unit) 1210. The input to the motion vector (optical flow) estimation is the current picture and one or more reference pictures (stored in the reference picture buffer). In Figure 12, the picture is referred to as a "frame," a term sometimes used for pictures in video. The optical flow estimation unit 1210 outputs an optical flow 1215. In different implementations, the motion estimation unit may output motion information that already has a different spatial resolution, for example, for some N×N block or for each pixel of the original resolution, which may be referred to as a dense optical flow. The motion vector information is sent to the decoding side (embedded in bitstream 1250) and is intended to be used for motion compensation. To obtain a motion-compensated region, each pixel in the region should have a defined motion vector. Transmitting motion vector information for each pixel at the original resolution can be too costly. To reduce signaling overhead, a motion designation (or segmentation) module 1220 is used. The corresponding module 1270 on the decoding side performs a motion generation (densification) task to reconstruct the motion vector field 1275. The motion designation (or segmentation) module 1220 outputs motion information (e.g., motion vectors, and / or possibly a reference picture) and segmentation information. This information is added (encoded) to the bitstream.

[0170] In this embodiment, the motion segmentation unit 1220 and the motion generation unit 1270 include only a downsampling layer dwnspl and a corresponding upsampling layer upspl, as shown in Figure 13. Nearest-neighbor method can be used for downsampling and upsampling, and average pooling can be used for downsampling. Feature map data from layers of different spatial resolutions are selected by the encoder and transmitted in a bitstream as selected information 1120, along with segmentation information 1130 that instructs the decoder how to interpret and utilize the selected information 1120. The motion segmentation (sparsed) net 1220 is shown in Figure 13 as network 1310. Thus, a dense optical flow 1215 is inserted into the motion segmentation (sparsed) net 1310. The net 1310 includes three downsampling layers and signal selection logic 1320 that selects the information to be included in the bitstream 1350. Its function is similar to that already described with reference to the more general Figure 9.

[0171] In the embodiments described above, signaling information related to layers other than the output layer improves the system's scalability. Such information may be related to hidden layers. Embodiments and examples of utilizing the provided scalability and flexibility are presented below. In other words, some approaches are provided regarding how to select layers and how the information may appear.

[0172] Some embodiments of this specification describe an image or video compression system using an autoencoder architecture in which the encoding portion includes one or more dimensional (or spatial resolution) reduction steps (implemented by layers incorporating downsampling operations). Along with the reduction (encoding) side, the reconstruction (decoding) side is learned, where the autoencoder attempts to produce a representation from the reduced encoding that is as close as possible to its original input, which typically means one or more resolution increase steps on the decoding side (implemented by layers incorporating upsampling operations).

[0173] Hereafter, below the encoder, it means the encoding portion of an autoencoder that generates a latent signal representation contained in the bitstream. Such an encoder is, for example, 101 or 121 described above. Below the decoder, it means the generating portion of an autoencoder that perceives the latent signal representation obtained from the bitstream. Such a decoder is, for example, decoder 104 or 144 described above.

[0174] As already explained with reference to Figure 11, the encoder selects a portion (or more portions) of feature map information (selected information 1120) from layers of different spatial resolutions according to the signal selection logic 1100, and transmits the selected information 1120 in the bitstream 1150. The segmentation information 1130 indicates which layer the selected information was obtained from and which portion of the feature map of the corresponding layer it was obtained from.

[0175] According to one embodiment, the processing performed by layer j among the multiple N cascade layers is as follows: - Using the feature map elements output by the j-th layer, determine the first cost resulting from reconstructing a portion of the reconstructed picture, - Using the feature map elements output by the (j-1)th layer, determine the second cost resulting from reconstructing a portion of the picture, - If the first cost is higher than the second cost, select the (j-1)th layer and select information about the portion of the (j-1)th layer.

[0176] The decision of which layer to select may be based on or a function of distortion. For example, in motion vector field coding, the reconstructed picture (or picture portion) may be a motion-compensated picture (or picture portion).

[0177] In this exemplary implementation, to select the chosen information, the encoder includes a cost calculation unit (module) that estimates the cost of transmitting motion information from a specific resolution layer at a given location. The cost is calculated by combining an estimate of the amount of bits required to transmit the motion information, multiplied by the Lagrangian multiplier, and the distortion caused by motion compensation using the selected motion vector. In other words, according to one embodiment, rate-distortion optimization (RDO) is performed.

[0178] In other words, in some embodiments, the first and second costs include amounts of data and / or distortion. For example, the amount of data includes the amount of data required to transmit data related to the selected layer. This may be motion information or other information. It may also be or include the overhead caused by residual coding. Distortion is calculated by comparing the reconstructed picture with the target picture (the original picture being coded or a portion of such a picture). Note that RDO is only one possibility. This disclosure is not limited to such an approach. Furthermore, complexity or other factors may be included in the cost function.

[0179] Figure 14 shows the first part of the cost calculation. Specifically, the cost calculation (or estimation) unit 1400 obtains the optical flow L1 downsampled by the downsampling layer (downspl4) of the motion segmentation unit 1140. The cost calculation unit 1400 then upsamples the optical flow to its original resolution 1415, for example, in this case upsampled by 4 in each of the two directions (x and y). Next, motion compensation 1420 is performed using the upsampled motion vector output from 1410 and the reference picture(s) 1405 to obtain a motion-compensated frame(s) or a portion of the motion-compensated frame(s) 1420. Then, the distortion is calculated 1430 by comparing the motion-compensated picture(s) 1420 with the target picture 1408. The target picture 1408 may be, for example, the coded picture(s) (original picture). In some exemplary implementations, the comparison may be performed by calculating the mean squared error (MSE) or sum of absolute differences (SAD) between the target picture 1408 and the motion-compensated picture 1420. However, other types of measurements / metrics, such as more advanced metrics targeting subjective perception like MS-SSIM or VMAF, may be used as an alternative or in addition. The calculated strain 1430 is then provided to the cost calculation module 1460.

[0180] Furthermore, the rate estimation module 1440 calculates an estimate of the bit amount for each motion vector. The rate estimate may include not only the bits used to signal the motion vector, but also (in some embodiments) bits used to indicate segmentation information. The bit count thus obtained may be normalized, for example, per pixel (feature map element) 1450. The resulting rate (bit amount) is provided to the cost calculation module 1460. To obtain the rate (bit amount) estimate, the evaluation of the bit amount for each motion vector transmission is performed, for example, using the motion information coding module (e.g., by performing coding and recording the resulting bit amount), or in some simplified implementation, using the length of its x or y component motion vector as a rough estimate. Other estimation techniques may be applied. To take segmentation information into account, the segmentation information may be evaluated by the segmentation information coding module (e.g., by generating and coding the segmentation information and counting the resulting number of bits), or in a simpler implementation, by adding one bit to the total bit amount.

[0181] The next step in cost calculation in this example is cost calculation 1460, followed by downsampling 1470 (downspl 4) by 4 to the resolution of the corresponding downsampling layer of the motion segmentation unit 1100. In some cases, only one motion vector can be sent for each point (picture sample value). Therefore, the resulting cost tensor can have a corresponding size (dimension). Thus, the bit evaluation value can be normalized by the square of the downsampling filter shape (e.g., 4 × 4).

[0182] Next, using the Lagrange multiplier, the cost estimation unit 1460 is, for example, given by the formula Cost = D + λ * R, or Cost = R + β * D The cost is calculated using 1470, where D represents the distortion (calculated by 1430), R is the R-bit estimate (rate estimate output by 1440 or 1450), and λ and β are Lagrangian multipliers. Downsampling 1470 outputs the cost tensor 1480. As is known in the art, the Lagrangian multipliers as well as λ and β can be obtained empirically.

[0183] As a result, a tensor 1480 is obtained, which has cost estimates for each position in the feature map (in this case, the W × H position of the dense optical flow). Using successive averaging pooling and nearest neighbor upsampling results in averaging of motion vectors in an N × N (e.g., 4 × 4) region, where N × N is the average pooling filter shape and scaling factor for the upsampling operation. During upsampling using nearest neighbor, values ​​from the low-resolution layer are duplicated (repeated) at all points in the high-resolution layer corresponding to the filter shape. This corresponds to a translational motion model.

[0184] Various implementation forms of the cost selection unit are possible. For example, FIG. 15 shows another exemplary implementation form. In this example, unlike FIG. 14, the motion vector field obtained after downsampling the dense optical flow by 4 in each of the x and y dimensions at 1501 is not upsampled at 1415. Rather, it is provided directly for motion compensation 1510 and estimation rate 1540. Instead, the reference picture 1505 and the target picture 1505 can be downsampled at 1515, 1518 to the corresponding resolution before motion compensation 1510 and distortion evaluation 1530. This leads to excluding the step of initial motion field upsampling 1415 to the original resolution in FIG. 14, as well as excluding the final cost downsampling step 1470 in FIG. 14. Thereby, bit normalization 1450 is also not required. This implementation form may require less memory to store tensors during processing, but may give less accurate results. Note that in order to accelerate or reduce the complexity of RDO, it is conceivable to downsample the denser optical flow as well as the reference picture and the target picture more than what is done by L1. However, the accuracy of such RDO may be further reduced.

[0185] By applying the cost estimation units (1400, 1500) to each downsampling layer of the motion segmentation units (1220, 1310), costs are obtained using different levels of motion vector averaging (different spatial resolutions). As a next step, the signal selection logic 1100 uses the cost information from each downsampling layer to select motion information of different spatial resolutions. To achieve that the signal selection logic 1100 performs a pairwise comparison of costs from sequential (cascaded) downsampling layers, the signal selection logic 1100 selects the minimum cost at each spatial position and propagates it to the next downsampling layer (in the sequence of processing). FIG. 16 is a diagram showing an exemplary architecture of the signal selection unit 1600.

[0186] The dense optical flow 610 enters the three downsampling layers downspl4, downspl2, and downspl2, similar to that shown in FIG. 11. The signal selection logic 1600 in FIG. 16 is an exemplary implementation of the signal selection logic 1100 in FIG. 11. In particular, the LayerMv tensor 611 is a subsampled motion vector field (feature map) that enters the cost calculation unit 613. Also, the LayerMv tensor 611 enters the layer information selection unit 614 of the first layer. The layer information selection unit 614 provides the selected motion vector to the bitstream if there is a selected motion vector on this (first) layer. Its function will be further described below.

[0187] The cost calculation unit 613 calculates the cost, for example, as described for the cost calculation unit 1400 with reference to FIG. 14. This outputs a cost tensor, which is then downsampled by 2 to match the resolution at which the second layer operates. After processing by the second downsampling layer downspl 2, the LayerMV tensor 621 is provided to the next layer (the third layer) and the cost calculation unit 623 of the second layer. The cost calculation unit 623 operates in the same manner as the cost calculation unit 1400. Instead of upsampling / downsampling by 4 as in the example described with reference to FIG. 14, downsampling by 2 in each direction is applied, as will be apparent to those skilled in the art.

[0188] To perform a pairwise comparison of cost tensors from cost calculation units 613 and 623, the cost tensor from the previous (first) downsampling layer is downsampled (by 2) to the current resolution layer (second). Next, a pooling operation 625 is performed between the two cost tensors. In other words, the pooling operation 625 maintains the low cost in the element-wise cost tensor. The selection of the layer with the low cost is captured for each element index of the pooling operation result. For example, for a particular tensor element, if the cost of the first tensor is lower than the cost of the corresponding element of the second tensor, the index is equal to 0; otherwise, the index is equal to 1.

[0189] To ensure gradient propagation for training purposes, soft arg max can be used to obtain a pooled index with gradients. If gradient propagation is not required, normal pooling with indices can be used. As a result of the pooling operation 625, an index (LayerFlag tensor) indicating whether a motion vector from the current or previous resolution layer was selected, along with the motion vector (LayerMv tensor) from the corresponding downsampling layer of the motion segmentation unit, is transferred to the layer information selection unit 624 of the current (here, the second) layer. The best pooled cost tensor is propagated to the next downsampling level (downspl2), and then the operation is repeated for the third layer.

[0190] In particular, the LayerMv 621 output from the second layer is further downsampled (downspl 2) by the third layer, and the resulting motion vector field LayerMv 631 is provided to the cost calculation unit 633 of the third layer. The calculated cost tensor is propagated from the second layer and compared element-wise with the downsampled cost tensor provided by the MinCost pooling unit 625 635. After processing by the MinCost pooling unit 635, an index (LayerFlag tensor) indicating whether the motion vector from the current (third) or previous (second) resolution layer was selected, along with the motion vector (LayerMv tensor) from the corresponding downsampling layer of the motion segmentation unit, is transferred to the layer information selection unit 634 of the current (here, third) layer. In this example, only three layers exist for illustrative purposes. However, generally, there may be more than three layers, and the additional layers and the signal selection logic for these layers have similar functions to those shown for the second and third layers.

[0191] To collect the pooled information from each spatial resolution layer, the following process is performed in reverse order from the lowest resolution layer to the highest resolution layer, using layer information selection units 634, 624, and 614. First, the TakeFromPrev tensor, which is the same size as the lowest resolution layer (here, the third layer), is initialized with zeros 601. Then, the same operation is repeated for layers of different resolutions, as follows: The values ​​of the LayerFlag tensor (in the current layer) at positions where the value of the tensor (NOT TakeFromPrev) is equal to 1 are selected to be sent in the bitstream as segmentation information. The (NOT TakeFromPrev) tensor is the element-wise negation of the TakeFromPrev tensor. In the third (here, last) layer, the (NOT TakeFromPrev) tensor therefore has all values ​​set to 1 (negated zeros set by 601). Thus, the segmentation information 1130 (LayerFlag) of the last (here, third) layer is always sent.

[0192] The TakeFromCurrent tensor is obtained using the logical operation TakeFromCurrent=(NOT TakeFromPrev)AND LayerFlag. The flag in this tensor, TakeFromCurrent, indicates whether motion vector information is selected to be sent as a bitstream from the current resolution layer. The layer information selection unit (634,624,614) selects motion vector information from the corresponding downsampling layer of the motion segmentation unit by taking the value of the LayerMv tensor, and the value of the TakeFromCurrent tensor is equal to 1. This information is sent as selected information 1120 in the bitstream.

[0193] For the third processing layer (the first in reverse order) corresponding to the lowest resolution, TakeFromPrev is initialized to zero, then all flags are sent, and then all values ​​of (NOT TakeFromPrev) are equal to 1. For the last processing layer corresponding to the highest resolution layer, the LayerFlag flag does not need to be sent. For all positions where motion information was not selected from the previous layer, it is assumed that the position should be selected from the current layer or the next (highest resolution) layer.

[0194] Note that the cost calculation shown in Figure 16 is a parallelizable method that can be executed on a GPU / NPU. This method is also trainable, as it transfers gradients, making it usable in an end-to-end trainable video coding solution.

[0195] Note that the reverse processing order is similar to what the decoder does when analyzing segmentation information and motion vector information, as shown below when describing the decoder's function.

[0196] Another exemplary implementation of the signal selection logic 1700 is shown in Figure 17. Compared to Figure 16, the block diagram in Figure 17 introduces multiple coding options in the same resolution layer. This is indicated by options 1 to N within the layer 1 cost calculation unit 710. Note that, generally, one or more or all layers may contain more options. In other words, any of the cost calculation units 613, 623, and 633 can provide more options. These options can be one or more or all of the following, for example, different reference pictures used for motion estimation / compensation, single-hypothesis prediction, bi-hypothesis prediction, or multi-hypothesis prediction, different prediction methods, e.g., inter-frame prediction or intra-frame prediction, direct coding without prediction, multi-hypothesis prediction, presence or absence of residual information, and level of residual quantization. The cost is calculated for each coding option in the cost calculation unit 710. The best option is then selected using minimum cost pooling 720. The indicator (e.g., index) 705 of the best selected option is sent to the layer information selection module 730, and then, if the corresponding point in the current layer is selected to transmit information, the indicator BestOpt is transferred in the bitstream. In the given example, the options are shown only for the first layer, but it should be understood that similar option selection logic can be applied to other layers of different resolutions, or to all layers.

[0197] The approach described above is also suitable for segmenting and transferring logical information such as flags or switchers that control the picture reconstruction process, and for information that is intended to remain immutable after decoding and be kept the same as the encoded side. In other words, instead of the motion vector field (dense optical flow) processed in the exemplary implementation in Figure 16, any one or more other parameters, including segmentation, can be encoded in a similar manner. This could be one or more or all of the following: an indicator showing different reference pictures used for motion estimation / compensation, a single-hypothesis, bi-hypothesis, or multi-hypothesis prediction indicator, different prediction methods, e.g., inter-frame or intra-frame prediction, an indicator of direct coding without prediction, multi-hypothesis prediction, presence or absence of residual information, quantization level of residuals, parameters of an in-loop filter, etc.

[0198] Further modifications of the above embodiments and examples According to the first modification, the downsampling layer of the motion segmentation unit 1310 and / or the upsampling layer of the motion generation unit 1360 include a convolution operation. This is shown in Figure 18. As seen in Figure 18, compared to Figure 13, the downsampling layer "dwnspl" and the upsampling layer "upspl" are replaced by the downsampling convolutional layer "conv↓" in the motion segmentation unit 1810 and the upsampling convolutional layer "conv↑" in the motion generation unit 1860, respectively. Part of the advantage of the convolutional rescaling (downsampling, upsampling) layer is that it enables learnable downsampling and upsampling processes. This allows, for example, in the case of use for increasing the density of motion information, to find the optimal upsampling transform, and thus can reduce the blocking effect caused by motion compensation using block-averaged motion vector information, as described in the embodiments and examples above. The same is applicable to texture restoration processes, e.g., original image intensity values ​​or predictive residual generation processed by cascading layers.

[0199] In the example shown in Figure 18, all downsampling and upsampling layers are convolutional layers. Generally, this disclosure is not limited to such implementations. Generally, within a segmentation unit (1310, 1810) and / or a generation unit (1360, 1860), a subset (one or more) of downsampling and corresponding upsampling operations may be implemented as convolutions.

[0200] The examples described herein are provided for dense optical flow / motion vector field processing and therefore refer to motion segmentation units (1310, 1810) and / or motion generation units (1360, 1860), but it should be noted that this disclosure is not limited to such data / feature maps. Rather, instead of, or in addition to, motion vector fields in any of the embodiments and examples herein, any coding parameters, or textures such as image samples, or prediction residuals (prediction errors), etc., can be processed.

[0201] For example, it should be noted that an encoder that uses motion information averaging as downsampling can be used in combination with a decoder that has a convolutional upsampling layer. Furthermore, an encoder with a convolutional layer aimed at finding a better latent representation can be combined with a motion generation network (decoder) that implements a nearest-neighbor-based upsampling layer. Further combinations are possible. In other words, the upsampling and downsampling layers do not need to be of the same type.

[0202] According to a second modification, which can be combined with any of the embodiments and examples described above (as well as the first modification), the network processing includes one or more additional convolutional layers between the cascaded layers having different resolutions as described above. For example, the motion segmentation unit 1310 and / or motion generation unit 1360 further include one or more intermediate convolutional layers between some or all of the downsampling and upsampling layers. This is shown in Figure 19, which illustrates exemplary implementations of such motion segmentation network (module) 1910 and motion generation network (module) 1860. Note that the terms “module” and “unit” are used interchangeably herein to indicate functional units. In this particular embodiment, units 1910 and 1960 are more specifically network structures having multiple cascaded layers.

[0203] For example, motion segmentation unit 1910 has an additional convolutional layer "conv" before each downsampling layer "conv↓" (which may be other types of downsampling) compared to motion segmentation unit 1310. Furthermore, motion generation unit 1960 has an additional convolutional layer "conv" before each upsampling layer "conv↑" (which may be other types of upsampling) compared to motion generation unit 1360.

[0204] This can further reduce blocking artifacts caused by the sparsification of motion information and increase the generalization effect of finding a better latent representation. As with respect to the first modification, encoders and decoders from the different embodiments / modifications described above can be combined into a single compression system. For example, it is possible to have only an encoder with additional layers between downsampling layers and a decoder without such additional layers, and vice versa. Alternatively or additionally, it is possible to have different numbers and positions of such additional layers in the encoder and decoder.

[0205] According to the third modification, direct connections to the input and output signals are provided, as also shown in Figure 19. Note that the second and third modifications, although shown in the same figure, are independent. They may be applied together or separately to the embodiments and examples described above, as well as to other modifications. Direct connections are indicated by dashed lines.

[0206] In some embodiments, in addition to bottleneck information from the autoencoder's latent representation (output of the lowest resolution layer), information from higher resolution layers is added to the bitstream. To optimize signaling overhead, only a portion of the information from different resolution layers is inserted into the bitstream and controlled by signal selection logic. On the receiving (decoder) side, corresponding signal supply logic supplies information from the bitstream to layers of different spatial resolutions, as will be described in more detail below. Furthermore, information from the input signal prior to the downsampling layer can be added to the bitstream, thereby further increasing variability and flexibility. For example, the coding may be aligned with real-world object boundaries and segments with high spatial resolution, tailored to the characteristics of a particular sequence.

[0207] According to the fourth modification, the shapes of the downsampling and upsampling filters can have shapes other than squares, such as rectangles with horizontal or vertical orientations, asymmetric shapes, or any other shape by employing a mask operation. This allows for further increases in the variability of the segmentation process to better capture actual object boundaries. This modification is shown in Figure 20. In the motion segmentation unit 2010, a first downsampling layer, which may be the same as in any of the preceding embodiments, is followed by two further downsampling layers employing filters selected from a set of filter shapes. This modification is not limited to processing motion vector information.

[0208] Generally, in downsampling by a layer, it is applied to use a first filter to downsample an input feature map to obtain a first feature map and use a second filter to downsample the input feature map to obtain a second feature map. Cost calculation includes determining a third cost resulting from reconstructing a part of a reconstructed picture using the first feature map and determining a fourth cost resulting from reconstructing a part of the reconstructed picture using the second feature map. And in the selection, if the third cost is smaller than the fourth cost, the first feature map is selected, and if the third cost is larger than the fourth cost, the second feature map is selected. In this example, the selection was out of two filters. However, the present disclosure is not limited to two filters. Rather, for example, the cost of all selectable filters can be estimated and the filter that minimizes the cost can be selected, so that the selection from a predetermined number of filters may be performed in a similar manner.

[0209] The shapes of the first filter and the second filter can be any of a square, a horizontally oriented rectangle, and a vertically oriented rectangle. However, the present disclosure is not limited to these shapes. Generally, any filter shape can be designed. The filter can further include a filter defined in any desired shape. Such a shape can be indicated by obtaining a mask, the mask is composed of flags, the mask represents any filter shape, and one of the first filter and the second filter (generally, any of the selectable filters from the filter set) has any filter shape.

[0210] In exemplary implementations, to introduce variability, the encoder further includes pooling cost tensors obtained with the help of filters having different shapes. The index of the selected filter shape is signaled in the bitstream as part of the segmentation information, similar to how it is described above for the motion vector. For example, in the case of a selection between a horizontally oriented rectangular shape and a vertically oriented rectangular shape, a corresponding flag may be signaled in the bitstream. For example, the method of selecting multiple encoding options described with reference to Figure 17 may be used for selecting different filter shapes at the same resolution layer.

[0211] According to the fifth modification, motion models from a predetermined set of different motion models may be selected in the same resolution layer. In the previous embodiments, specific cases of downsampling filters and / or upsampling filters were described. In such cases, motion information may be averaged across square blocks representing translational motion models. In this fifth modification, other motion models may be employed in addition to translational motion modes. Such other motion models are as follows: - Affine motion model, - Higher-order motion models, - CNN layers specifically trained to represent specific motion models, such as zoom, rotation, affine, perspective, etc. It may include one or more of these.

[0212] In an exemplary implementation of the fifth variation, the autoencoder further includes a set of CNN layers and / or “handcrafted” layers representing a model other than the translational motion model. Such an autoencoder (and decoder) is shown in Figure 21. In Figure 21, layers containing a set of filters, indicated as a “conv fit set,” are provided on both the encoder and decoder sides.

[0213] For example, in each spatial layer, the encoder selects one or more appropriate filters from a set of filters that correspond to a given motion model and inserts the instruction into the bitstream. On the receiving end, the signal supply logic interprets the indicator and performs convolution in a specific layer using the corresponding filter(s) from the set.

[0214] The examples of the methods described above use motion information, particularly motion vectors, as exemplary inputs for encoding. It should be noted again that these methods are also applicable to compressing different types of image or video information, such as direct image sample values, predicted residual information, and intra-frame and inter-frame predicted parameters.

[0215] According to the sixth modification, the RDO illustrated above with reference to Figure 16 or Figure 17 may be applied to a conventional block-based codec.

[0216] Conventional video coding methods, such as state-of-the-art video coding standards like AVC, HEVC, VVC, or EVC, use a block-based coding concept, where a picture is recursively divided into square or rectangular blocks. For these blocks, signal reconstruction parameters are estimated or evaluated by the encoder and sent to the decoder in a bitstream. Typically, the encoder aims to find the optimal reconstruction parameters for the set of blocks representing the picture in terms of rate distortion cost, attempting to maximize reconstruction quality (i.e., minimize distortion of the original picture) and minimize the amount of bits required to send parameters for the reconstruction process. This task of parameter selection (or coding mode determination) is a complex and resource-intensive task, which is also a major cause of encoder complexity. Due to processing time constraints, for example in real-time applications, the encoder may sacrifice the quality of mode determination, which affects the quality of the reconstructed signal. Optimizing the mode determination process is always a desirable technical improvement.

[0217] One of the coding mode determinations is whether or not the current block (or coding unit (CU)) should be divided into multiple blocks according to the division method.

[0218] According to the sixth modification, the motion segmentation unit 1310 (or 1810) described above is adapted for segmentation mode determination based on cost minimization (e.g., rate-distortion optimization criteria). Figure 22 shows an example of such optimization.

[0219] As seen in Figure 22, instead of downsampling layers, a block partition structure is used to represent information at different spatial resolutions. For each block of a given size N×N (considering square blocks) in the picture or a portion of the picture, the cost calculation unit calculates a distortion tensor and further downsamples the resolution to 1 / 16th (to match the original resolution). In the given example in Figure 22, the first block size is 16×16, and for example, downsampling is performed by an average pooling operation to obtain a tensor, whose element represents the average distortion for each 16×16 block. In the first layer, the picture is divided into 16×16 blocks 2201 at the initial highest resolution. In the second layer, the resolution is reduced so that the block size 2202 in the picture is 32×32 (corresponding to combining the four blocks from the previous layer). In the third layer, the resolution is reduced again so that the block size 2203 is 64×64 (corresponding to combining the four blocks from the previous layer). It should be noted that the combination of the four blocks from the previous layer in this case can be considered a subsampling of block-related information. This is because, in the first layer, block-related information is provided for each 16x16 block, whereas in the second layer, block-related information is provided only for 32x32 blocks, i.e., four times fewer parameters are provided. Similarly, in the third layer, block-related information is provided only for 64x64 blocks, i.e., four times fewer parameters are provided than in the second layer and sixteen times fewer parameters than in the first layer.

[0220] In this context, block-related information includes any information coded per block, such as prediction mode-specific information like prediction mode, motion vector, prediction direction, and reference picture, as well as filtering parameters, quantization parameters, transformation parameters, or other settings that may vary at the block (coding unit) level.

[0221] Then, the cost calculation units 2211, 2212, and 2213 of the first, second, and third layers respectively calculate the cost based on the block reconstruction parameters for block sizes 2201, 2202, and 2203, and the input picture of size W × H.

[0222] The output cost tensor is obtained as the average distortion per block, combined with an estimate of the bits required to transmit the coding parameters for the N×N blocks (e.g., in the first layer 16×16) using the Lagrange multiplier. An exemplary structure of the cost calculation unit 2300 for N×N blocks (which may correspond to each or any of the cost calculation units 2211, 2212, and 2213) is shown in Figure 23.

[0223] Figure 23 shows an exemplary block diagram of the cost calculation unit 2300 for a typical N×N block size 230x. The cost calculation unit 2300 retrieves block reconstruction parameters (block-related parameters) associated with the N×N blocks 2310. This retrieval may correspond to fetching parameters (parameter values) from memory or other sources. For example, block-related parameters may be a specific prediction mode, such as inter-prediction mode. In unit 2310, the block reconstruction parameters are retrieved, and in the reconstruction unit 2320, a portion of the picture is reconstructed using these parameters (in this example, all blocks are reconstructed using inter-prediction mode). Next, the distortion calculation unit 2330 calculates the distortion of the reconstructed portion of the picture by comparing it with the corresponding portion of the target picture, which may be the original picture being encoded. Since the distortion may be calculated per sample, downsampling of the distortion (to one value per N×N block) 2340 may be performed to obtain it on a block basis. In the lower branch, the rate or number of bits required to code the picture is estimated 2360. In particular, the bit estimation unit 2360 can estimate the number of bits signaled per N×N block. For example, it may calculate the number of bits per block required for the interprediction mode. Once the estimated distortion and bit amount (or rate) are obtained, the cost can be calculated, for example, based on the Lagrangian optimization described above. The output is the cost tensor.

[0224] Throughout this specification, it should be noted that when only 2D images of samples, such as grayscale images, are observed, the term "tensor" as used herein can refer to a matrix. However, multiple channels, such as color or depth channels for a picture, may exist so that the output can have more dimensions. General feature maps can also be provided in more dimensions than two or three.

[0225] The same cost evaluation procedure is performed for the first layer (with a 16x16 block granularity) and for the next level of the quadtree split into blocks of size 32x32 samples. To determine whether it is better to use one 32x32 block or four 16x16 blocks for the reconstruction parameters (block-related parameters), the cost tensor evaluated for the 16x16 blocks is downsampled by 2 (see Figure 22). The minimum cost pooling operation 2222 then provides the best decision for each 32x32 block. The index of the pooled costs is passed to the layer information selection unit 2232 to be sent in a bitstream as split_flags. The reconstruction parameters blk_rec_params for the best selected block according to the pooled index are also passed to the layer information selection unit 2231. The pooled cost tensor is further passed to the next quadtree aggregation level, i.e., MinCost pooling 2223, for blocks of size 64x64 (downsampled by 2). The MinCost pooling unit 2223 also receives the cost calculated for the 64x64 block resolution 2203 in the cost calculation unit 2213. This passes the index of the pooled cost as split_flags to the layer information selection unit 2233, as shown in the bitstream. It also passes the reconfiguration parameters blk_rec_params for the best selected block according to the pooled index to the layer information selection unit 2233.

[0226] To collect pooled information from each block aggregation level, layer information selection units 2233, 2232, and 2231 are used in the manner described above with reference to Figure 16, and processing is performed in reverse order from the highest (in this example) aggregation level (64 × 64 samples) to the lowest (in this example) aggregation level (16 × 16 samples).

[0227] The result is a bitstream encoding the quadtree partition obtained by optimization, along with the encoded values ​​of the acquired partitions (blocks) and, if applicable, further coding parameters. The partition flags of the block partitions can be determined using the method described above. Conventional methods based on the evaluation of each or some of the possible coding modes can be used to obtain the reconstruction parameters for each block.

[0228] Figure 24 shows an example of the seventh modification. The seventh modification is an extension of the sixth modification described above with reference to Figures 22 and 23. The seventh modification represents a way in which the evaluation of coding modes is incorporated into the design. In particular, as can be seen from the figure, the cost calculation unit 710 can evaluate N options. Note that the term "N" here is a placeholder for some integer. The "N" indicating the number of options does not have to be the same as the "N" in "N × N" which indicates a typical block size. In the cost calculation unit 710, for the same level of block partitioning, for example for a block of size 16 × 16 samples (as in the first layer), the encoder iterates over all possible (or a limited set of them) coding modes for each block.

[0229] Consider an encoder having N options for coding each 16x16 block, denoted as blk_rec_params 0, blk_rec_params 1, ..., blk_rec_params N. The parameter combination blk_rec_params k (where k is an integer from 0 to N) may be a combination of, for example, some prediction modes (e.g., from inter and intra), some transformations (e.g., from DCT and KLT), some filtering order or set of filter coefficients (of predefined filters). In some implementations, blk_rec_params k may be the value of a single parameter k if only one parameter is optimized. As will be apparent to those skilled in the art, any one or more parameters can be optimized by checking their usage cost.

[0230] For each given set of block reconstruction parameters (blk_rec_params_k), the cost calculation unit 2410 computes a tensor representing the cost of each block. Then, using minimum cost pooling 2420, the best coding mode for each block is selected and forwarded to the layer information selection unit 2430. The best pooled cost tensor is further downsampled by 2x and forwarded to the next quadtree aggregation level (in this case, the second layer corresponding to an aggregation with a block size of 32x32). Then, a splitting (partitioning) decision is made, similar to the sixth modification described above. In Figure 24, options 0..N are evaluated only at the first layer (aggregation level 16x16). However, this disclosure is not limited to this approach. Rather, the evaluation of options 0..N may be performed at each aggregation level.

[0231] For example, at the next level of the quadtree aggregation (32x32, 64x64), the encoder evaluates (by calculating the cost in each cost unit) and pools the best coding modes for each block (by each MinCost pooling unit) (not shown in the diagram for clarity), which are then compared to the previous aggregation level. The decision regarding the best mode and the corresponding reconfiguration parameters set accordingly is provided to a layer information selection unit (such as layer information selection unit 2430 shown for the first layer). To collect the pooled information from each block aggregation level, the layer information selection unit is used to process in reverse order from the upper aggregation level (64x64) to the lower aggregation level (16x16), as described in the sixth variation.

[0232] Different block shapes can be used to represent more advanced partitioning methods such as binary trees, ternary trees, asymmetric and geometric partitions. Figure 25 illustrates such partitions of blocks. In other words, optimization is not necessarily performed only across different block sizes, but may also be performed across different partitioning types (e.g., by corresponding options). Figure 25 shows the following example: -Quadrarrow partitioning 2510. In quadrarrow partitioning, a block is divided into four blocks of the same size. -(Symmetric) Binary Tree Partition 2520. In a symmetric binary partition, a block is divided into two blocks of the same size. The partition can be vertical or horizontal. Vertical or horizontal is an additional parameter of the partition. - Asymmetric binary partitioning 2530. In asymmetric binary partitioning, a block is divided into two blocks of different sizes. The size ratio may be fixed (to save overhead caused by signaling) or variable (in which case some ratio options may also be optimized, i.e., configurable). -Triangular tree partitioning 2540. In a three-valued partition, a block is divided into three partitions by two vertical lines or two horizontal lines. Vertical or horizontal are additional parameters of the partition.

[0233] This disclosure is not limited to these exemplary partitioning modes. Triangular partitions or any other type of partitioning may be employed.

[0234] In the seventh variation, a hybrid architecture applicable to common video coding standards is supported and empowered by a powerful (neural) network-based approach. The technical advantages of the described method are that it can provide a highly parallelizable GPU / NPU-friendly scheme that can enable faster computations for the mode decision process. It can enable global picture optimization because multiple blocks are considered at the same decision level, and incorporate a learnable portion to speed up the decision, for example, to evaluate the amount of bits required for reconstruction parameter coding.

[0235] In summary, processing with a cascaded layer structure according to the sixth or seventh variation involves processing data about the same picture segmented (i.e., divided / partitioned) into blocks having different block sizes and / or shapes in different layers. Selecting layers involves selecting layers based on the cost calculated for a given set of coding modes.

[0236] In other words, different layers can process picture data with different block sizes. Therefore, a cascaded layer includes at least two layers that process different block sizes from one another. When we refer to "block" here, what we mean is a unit, i.e., the part of the picture on which coding is performed. Blocks are sometimes also called coding units or processing units.

[0237] A predetermined set of coding modes corresponds to a combination of coding parameter values. Different block sizes may be evaluated in a single set of coding modes (a combination of values ​​for one or more coding parameters). Alternatively, the evaluation may include various combinations of block size and partition shape (as shown in Figure 25). However, the disclosure is not limited thereto, and there may be multiple predetermined sets of coding modes (combinations of coding parameter values) for each block, which may further include coding modes such as intra / inter prediction type, intra prediction mode, residual skip, and residual data, for example, as specifically mentioned in the seventh modification.

[0238] For example, the process involves determining the cost for different sets of coding modes (combinations of coding parameter values) for at least one layer, and selecting one of the sets of coding modes based on the determined cost. Figure 24 shows the case where only the first layer performs such selection. However, it is not limited to this. At high speed, each cost calculation unit may have the same structure as the first cost calculation unit 2410, including 0..N options. This is not shown in the figure for the sake of simplicity.

[0239] As mentioned above, this is a GPU-friendly RDO that can be performed by a codec and selects the best coding mode for each block. In Figure 24, the input image (picture) is identical in each layer. However, the coding of the picture (for cost calculation) is performed in each layer with different block sizes. Apart from the block size, for one or more of the block sizes, further coding parameters may be tested and selected based on the RDO.

[0240] In particular, in these variations, the data instructions related to the selected layer include a selected set of coding modes (e.g., blk_rec_params).

[0241] In summary, in some embodiments, an encoder is provided whose structure corresponds to a neural network autoencoder for coding video or image information. Such an encoder may be configured to analyze input image or video information by a neural network including layers of different spatial resolutions, to transfer a latent representation corresponding to the lowest resolution layer output in the bitstream, and to transfer outputs other than the lowest resolution layer in the bitstream.

[0242] decryption The encoder described above provides a bitstream containing selected layer feature data and / or segmentation information. Correspondingly, the decoder processes the data received from the bitstream in multiple layers. Furthermore, the selected layer receives additional (direct) input from the bitstream. This input may be some feature data information and / or segmentation information.

[0243] Accordingly, embodiments focusing on information related to the selected layer, which is feature data, are described below. Other embodiments described focus on information related to the selected layer, which is segmentation information. There are also hybrid embodiments in which the bitstream carries both feature data and segmentation information, and the layers process both feature data and segmentation information.

[0244] As a simple example, a decoder for a neural network autoencoder may be provided for video or image information coding. The decoder may be configured to read latent representations corresponding to low-resolution layer inputs from a bitstream, obtain layer input information based on the corresponding information read from the bitstream for the layers other than the low-resolution layer(s), obtain combined inputs for the layers based on the layer information obtained from the bitstream and the output from the previous layer, feed the combined inputs to the layers, and synthesize an image based on the layer outputs.

[0245] Here, the term "low resolution" refers to a layer that processes feature maps with low resolution, such as a latent space feature map given from a bitstream. Low resolution may actually be the lowest resolution of the network.

[0246] The decoder may be further configured to obtain segmentation information based on corresponding information read from the bitstream, and to obtain combined inputs for the layer based on the segmentation information. The segmentation information may be a quadtree, a dual (binary)tree, a ternary tree data structure, or a combination thereof. The layer input information may correspond to, for example, motion information, image information, and / or predicted residual information.

[0247] In some examples, information obtained from the bitstream corresponding to the layer input information is decoded using a super-prior neural network. Information obtained from the bitstream corresponding to the segmentation information can also be decoded using a super-prior neural network.

[0248] Decoders can be readily applied to decoding motion vectors (e.g., motion vector fields or optical flows). Some of these motion vectors may be similar or correlated. For example, in a video showing an object moving across a certain background, there may be two groups of similar motion vectors. The first group of motion vectors may be used in predicting the pixels representing the object, and the second group may be used in predicting the background pixels. Therefore, instead of signaling all motion vectors in the encoded data, it may be beneficial to signal groups of motion vectors to reduce the amount of data representing the encoded video. This may allow signaling representations of motion vector fields that require less data.

[0249] Figure 9 shows the bitstream 930, which is received by the decoder and generated by the encoder as described above. On the decoding side, the decoder portion of the system 900 includes, in some embodiments, a signal supply logic 940 that interprets segmentation information obtained from the bitstream 930. According to the segmentation information, the signal supply logic 940 identifies a specific (selected) layer, spatial size (resolution), and location for the portion of the feature map where the corresponding selected information (also obtained from the bitstream) should be placed.

[0250] It should be noted that in some embodiments, segmentation information is not necessarily processed by the cascade network. It may be provided independently or derived from other parameters in the bitstream. In other embodiments, feature data is not necessarily processed within the cascade network, but segmentation information is. Therefore, the two sections, “Decoding with Feature Information” and “Decoding with Segmentation Information,” describe examples of such embodiments, as well as combinations of such embodiments.

[0251] For embodiments of both sections, it should be noted that the encoder-side modifications described above (1 to 7) are applied correspondingly to the decoder-side. For clarity, additional features of the modifications are not copied in both sections. However, as will be apparent to those skilled in the art, they can be applied alternatively or in combination to the decoding approaches of both sections.

[0252] Decoding using feature information In this embodiment, a method for decoding data from a bitstream for picture or video processing is provided, as shown in Figure 33. Correspondingly, an apparatus for decoding data from a bitstream for picture or video processing is provided. The apparatus may include processing circuitry configured to perform the steps of the method.

[0253] The method includes obtaining two or more sets of feature map elements from a bitstream, where each set of feature map elements is associated with a feature map. The acquisition can be performed by parsing the bitstream. In some exemplary implementations, the bitstream parsing may also include entropy decoding. This disclosure is not limited to any particular method of obtaining data from a bitstream.

[0254] The method further includes step 3320, which inputs each of two or more sets of feature map elements into two or more feature map processing layers among multiple cascade layers.

[0255] Cascaded layers may form part of a processing network. In this disclosure, the term “cascaded” means that the output of one layer is later processed by another layer. Cascaded layers do not need to be directly adjacent (the output of one cascaded layer may directly enter the input of a second cascaded layer). Referring to Figure 9, data from bitstream 930 is input to signal supply logic 940, which supplies sets of feature map elements to the appropriate layers (indicated by arrows) 953, 952, and / or 951. For example, a first set of feature elements is inserted into the first layer 953 (the beginning of the processing sequence), and a second set of feature elements is inserted into the third layer 951. The set does not need to be inserted into the second layer. The number of layers and their positions (in the processing sequence) may vary, and this disclosure is not limited to any particular set.

[0256] The method further includes obtaining the decoded data for picture or video processing as a result of processing by multiple cascaded layers 3330. For example, the first set is a set of latent feature map elements processed by all layers of the network. The second set is an additional set provided to another layer. Referring to Figure 9, the decoded data 911 is obtained after the first set has been processed by three layers 953, 952, and 951 (in that order).

[0257] In a typical implementation, feature maps are processed in each of two or more feature map processing layers, and the feature maps processed in each of these layers have different resolutions. For example, the first feature map processed by the first layer has a different resolution than the second feature map processed by the second layer.

[0258] In particular, the processing of feature maps in two or more feature map processing layers includes upsampling. Figure 9 shows a network in which the decoding unit includes three (directly) cascaded upsampling layers 953, 952, and 951.

[0259] In an exemplary implementation, the decoder includes only upsampling layers of different spatial resolutions, and a nearest neighbor approach is used for upsampling. The nearest neighbor approach repeats the low-resolution values ​​in the high-resolution areas corresponding to a given shape. For example, if one low-resolution element corresponds to four high-resolution elements, the value of one element is repeated four times in the high-resolution region. In this case, the term "corresponding" means describing the same area in the highest-resolution data (initial feature map, initial data). Such a method of upsampling allows information to be transmitted from the low-resolution layer to the high-resolution layer without modification, which may be suitable for some kind of data, such as logical flags or indicator information, or information that is desired to remain the same as it was acquired on the encoder side without modification by some convolutional layers. Examples of such data include prediction information, such as motion information which may include motion vectors estimated on the encoder side, a reference index indicating which particular picture from a reference picture set should be used, a prediction mode indicating whether a single reference frame or multiple reference frames should be used, or a combination of different predictions such as combined intra-interpretation, and the presence or absence of residual information.

[0260] However, this disclosure is not limited to upsampling performed by the nearest neighbor approach. Alternatively, upsampling may be performed by applying some interpolation or extrapolation, or by applying convolution, etc. These approaches may be particularly suitable for upsampling data that is expected to have smooth characteristics, such as motion vectors or residuals or other sample-related data.

[0261] In Figure 9, the encoders (e.g., symbols 911-920) and decoders (e.g., symbols 940-951) have correspondingly the same number of downsampling and upsampling layers, where the nearest neighbor method may be used for upsampling and average pooling may be used for downsampling. The shape and size of the pooling layers are matched to the scale factor of the upsampling layers. In some other possible implementations, a different method of pooling, such as max pooling, may be used.

[0262] As already illustrated in several encoder embodiments, the data for picture or video processing may include a motion vector field. For example, Figure 12 shows the encoder and decoder side. On the decoder side, the bitstream 1250 is parsed and motion information 1260 (which may include segmentation information as described below) is obtained from it. The obtained motion information is then provided to the motion generation network 1270, which can increase the resolution of the motion information, i.e., increase the density of the motion information. The reconstructed motion vector field (e.g., dense optical flow) 1275 is then provided to the motion compensation unit 1280, which uses the reconstructed motion vector field to obtain predicted picture / video data based on a reference frame(s) and reconstructs motion-compensated frames based on it (for example, by adding the decoded residuals, as illustrated in the decoder portion of the encoder in Figure 5A or the reconstruction unit 314 in Figure 7B).

[0263] Figure 13 also shows the decoder-side motion generation (densification) network 1360. Network 1360 includes a signal supply logic 1370 that is functionally similar to the signal supply logic 940 in Figure 9, and three upsampling (processing) layers. The main difference from the embodiments described above with reference to Figure 9 is that in Figure 13, network 1360 is specialized for motion vector information processing that outputs a motion vector field.

[0264] As described above, according to one embodiment, the method further includes obtaining segmentation information relating to two or more layers from a bitstream. Subsequently, obtaining feature map elements from the bitstream is based on the segmentation information. Inputting the set of feature map elements into two or more feature map processing layers is based on the segmentation information. Some detailed examples of the use of segmentation information in analysis and processing are provided below in the section on decoding using segmentation information. For example, Figures 28 and 29 provide very specific (simply illustrative) layer processing options.

[0265] In some embodiments, the cascading layers further include multiple segmentation information processing layers. The method further includes processing segmentation information in the multiple segmentation information processing layers. For example, processing segmentation information in at least one of the multiple segmentation information processing layers includes upsampling. Such upsampling of segmentation information and / or feature maps includes nearest neighbor upsampling in some embodiments. Generally, upsampling applied to feature map information may differ from upsampling applied to segmentation information. Furthermore, since upsampling within the same network may differ, a single network (segmentation information processing or feature map processing) may include different types of upsampling layers. Such examples are shown, for example, in Figure 20 or Figure 21. Note that upsampling types other than nearest neighbor may include some interpolation approaches, such as polynomial approaches, e.g., bilinear, cubic, etc.

[0266] According to an exemplary implementation, the upsampling of segmentation information and / or feature maps involves (transposed) convolution. This corresponds to the first modification described above for the encoder. Figure 18 shows a motion generation unit 1869 on the decoder side that includes a convolution operation "conv↑" instead of nearest neighbor upsampling. This enables a learnable upsampling process, which, for example, when used for increasing the density of motion information, allows for finding the optimal upsampling transform and, as described above with reference to the encoder, can reduce the blocking effect caused by motion compensation by using block-averaged motion vector information. The same can be applied, for example, to a texture restoration process for generating original image intensity values ​​or predicted residuals. The motion generation unit 1869 also includes a signal supply logic that functionally corresponds to the signal supply logic 940 in Figure 9 or the signal supply logic 1370 in Figure 13.

[0267] Figure 30 shows a block diagram of exemplary decoder side-layer processing according to a first modification. In particular, the bitstream 3030 is analyzed, and the signal supply logic 3040 (functionally corresponding to signal supply logic 940 or 1370) provides a selection instruction to the convolution upsampling filter 300. In some embodiments, the convolution filter can be selected from a set of N filters (indicated as filters 1 to N). Filter selection indicates the selected filter and can be obtained based on information analyzed from the bitstream. Instructions for the selected filter can be given (generated and inserted into the bitstream) by the encoder based on an optimization approach such as RDO. In particular, RDO as illustrated in Figure 17 or Figure 24 may be applied (treating filter size / shape / order as one of the options, i.e., as the coding parameters to be optimized). However, the disclosure is not limited thereto, and in general, filters can be derived based on other signaled parameters (such as coding mode, interpolation direction, etc.).

[0268] In summary, the signal supply logic unit controls the inputs of different layers having different filter shapes and selectively bypasses the layer outputs to the next layer according to segmentation and motion information obtained from the bitstream. The convolutional filter unit 3000 corresponds to a convolution by one layer. Multiple such convolutional swamp filters can be cascaded as shown in Figure 18. It should be noted that this disclosure is not limited to variable or trainable filter configurations. In general, convolutional upsampling can also be performed using fixed convolution operations.

[0269] Aspects of this embodiment can be combined with aspects of other embodiments. For example, an encoder having motion information averaging in a downsampling layer can be used in combination with a decoder including a convolutional upsampling layer. An encoder having a convolutional layer aimed at finding a better latent representation can be combined with a motion generation network including a nearest-neighbor-based upsampling layer. Further combinations are possible. In other words, the implementation of the encoder and decoder does not need to be symmetrical.

[0270] Figure 32A shows two examples of reconstruction applying the nearest neighbor approach. In particular, Example 1 shows the case where the flag value of the segmentation information in the lowest resolution layer is set to (1). Correspondingly, the motion information shows one motion vector. Since the motion vector is already shown in the lowest resolution layer, no further motion vectors and further segmentation information are signaled in the bitstream. The network generates respective motion vector fields with high resolution (2x2) and highest resolution (4x4) by copying it from one signaled motion vector during upsampling by nearest neighbor. The result is a 4x4 area with all 16 motion vectors identical and equal to the signaled motion vector.

[0271] Figure 32B shows two examples of reconstruction applying a convolutional layer-based approach. Example 1 has the same input as Example 1 in Figure 32A. In particular, the segmentation information of the lowest resolution layer has a flag value set to (1). Correspondingly, the motion information shows a single motion vector. However, instead of simply copying a single motion vector, the motion vectors in the upper and top layers are not exactly identical after applying (possibly trained) convolutional layers.

[0272] Similarly, Example 2 in Figure 32A shows segmentation information 0 in the lowest resolution layer and segmentation information 0101 in the next (higher resolution) layer. Accordingly, two motion vectors for the positions indicated by the segmentation information are signaled in the bitstream as motion information. These are shown in the intermediate layers. As seen in the bottom layer, the signaled motion vectors are each copied four times to cover the highest resolution region. The remaining eight motion vectors in the highest resolution (bottom) layer are signaled in the bitstream. Example 2 in Figure 32B applies convolution instead of nearest neighbor copying. The motion vectors are no longer copied. The transitions between copied motion vectors in Figure 32A are somewhat smoother here, allowing for a reduction in blocking artifacts.

[0273] Similar to the second variant of the encoder described above, in the decoder, the multiple cascaded layers include convolutional layers that do not perform upsampling between layers with different resolutions. Note that encoders and decoders are not necessarily symmetrical in this respect: an encoder may have such additional layers, but a decoder may not, and vice versa. Of course, encoders and decoders may be designed symmetrically, and may have additional layers between the corresponding downsampling and upsampling layers of the encoder and decoder.

[0274] With regard to the combination of segmentation information processing and feature mapping processing, obtaining feature mapping elements from a bitstream is based on processed segmentation information processed by at least one of multiple segmentation information processing layers. The segmentation layer can analyze and interpret the segmentation information, as will be described in more detail below in the section on decoding using segmentation information. Note that the embodiments and examples described herein are applicable in combination with the embodiments in this section. In particular, the segmentation information layer processing described below with reference to Figures 26 to 32B may be performed in combination with the feature mapping processing described herein.

[0275] For example, inputting each of two or more sets of feature map elements into two or more feature map processing layers is based on processed segmentation information processed by at least one of multiple segmentation information processing layers. The acquired segmentation information is represented by a set of syntax elements, where the position of an element in the set of syntax elements indicates which feature map element position the syntax element relates to. The set of syntax elements is a bitstream portion that can be binaryized using an entropy code such as a fixed code, a variable-length code, or an arithmetic code, any of which may be context-adaptive. This disclosure is not limited to any particular coding or form of bitstream, as long as it has a predefined structure known to both the encoder and decoder sides. In this way, the analysis and processing of segmentation information and feature map information can be performed in association. For example, feature map processing includes, for each syntax element, (i) if the syntax element has a first value, parsing the feature map element from the bitstream at the location indicated by the syntax element's position in the bitstream, and (ii) otherwise (or more generally, if the syntax element has a second value), bypassing the parsing of the feature map element from the bitstream at the location indicated by the syntax element's position in the bitstream. The syntax elements may be binary flags that are ordered in the bitstream in the encoder and parsed in the correct order by the decoder through a specific layer structure of the processing network.

[0276] Note that options (i) and (ii) may also be provided for non-binary syntax elements. In this case, the first value means parsing, and the second value means bypassing. Syntax elements may take some additional values ​​other than the first and second values. These may also result in parsing or bypassing, or indicate a particular type of parsing, etc. The number of parsed feature map elements may correspond to a quantity of syntax elements equal to the first value.

[0277] According to an exemplary implementation, for each layer 1 < j < N of a plurality of N feature map processing layers, processing of a feature map by the j-th feature map processing layer includes analyzing a segmentation information element of the j-th feature map processing layer from a bitstream, obtaining a feature map processed by a previous feature map processing layer, analyzing a feature map element from the bitstream, and associating the analyzed feature map element with the obtained feature map. The position of the feature map element in the processed feature map is indicated by the analyzed segmentation information element and the segmentation information processed by a previous segmentation information processing layer. The association can be, for example, replacement of a previously processed feature map element, or combination, such as addition, subtraction, or multiplication. Some exemplary implementations are provided below. The analysis may depend on previously processed segmentation information, which provides the possibility of a very compact and efficient syntax.

[0278] For example, the method may include analyzing a feature map element from a bitstream when a syntax element has a first value, and bypassing the analysis of a feature map element from the bitstream when the syntax element has a second value or when the segmentation information processed by a previous segmentation information processing layer has the first value. This means that the analysis is bypassed when the relevant part has been analyzed in a previous layer. For example, a syntax element analyzed from a bitstream representing segmentation information is a binary flag. As described above, it may be beneficial for the processed segmentation information to be represented by a set of binary flags. The set of binary flags is a sequence of binary flags having a value of either 1 or 0 (corresponding to the first and second values described above).

[0279] In some embodiments, the upsampling of segmentation information in each segmentation information processing layer j further includes determining, for each p-th position in the acquired feature map indicated by the input segmentation information, the upsampled segmentation information indicates the feature map position that falls within the same region in the reconstructed picture as the p-th position. This provides a spatial relationship between the reconstructed image (or reconstructed feature map or generally data), the position in the subsampled feature map, and the corresponding segmentation flag.

[0280] As already stated above, similar to embodiments of encoders, the data for picture or video processing may include picture data (such as picture samples) and / or predicted residual data and / or predicted information data. When "residuals" are referred to in this disclosure, it should be noted that these may be pixel region residuals or transformed (spectral) coefficients (i.e., transformed residuals, residuals represented in a region different from the sample / pixel region).

[0281] Similar to the fourth modification described for the encoder side above, according to the exemplary implementation, a filter is used in the upsampling of the feature map, and the shape of the filter is one of a square, a horizontal rectangle, or a vertical rectangle. Note that the filter shape may be the same as the partition shape shown in Figure 25.

[0282] An exemplary decoder side-layer processing is shown in Figure 20. The motion generation network (unit) 2060 includes signaling logic and one or more (two in this case) upsampling layers that use filters (upsampling filters) which may be selected from a given or predefined set of filters. The selection may be performed on the encoder side, for example by RDO or by other configuration and signaled in the bitstream. On the decoder side, the filter selection instructions are parsed from the bitstream and applied. Alternatively, filters may be selected in the decoder without explicit signaling based on other coding parameters derived from the bitstream. Such parameters may be any parameters correlated with the content, such as prediction type, direction, motion information, residuals, loop filtering characteristics, etc.

[0283] Figure 31 shows a block diagram of an upsampling filter unit 3100 that supports the selection of one of N filters 1 to N. The signaling for filter selection may directly include the index of one of the N filters. It may include the filter orientation, filter order, filter shape, and / or coefficients. On the decoding side, the signaling logic interprets the filter selection flag (e.g., an orientation flag to distinguish between vertical and horizontal filters or further orientations) and supplies the feature map value(s) to the layer having the corresponding set of filter shapes. In Figure 31, a direct connection from the signaling logic to the selective bypass logic allows for no filter selection. The corresponding values ​​of the filter selection indicators may also be signaled or derived in the bitstream.

[0284] Generally, filters are used in feature map upsampling, and inputting information from a bitstream further involves obtaining information from the bitstream that indicates the filter shape and / or filter orientation and / or filter coefficients. There may be implementations where each layer has a set of filters from which to select, or where each layer has one filter, and the signaling logic determines which layer should be selected and which should be bypassed based on filter selection flags (indicators).

[0285] In some embodiments, a flexible filter shape can be provided, in that the information indicating the filter shape represents a mask composed of flags, where a flag having a third value indicates a non-zero filter coefficient, and a flag having a fourth value different from the third value indicates a zero filter coefficient. In other words, as already described for the encoder side, the filter shape can be defined by indicating the position of the non-zero coefficients. The non-zero coefficients can be derived based on predefined rules or signaled.

[0286] The above embodiments of the decoder may be implemented as a computer program product stored on a non-temporary medium, which, when executed on one or more processors, performs any step of the above method. Similarly, the above embodiments of the decoder may be implemented as a device for decoding an image or video, which includes a processing circuit configured to perform any step of the above method. In particular, a device for decoding data for picture or video processing from a bitstream may be provided, which includes an acquisition unit configured to acquire two or more sets of feature map elements from a bitstream, each set of feature map elements relating to a feature map; an input unit configured to input each of the two or more sets of feature map elements to two or more feature map processing layers of a plurality of cascaded layers; and a decoded data acquisition unit configured to acquire decoded data for picture or video processing as a result of processing by the plurality of cascaded layers. These units may be implemented in software or hardware, or a combination of both, as will be described in more detail below.

[0287] Decoding using segmentation information On the receiving end, the decoder of this embodiment performs analysis and interpretation of segmentation information. Thus, a method for decoding data for picture or video processing from a bitstream is provided, as shown in Figure 34. Correspondingly, an apparatus for decoding data for picture or video processing from a bitstream is provided. The apparatus may include processing circuitry configured to perform the steps of the method.

[0288] The method includes obtaining two or more sets of segmentation information elements from a bitstream 3410. This acquisition can be performed by parsing the bitstream. In some exemplary implementations, the bitstream parsing may also include entropy decoding. This disclosure is not limited to any particular method of obtaining data from a bitstream. The method further includes inputting each of the two or more sets of segmentation information elements into two or more segmentation information processing layers of a plurality of cascaded layers 3420. It should be noted that the segmentation information processing layers may be the same layers as the feature map processing layers or different layers. In other words, a single layer may have one or more functions.

[0289] Furthermore, in each of the two or more segmentation information processing layers, the method includes processing each set of segmentation information. Obtaining the decoded data for picture or video processing 3430 is based on the segmentation information processed by the multiple cascaded layers.

[0290] Figure 26 shows exemplary segmentation information for 3-layer decoding. Segmentation information can be thought of as selecting the layer from which feature map elements are analyzed or otherwise acquired (see the encoder-side description). Feature map element 2610 is not selected. Therefore, flag 2611 is set to 0 by the encoder. This means that feature map element 2610, which has the lowest resolution, is not included in the bitstream. However, flag 2611, which indicates that a feature map element is not selected, is included in the bitstream. For example, if the feature map element is a motion vector, this may mean that the motion vector 2610 of the largest block is not selected and is not included in the bitstream.

[0291] In the example shown in Figure 26, in feature map 2620, of the four feature map elements used to determine the feature map elements of feature map 2610, three feature map elements are selected for signaling (indicated by flags 2621, 2622, and 2624), while one feature map element 2623 is not selected. In an example using motion vectors, this could mean that three motion vectors are selected from feature map 2620, with each of their flags set to 1, while one feature map element is not selected, with its flag 2623 set to 0.

[0292] Next, the bitstream may contain all four flags 2621-2624 and three selected motion vectors. Generally, the bitstream may contain four flags 2621-2624 and three selected feature map elements. In feature map 2630, one or more elements that determine the unselected feature map elements of feature map 2620 may be selected.

[0293] In this example, when a feature map element is selected, none of the elements of the high-resolution feature map are selected. In this example, none of the feature map elements of feature map 2630 used to determine the feature map elements signaled by flags 2621, 2622, and 2624 are selected. In one embodiment, none of the flags of these feature map elements are included in the bitstream. Rather, only the flags of the feature map elements of feature map 2630 that determine the feature map element 2623 that has the flag are included in the bitstream.

[0294] In an example where the feature map elements are motion vectors, feature map elements 2621, 2622, and 2624 may each be determined by groups of four motion vectors in feature map 2630. In each group where the motion vectors are determined using flags 2621, 2622, and 2624, the motion vectors may have more similarity to one another than the four motion vectors in feature map 2620 that determine the motion vectors (feature map elements) in the unselected feature map 2630 (signaled by flag 2623).

[0295] Figure 26 illustrates the characteristics of the bitstream described above. Note that the decoder decodes (analyzes) such bitstream accordingly; that is, the decoder determines what information is included (signaled) based on the flag values ​​as explained above, and then analyzes / interprets the analyzed information accordingly.

[0296] In an exemplary implementation, segmentation information is organized as shown in Figure 27. For 2D information such as images or videos that are considered sequences of images, the feature maps of a certain layer may be represented in 2D space. The segmentation information includes indicators (binary flags) about the location in 2D space, indicating whether the feature map value corresponding to this location is presented in the bitstream.

[0297] In Figure 27, there is a starting layer (layer 0) for decoding segmentation information, for example, the lowest resolution layer, i.e., the latent representation layer. In this starting layer, each 2D position contains a binary flag. If such a flag is equal to 1, the selected information contains a feature map value for this position in this particular layer. On the other hand, if such a flag is equal to 0, there is no information for this position in this particular layer. This set of flags (or generally a tensor of flags, here a matrix of flags) is called TakeFromCurrent. The TakeFromCurrent tensor is upsampled to the next layer resolution, for example, using the nearest neighbor method. This tensor is called TakeFromPrev. The flags in this tensor indicate whether the corresponding sample position was filled in the previous layer (here, layer 0).

[0298] As the next step, the signaling logic reads the flag for the current resolution layer (LayerFlag). In this exemplary implementation, only positions that were not filled in the previous layer (not set to 1, not filled with feature map element values) are signaled. This can be expressed using a logical operation as TakeFromPrev==0 or !TakeFromPrev==1, where "!" represents a logical NOT operation (negation).

[0299] The number of flags required for this layer can be calculated as the number of zero (logically false) elements in the TakeFromPrev tensor, or the number of values ​​with 1 (logically true) in the inverse (!TakeFromPrev) tensor. For non-zero elements in the TakeFromPrev tensor, no flags are needed in the bitstream. This is indicated in the diagram by showing "-" at positions that do not need to be read. From an implementation standpoint, it may be easier to calculate the sum of the elements on the inverse tensor as sum(!TakeFromPrev). The signal supply logic can use this arithmetic to determine how many flags need to be parsed from the bitstream. The LayerFlag tensor is obtained by placing the read flags at positions where the value of TakeFromPrev is equal to 1. The TakeFromCurrent tensor for the current resolution layer (Layer 1 in this case) is then obtained as a combination of the TakeFromPrev and LayerFlag tensors by holding the flags at positions read from the bitstream for the current resolution layer and setting the values ​​of positions read in the previous resolution layer (positions marked with "-" in LayerFlag) to zero. This can be expressed and implemented using a logical AND operator, such as TakeFromCurrent = !TakeFromPrev AND LayerFlag. Then, to take into account the position read in the previous resolution layer, the TakeFromCurrent tensor is obtained using a logical OR operation as TakeFromCurrent = TakeFromCurrent OR TakeFromPrev. It should be understood that Boolean operations can be implemented using ordinary mathematical operations, e.g., multiplication for AND and addition for OR. This gives the advantage of preserving and transferring gradients, making it possible to use the above method in end-to-end training.

[0300] Next, the acquired TakeFromCurrent tensor is upsampled to the next resolution layer (layer 2 in this case), and the above process is repeated.

[0301] For generality and to simplify implementation, it is beneficial to unify the processing of all resolution layers without specifically considering the first resolution layer from which all flags are parsed from the bitstream. This can be achieved by initializing TakeFromPrev to zero before processing in the first (low-resolution) layer (Layer0) and repeating the above steps for each resolution layer.

[0302] To further reduce signaling overhead, in some further implementations, LayerFlags for the last resolution layer (here, the third layer, i.e., Layer 2) do not need to be transferred to the bitstream (included in the encoder and parsed in the decoder). This means that for the last resolution layer, the feature map values ​​are sent in the bitstream as selected information (see 1120 in Figure 11) about all positions in the last resolution layer that were not obtained in any of the previous resolution layers. In other words, for the last resolution layer, TakeFromCurrent = !TakeFromPrev, i.e., TakeFromCurrent corresponds to the negated TakeFromPrev. In this case as well, to maintain versatility in processing the last resolution layer, LayerFlag can be initialized to 1, and the same expression TakeFromCurrent = !TakeFromPrev AND LayerFlag can be used.

[0303] In some further possible implementations, the final resolution layer has the same resolution as the original image. If the final resolution layer does not have any additional processing steps, this means that some of the values ​​of the original tensor are sent, bypassing the compression in the autoencoder.

[0304] Below, an example of signal supply logic 2800 is described with reference to Figure 28. In Figure 28, the decoder's signal supply logic 2800 uses segmentation information (LayerFlag) to obtain and utilize selected information (LayerMv) transmitted in the bitstream. In particular, at each layer, the bitstream is parsed by the respective syntax interpretation units 2823, 2822, and 2821 (in this order) to obtain segmentation information (LayerFlag) and, if applicable, selected information (LayerMv). As described above, in order to enable the first layer (syntax interpretation 2823) to operate the same as the other layers, the TakeFromPrev tensor is initialized to all zeros at 2820. The TakeFromPrev tensor is propagated in the order of processing from the syntax interpretation of the previous layer (e.g., from 2823) to the syntax interpretation of the later layer (e.g., 2822). The propagation here involves a 2x upsampling, as described above with reference to Figure 27.

[0305] While interpreting the segmentation information (LayerFlag) for each resolution layer, a tensor called TakeFromCurrent is obtained (generated). This tensor TakeFromCurrent contains flags indicating whether or not feature map information (LayerMv) exists in the bitstream for each specific position in the current resolution layer. The decoder reads the feature map LayerMv values ​​from the bitstream and places them at positions where the flag in the TakeFromCurrent tensor is equal to 1. The total amount of feature map values ​​in the bitstream for the current resolution layer can be calculated based on the number of non-zero elements in TakeFromCurrent, or as sum(TakeFromCurrent) - the sum over all elements of the TakeFromCurrent tensor. As the next step, the tensor join logic of each layer 2813, 2812, and 2811 is joined (e.g., in 2812) by replacing the feature map values ​​at positions where the TakeFromCurrent tensor values ​​are equal to 1 with the feature map values ​​(LayerMv) sent in the bitstream as selected information, thereby combining the output of the previous resolution layer (e.g., generated in 2813 and upsampled to match the resolution of the next layer processing 2812) with the tensor join logic of each layer 2813, 2812, and 2811. As described above, in the first layer (tensor join 2813), the join tensor is initialized to all zeros in 2810 to enable the same operations as the other layers. After processing the LayerFlags from all layers (in 2811) and generating the output tensor of the last layer, in 2801, the join tensor is upsampled by 4 to obtain the original size of the dense optical flow which is W × H.

[0306] The exemplary implementation in Figure 28 provides a fully parallelizable scheme that can run on a GPU / NPU and leverage parallelism. This fully trainable scheme, which transfers gradients, makes it possible to use it in end-to-end trainable video coding solutions.

[0307] Figure 29 shows another possible exemplary implementation of the signal supply logic 2900. This implementation generates a LayerIdx tensor (referred to as LayerIdxUp in Figure 29) containing indexes of layers of different resolutions that indicate which layer should be used to obtain motion information (included in the encoder and parsed in the decoder) to be transferred in the bitstream. In each syntax interpretation block (2923, 2922, 2923), the LayerIdx tensor is updated by adding a TakeFromCurrent tensor multiplied by the upsampling layer index numbered from the highest resolution to the lowest resolution. The LayerIdx tensor is then upsampled and transferred (passed) to the next layer in processing order, for example, from 2923 to 2922, and from 2922 to 2921. To ensure similar processing in all layers, the tensor LayerIdx is zero-initialized in 2920 and passed to the syntax interpretation 2923 of the first layer.

[0308] After the last (in this case, the third) layer, the LayerIdx tensor is upsampled to its original resolution (upsampling 2995 by 4). As a result, each position in LayerIdx contains the index of the layer from which to extract motion information. The positions in LayerIdx correspond to the original resolution of the feature map data (in this case, dense optical flow) at the same resolution, and in this example, they are 2D (matrix). Thus, for each position in the reconstructed optical flow, LayerIdx specifies from where (from which layer's MayerMV) the motion information should be extracted.

[0309] Motion information (LayerMv, also referred to as LayerMvUp in Figure 29) is generated as follows: In each spatial resolution layer, the tensor coupling blocks (2913, 2912, 2911) combine the LayerMv obtained from the bitstream (after passing through their respective syntax interpretation units 2923, 2922, 2921) with an intermediate tensor and an intermediate TakeFromCurrent Boolean tensor based on segmentation information (LayerFlag) obtained from the bitstream, in the manner described above. The intermediate tensor can be initialized with zero (see initialization units 2910, 2919, 2918) or any other value. The initial values ​​are not important because, after all steps are finally completed, these values ​​are not selected for the dense optical flow reconstruction 2990 by this method. The concatenated tensors containing motion information (output from 2913, 2912, and 2911 respectively) are upsampled and concatenated with the concatenated tensors of the previous spatial resolution layer (2902, 2901). The concatenation is performed along an additional dimension corresponding to the motion information obtained from layers of different resolutions (i.e., the 2D tensor 2902 before concatenation becomes a 3D tensor after concatenation, and the 3D tensor 2901 before concatenation remains a 43D tensor after concatenation, but the tensor size increases). Finally, after all upsampling steps for LayerIdxUp and LayerMvUp are complete, the reconstructed dense optical flow is obtained by selecting motion information from LayerMvUp using the values ​​of LayerIdxUp as an axis index in LayerMvUp, where the axis is the dimension increased during the LayerMvUp concatenation step. In other words, the added dimension in LayerMvUp is the dimension across the number of layers, and LayerIdxUp selects the appropriate layer for each position.

[0310] The specific exemplary implementations described above are not limiting to this disclosure. Generally, segmentation can be performed in a variety of possible ways and signaled within a bitstream. Generally, obtaining a set of segmentation information elements is based on segmentation information processed by at least one segmentation information processing layer among a plurality of cascaded layers. Such a layer may include syntax interpretation units (2823, 2822, 2821) that parse / interpret the semantics of the analyzed segmentation information LayerFlag, as shown in Figure 28.

[0311] More specifically, the input set of segmentation information elements is based on processed segmentation information output by at least one of the multiple cascaded layers. This is illustrated, for example, in Figure 28 by the passing of the TakeFromPrev tensor between the syntax interpretation units (2823, 2822, 2821). As already explained in the description of the encoder side, in some exemplary implementations, the segmentation information processed in two or more segmentation information processing layers differs in resolution.

[0312] Furthermore, the processing of segmentation information in two or more segmentation information processing layers includes upsampling, as already illustrated with reference to Figures 9, 13, and others. For example, the upsampling of segmentation information includes nearest neighbor upsampling. Note that in this embodiment and previous embodiments, the disclosure is not limited by the application of nearest neighbor upsampling. Upsampling may include interpolation rather than a simple copy of adjacent sample (element) values. The interpolation may be linear or polynomial, or any known interpolation such as cubic upsampling. With respect to copying, note that the copy performed by the nearest neighbor is a copy of the element value from a given (available) nearest neighbor, e.g., from above or to the left. If there are adjacent elements at the same distance from the position to be filled, it may be necessary to predefine such adjacent elements from which the copy will be made.

[0313] As described above for the first variant, in some exemplary implementations, the upsampling includes transposed convolution. In addition to applying convolutional upsampling to feature map information, or instead, convolutional upsampling may also be applied to segmentation information. Note that the type of upsampling performed on segmentation information is not necessarily the same as the type of upsampling applied to feature map elements.

[0314] Generally, for each segmentation information processing layer j of multiple N segmentation information processing layers from multiple cascade layers, the input is: If -j=1, the initial segmentation information is input from the bitstream (and / or based on initialization, e.g., initialization to 0 in 2820); otherwise, the segmentation information processed by the (j-1)th segmentation information processing layer is input. - This includes outputting the processed segmentation information.

[0315] This is the segmentation information related to the input layer and is not necessarily (still potentially) the entire segmentation information from the bitstream. The upsampled segmentation information in the j-th layer is the segmentation information upsampled in the j-th layer, i.e., the segmentation information output by the j-th layer. Generally, the processing by the segmentation layer includes upsampling (TakeFromPrev) and including new elements from the bitstream (LayerFlag).

[0316] For example, the processing of the input segmentation information by each layer j < N of a plurality of N segmentation information processing layers further includes parsing the segmentation information element (LayerFlag) from the bitstream and associating the parsed segmentation information element with the segmentation information (TakeFromPrev) output by the previous layer (e.g., in the syntax interpretation unit 282x of FIG. 28). The position of the parsed segmentation information element (LayerFlag) within the relevant segmentation information is determined based on the segmentation information output by the previous layer. As can be seen in FIGS. 28 and 29, there can be various different ways of associating and propagating the position information. The present disclosure is not limited to any particular implementation form.

[0317] For example, the amount of segmentation information elements analyzed from the bitstream is determined based on the segmentation information output by the preceding layer. In particular, if a region has already been covered by segmentation information from the previous layer, that region does not need to be covered again on the subsequent layer. It should be noted that this design provides an efficient analysis approach. Each position in the resulting reconstructed feature map data corresponding to a position in the resulting reconstructed segmentation information is associated only with segmentation information belonging to a single layer (out of N processing layers). This means there is no duplication. However, this disclosure is not limited to this approach. Segmentation information may be redundant, although this may lead to maintaining some redundancy.

[0318] As already shown in Figure 27, in some embodiments, the analyzed segmentation information elements are represented by a set of binary flags. The ordering (syntax) of the flags in the bitstream can convey the association between the flags and the layers to which they belong. The ordering (sequence) may be given by a predetermined order of processing in the encoder and corresponding processing in the decoder. These are illustrated, for example, in Figures 16 and 28.

[0319] In some exemplary embodiments, for example, the embodiments described above with reference to the seventh variant, obtaining decoded data for picture or video processing includes determining at least one of the following parameters based on segmentation information. The segmentation information may determine the analysis of additional information such as coding parameters, which may include, as with motion information, intra-picture or inter-picture prediction mode, picture reference index, single-reference or multiple-reference prediction (including dual prediction), presence or absence prediction residual information, quantization step size, motion information prediction type, motion vector length, motion vector resolution, motion vector prediction index, motion vector difference size, motion vector difference resolution, motion interpolation filter, in-loop filter parameters, and / or post-filter parameters. In other words, the segmentation information may specify from which processing layer the coding parameters may be obtained when processed by the segmentation information processing layer. For example, in the encoder approach described above in Figure 22 or Figure 23, the reconstruction (coding) parameters may be received from the bitstream instead of (or in addition to) the motion information (LayerMv). Such reconstruction (coding) parameters blk_rec_params can be analyzed in the decoder in the same manner as illustrated in Figures 28 and 29 for motion information.

[0320] Generally, segmentation information is used for the analysis and input of feature map elements (either motion information or the aforementioned reconstruction parameters or sample-related data). The method may further include obtaining a set of feature map elements from a bitstream and inputting the set of feature map elements into feature map processing layers among several layers, based on the segmentation information processed by the segmentation information processing layer. Furthermore, the method may further include obtaining decoded data for picture or video processing based on the feature maps processed by several cascaded layers. In particular, in some embodiments, at least one of the several cascaded layers is both a segmentation information processing layer and a feature map processing layer. As described above, the network may be designed using separate segmentation information processing layers and feature map processing layers, or using a combined layer having both functions. In some implementations, each of the several layers is either a segmentation information processing layer or a feature map processing layer.

[0321] The methods described above can be embodied as a computer program product stored on a non-temporary medium, which, when executed on one or more processors, causes the processors to perform any of the steps of these methods. Similarly, a device is provided for decoding an image or video, which includes a processing circuit configured to perform any of the method steps of the methods described above. The functional structure of the device also provided by this disclosure may correspond to the functions provided by the embodiments and steps described above. For example, a device is provided for decoding data for picture or video processing from a bitstream, which includes an acquisition unit configured to acquire two or more sets of segmentation information elements from a bitstream; an input unit configured to input each of the two or more sets of segmentation information elements to two or more segmentation information processing layers of a plurality of cascaded layers, respectively; a processing unit configured to process each set of segmentation information in each of the two or more segmentation information processing layers; and a decoded data acquisition unit configured to acquire the decoded data for picture or video processing based on the segmentation information processed in the plurality of cascaded layers. These units and any further units may perform all the functions of the method described above.

[0322] Overview of some embodiments Embodiments relating to encoding using feature information or segmentation information According to one aspect of the present disclosure, a method is provided for encoding data for image or video processing into a bitstream, the method comprising processing the data, the processing comprising generating feature maps in a plurality of cascaded layers, each feature map including its respective resolution, and at least two of the generated feature maps having different resolutions from each other; selecting layers from the plurality of layers that do not have a layer that generates the lowest resolution feature map; and generating a bitstream that includes inserting information relating to the selected layer into the bitstream.

[0323] Such a method can give improved efficiency to such encoding because it allows data from different layers to be encoded, and therefore allows features of different resolutions or other types of layer-related information to be included in the bitstream.

[0324] According to one aspect of the present disclosure, a device for encoding data for image or video processing into a bitstream includes a processing unit configured to process the data, the processing of which includes generating feature maps of different resolutions in a plurality of cascaded layers, each feature map having its own resolution; a selection unit configured to select from the plurality of layers a layer different from the layer that generates the lowest resolution feature map; and a generation unit configured to generate a bitstream, which includes inserting data instructions relating to the selected layer into the bitstream. The processing unit, selection unit, and generation unit may be implemented by processing circuits such as one or more processors, or by any combination of software and hardware.

[0325] Such a device may provide improved efficiency in such decoding because it allows data from different layers to be decoded and used for reconstruction, and thus enables the use of features of different resolutions or other types of layer-related information.

[0326] In exemplary implementations, processing further involves downsampling by one or more of the cascaded layers. The application of downsampling allows for a reduction in processing complexity on the one hand, and also reduces the amount of data provided within the bitstream on the other hand. Furthermore, layers processing different resolutions can thus focus on features at different scales. Therefore, a network processing pictures (still or video) can operate efficiently.

[0327] For example, one or more downsampling layers include mean pooling or max pooling for downsampling. Mean pooling and max pooling operations are part of multiple frameworks, and they provide efficient means for downsampling with low complexity.

[0328] In another example, convolution is used in downsampling. Convolution may provide some more advanced way to downsample using kernels that can be appropriately selected for a particular application, or may even be trainable. This enables a trainable downsampling process that allows finding a better latent representation of motion information, while retaining the advantages of representing and transferring information at different spatial resolutions, increasing adaptability.

[0329] In the example implementation, information related to the selected layer includes elements of the layer's feature map.

[0330] By providing features with different resolutions, the scalability of encoding / decoding is increased, and the bitstream thus generated can offer more flexibility in meeting optimization criteria such as rate, distortion, and complexity, ultimately providing the potential for increased coding efficiency.

[0331] In any of the above examples, for example, the information related to the selected layer includes information indicating from which layer and / or from which part of the feature map of that layer the elements of that layer's feature map were selected.

[0332] Signaling segmentation information can provide efficient coding of feature maps from different layers, such that each area of ​​the original (coded) feature map (data) can be covered by information from only one layer. This is not limiting to the present invention, but in some cases, the present invention can also provide inter-layer overlap for specific regions within the coded feature map (data).

[0333] The method described above, in an exemplary implementation, includes the step of acquiring data to be encoded, the processing of the data to be encoded includes processing by each layer j of a plurality of N cascaded layers, acquiring the data to be encoded as layer input when j=1, and otherwise acquiring the feature map processed by the (j-1)th layer as layer input, and processing the acquired layer input, the processing including downsampling, and outputting the downsampled feature map.

[0334] In response to this, the above-described device, in an exemplary implementation, has a processing unit configured to acquire data to be encoded and to perform processing on the data to be encoded, the processing of the data to be encoded includes processing by each layer j of a plurality of N cascaded layers, acquiring the data to be encoded as a layer input when j=1, and otherwise acquiring a feature map processed by the (j-1)th layer as a layer input, processing the acquired layer input, the processing includes downsampling, and outputting the downsampled feature map.

[0335] The method according to any of the previous examples, in some embodiments, is to select information for insertion into the bitstream, the information being related to a first region in the feature map processed by layer j>1, the first region corresponding to a region in the feature map encoded in a layer smaller than j containing a plurality of elements or in the initial data, and excluding from the selection a region corresponding to the first region from the selection in the feature map processed by layer k (k is an integer greater than or equal to 1 and k < j).

[0336] The apparatus according to any of the foregoing examples, in some embodiments, is to select information for insertion into the bitstream, the information being related to a first region in the feature map processed by layer j>1, the first region corresponding to a region in the feature map encoded in a layer smaller than j containing a plurality of elements or in the initial data, and excluding from the selection a region corresponding to the first region from the selection in the feature map processed by layer k (k is an integer greater than or equal to 1 and k < j), and further includes a processing circuit configured to perform the above.

[0337] Such a selection in a layer does not cover the area of the original feature map covered by other layers and can be particularly efficient with respect to coding overhead.

[0338] In any of the above examples, for example, the data to be encoded includes image information and / or prediction residual information and / or prediction information.

[0339] Alternatively, the information related to the selected layer includes prediction information.

[0340] In any of the above examples, for example, the data related to the selected layer includes an indication of the position of the feature map elements within the feature map of the selected layer.

[0341] Such a display enables appropriately associating feature map elements of different resolutions with the input data region.

[0342] In any of the above examples, for example, the positions of selected and unselected feature map elements are indicated by multiple binary flags based on the positions of the flags in the bitstream.

[0343] Binary flags provide a particularly efficient way to code segmentation information.

[0344] According to one embodiment, in the method or apparatus described above, the processing by layer j of a plurality of N cascade layers includes: determining a first cost arising from reconstructing a portion of the reconstructed picture using feature map elements output by the j-th layer; determining a second cost arising from reconstructing a portion of the picture using feature map elements output by the (j-1)-th layer; and, if the first cost is higher than the second cost, selecting the (j-1)-th layer and selecting information about the portion in the (j-1)-th layer.

[0345] Providing optimization that includes distortion offers an efficient means of achieving the desired quality.

[0346] For example, the first and second costs include data volume and / or distortion. Optimization by considering the rate (the amount of data generated by the encoder) and the distortion of the reconstructed picture allows for flexible meeting of various application or user requirements.

[0347] Alternatively or additionally, the data to be encoded is a motion vector field. The methods described above are readily applicable to compress motion vector fields such as dense optical flow or subsampled optical flow. The application of these methods may provide efficient coding of motion vectors (with respect to rate and distortion or other criteria) and allow for further reduction of the bitstream size of the encoded picture or video data.

[0348] In some embodiments, the prediction information includes a reference index and / or a prediction mode. Further information regarding the prediction may be processed in addition to, or instead of, the motion vector field. The reference index and prediction mode, like the motion vector field, may be correlated with the picture content; therefore, encoding feature map elements with different resolutions can improve efficiency.

[0349] For example, the amount of data includes the amount of data required to transmit data related to the selected layer. In this way, the overhead generated by providing information related to layers other than the output layer may be taken into consideration during optimization.

[0350] Alternatively, distortion is calculated by comparing the reconstructed picture with the target picture. Such an end-to-end quality comparison ensures that distortion in the reconstructed image is properly accounted for. Therefore, optimization allows for the selection of an efficient coding approach, which can more accurately meet the quality requirements imposed by the application or user.

[0351] In any of the above examples, for instance, the process includes additional convolutional layers between cascaded layers having different resolutions.

[0352] Providing such additional layers in a cascaded layer network allows for the introduction of additional processing, such as various types of filtering, to improve the quality or efficiency of coding.

[0353] According to an exemplary implementation, the processing circuit of a method or apparatus based on the above-described embodiment includes, in layered downsampling, obtaining a first feature map by downsampling the input feature map using a first filter and obtaining a second feature map by downsampling the input feature map using a second filter, determining a third cost resulting from reconstructing the portion of the picture reconstructed using the first feature map, determining a fourth cost resulting from reconstructing the portion of the picture reconstructed using the second feature map, and, in selecting, selecting the first feature map if the third cost is less than the fourth cost.

[0354] Applying different downsampling filters can be helpful in adapting to different characteristics of the content.

[0355] For example, the shapes of the first and second filters may be a square, a horizontally oriented rectangle, or a vertically oriented rectangle.

[0356] While these filters are still of a simple shape, they can offer additional improvements in terms of adapting to object boundaries.

[0357] A step performed by a method step or a processing circuit of an apparatus may further include obtaining a mask, the mask consisting of flags, the mask representing an arbitrary filter shape, and one of the first and second filters having the arbitrary filter shape.

[0358] This provides the flexibility to design filters of any shape.

[0359] A step performed by a method step or a processing circuit of an apparatus may further include processing data relating to the same picture segmented into blocks having different block sizes and shapes in different layers, wherein the selection includes selecting layers based on a cost calculated for a given set of coding modes.

[0360] In some exemplary implementations, the process involves determining the cost for different sets of coding modes for at least one layer, and then selecting one of the sets of coding modes based on the determined cost.

[0361] Applying optimization to coding modes can enable efficient rate distortion optimization, and therefore, improved coding efficiency.

[0362] For example, the instructions for the data related to the selected layer include a selected set of coding modes.

[0363] According to one aspect of the present disclosure, a computer program stored on a non-temporary medium includes code that, when executed on one or more processors, performs any of the steps of the methods presented above.

[0364] According to one aspect of the present disclosure, a device for encoding an image or video is provided, which includes a processing circuit configured to perform a method according to any of the examples presented above.

[0365] Any of the above-described devices can be realized on an integrated chip. The present invention can be implemented in hardware (HW) and / or software (SW). Furthermore, the HW-based implementation can be combined with the SW-based implementation.

[0366] Please note that this disclosure is not limited to any particular framework. Furthermore, this disclosure is not limited to image or video compression, but may also apply to object detection, image generation, and recognition systems.

[0367] For clarity, any one of the embodiments described above may be combined with one or more of the other embodiments described above to create a new embodiment within the scope of this disclosure.

[0368] Embodiments relating to decoding using feature map elements According to one embodiment, a method is provided for decoding data from a bitstream for picture or video processing, the method comprising: obtaining two or more sets of feature map elements from a bitstream, each set of feature map elements relating to a feature map; inputting each of the two or more sets of feature map elements into two or more feature map processing layers of a plurality of cascaded layers; and obtaining the decoded data for picture or video processing as a result of processing by the plurality of cascaded layers.

[0369] Such a method may provide improved efficiency because it allows data from different layers to be used in decoding, and therefore allows features or other types of layer-related information to be analyzed from the bitstream.

[0370] For example, feature maps are processed in each of two or more feature map processing layers, and the feature maps processed in each of the two or more feature map processing layers differ in resolution.

[0371] In some embodiments, the processing of feature maps in two or more feature map processing layers includes upsampling.

[0372] The application of upsampling allows for a reduction in processing complexity (since the first layer has a lower resolution), and also reduces the amount of data given in the bitstream and analyzed in the decoder. Furthermore, layers processing different resolutions can thus focus on features at different scales. Therefore, a network processing pictures (still or video) can operate efficiently.

[0373] In an exemplary implementation, the method further includes the step of obtaining segmentation information relating to two or more layers from a bitstream, where obtaining feature map elements from the bitstream is based on the segmentation information, and inputting sets of feature map elements into two or more feature map processing layers is also based on the segmentation information.

[0374] Using segmentation information can provide efficient decoding of feature maps from different layers such that each area of ​​the original resolution (to be reconstructed) can be covered by information from only one layer. While not limiting the invention, the invention may, in some cases, also provide inter-layer overlap of specific areas within the feature map (data). For example, multiple cascaded layers further comprise multiple segmentation information processing layers, and the method further comprises processing segmentation information in multiple segmentation information processing layers. Such an approach provides the possibility of controlling the analysis of feature elements from different layers.

[0375] In some embodiments, the processing of segmentation information in at least one of a plurality of segmentation information processing layers includes upsampling. The hierarchical structure of the segmentation information provides a small amount of side information to be inserted into the bitstream, and thus can improve efficiency and / or processing time.

[0376] For example, the upsampling of segmentation information and / or feature maps includes nearest neighbor upsampling. Nearest neighbor upsampling has low computational complexity and can be easily implemented. Moreover, it is particularly efficient for logical representations such as flags.

[0377] In some embodiments and examples, the upsampling of segmentation information and / or feature maps includes transposed convolution. The use of convolution may help reduce blocking artifacts and may enable a trainable solution in which the upsampling filter is selectable.

[0378] In an exemplary implementation, obtaining feature map elements from a bitstream is based on processed segmentation information processed by at least one of multiple segmentation information processing layers.

[0379] In an exemplary implementation, inputting each of two or more sets of feature map elements into two or more feature map processing layers is based on processed segmentation information processed by at least one of multiple segmentation information processing layers.

[0380] According to one embodiment, the acquired segmentation information is represented by a set of syntax elements, the position of an element within the set of syntax elements indicates which feature map element position the syntax element is associated with, and the feature map processing includes, for each syntax element, parsing from the bitstream the element of the feature map at the position indicated by the syntax element's position in the bitstream if the syntax element has a first value, and otherwise bypassing the parsing from the bitstream the element of the feature map at the position indicated by the syntax element's position in the bitstream.

[0381] Such a relationship between the segmentation information and the feature map information enables both efficient coding of frequency information and analysis in a hierarchical structure considering different resolutions.

[0382] For example, the processing of the feature map by each layer 1 < j < N of a plurality of N feature map processing layers includes analyzing the segmentation information element of the j-th feature map processing layer from the bitstream, obtaining the feature map processed by the preceding feature map processing layer, analyzing the feature map element from the bitstream, and associating the analyzed feature map element with the obtained feature map. The position of the feature map element in the processed feature map is indicated by the analyzed segmentation information element and the segmentation information processed by the preceding segmentation information processing layer.

[0383] In particular, when the syntax element has a first value, analyzing the element of the feature map from the bitstream, and bypassing the analysis of the element of the feature map from the bitstream when the syntax element has a second value or when the segmentation information processed by the preceding segmentation information processing layer has the first value.

[0384] For example, the syntax element analyzed from the bitstream representing the segmentation information is a binary flag. In particular, the processed segmentation information is represented by a set of binary flags.

[0385] Providing binary flags enables efficient coding. On the decoder side, the processing of logical flags can be performed with low complexity.

[0386] In an exemplary implementation, the upsampling of segmentation information in each segmentation information processing layer j further includes determining, for each p-th position in the acquired feature map indicated by the input segmentation information, the indication of the feature map position contained in the same region within the reconstructed picture as the p-th position as the upsampled segmentation information.

[0387] For example, data for picture or video processing includes a motion vector field. Since a dense optical flow or motion vector field with a resolution similar to the resolution of the picture is desirable for modeling motion, the hierarchical structure of the present invention is readily applicable and efficient for reconstructing such motion information. Using layering and signaling, a good trade-off between rate and distortion can be achieved.

[0388] For example, data for picture or video processing includes picture data and / or predictive residual data and / or predictive information data. This disclosure can be used for various different parameters. However, picture data and / or predictive residual data and / or predictive information data may still have some redundancy in the spatial domain, and the layered approach described herein can provide efficient decoding from the bitstream using different resolutions.

[0389] In some embodiments and examples, a filter is used in the upsampling of the feature map, and the shape of the filter is one of a square, a horizontal rectangle, or a vertical rectangle.

[0390] Applying different upsampling filters can be helpful in adapting to different characteristics of the content. For example, when filters are used in feature map upsampling, inputting information from a bitstream further involves obtaining information from the bitstream that indicates the filter shape and / or filter coefficients.

[0391] In response, the decoder can provide better reconstruction quality based on the information from the encoder transmitted in the bitstream.

[0392] For example, the information indicating the filter shape represents a mask composed of flags, where a flag with a third value indicates a non-zero filter coefficient, and a flag with a fourth value different from the third value indicates a zero filter coefficient, thus the mask represents the filter shape. This provides flexibility for designing filters of any shape.

[0393] For example, multiple cascaded layers include convolutional layers that do not involve upsampling between layers with different resolutions.

[0394] Providing such additional layers in a cascaded layer network allows for the introduction of additional processing, such as various types of filtering, to improve the quality or efficiency of coding.

[0395] According to one embodiment, a computer program product stored in a non-temporary medium is provided, and when this computer program product is executed on one or more processors, it performs one of the methods described above.

[0396] According to one embodiment, a device for decoding an image or video is provided, which includes a processing circuit configured to perform a method according to any of the embodiments and examples described above.

[0397] According to one embodiment, a device is provided for decoding data for picture or video processing from a bitstream, the device comprising: an acquisition unit configured to acquire two or more sets of feature map elements from a bitstream, wherein each set of feature map elements relates to a feature map; an input unit configured to input each of the two or more sets of feature map elements to two or more feature map processing layers of a plurality of cascaded layers; and a decoded data acquisition unit configured to acquire the decoded data for picture or video processing as a result of processing by the plurality of cascaded layers.

[0398] Any of the above-described devices can be realized on an integrated chip. The present invention can be implemented in hardware (HW) and / or software (SW). Furthermore, the HW-based implementation can be combined with the SW-based implementation.

[0399] Please note that this disclosure is not limited to any particular framework. Furthermore, this disclosure is not limited to image or video compression, but may also apply to object detection, image generation, and recognition systems.

[0400] Some exemplary implementations in hardware and software A corresponding system capable of deploying the encoder-decoder processing chain described above is shown in Figure 35. Figure 35 is a schematic block diagram illustrating exemplary coding systems, e.g., video, image, audio, and / or other coding systems (or short coding systems) that may utilize the technology of the present application. The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video coding system 10 represent examples of devices that may be configured to perform the technology described in the various examples of the present application. For example, video coding and decoding may employ a processing network such as a neural network, or generally, one of those described in the embodiments and examples above.

[0401] As shown in Figure 35, the coding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 in order to decode the encoded picture data 13.

[0402] The source device 12 includes an encoder 20 and, optionally, a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.

[0403] The picture source 16 may include, or could include, any kind of picture capturing device, e.g., a camera for capturing real-world pictures, and / or any kind of picture generating device, e.g., a computer graphics processor for generating computer-animated pictures, or any other kind of device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may be any kind of memory or storage for storing any of the aforementioned pictures.

[0404] Unlike the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be referred to as the original picture or original picture data 17.

[0405] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain a preprocessed picture 19 or preprocessed picture data 19. The preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 18 may be an optional component. A neural network may be used for preprocessing.

[0406] The video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21.

[0407] The communication interface 22 of the source device 12 may be configured to receive encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0408] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and may further include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34, i.e., optionally.

[0409] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or any further processed version thereof) for example directly from the source device 12, or from any other source, such as a storage device, such as an encoded picture data storage device, and to supply the encoded picture data 21 to the decoder 30.

[0410] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or any type of private and public network or any combination thereof.

[0411] The communication interface 22 may be configured, for example, to package the encoded picture data 21 into an appropriate format, such as a packet, and / or to process the encoded picture data using any kind of transmit encoding or processing for transmission over a communication link or communication network.

[0412] A communication interface 28 forming a counterpart to communication interface 22 may be configured, for example, to receive transmitted data and process the transmitted data using any kind of corresponding transmit decoding or processing and / or depackaging to obtain encoded picture data 21.

[0413] Both communication interfaces 22 and 28 may be configured as unidirectional or bidirectional communication interfaces, as indicated in Figure 35 by arrows pointing from source device 12 to destination device 14 for communication channel 13, for example, and may be configured to send and receive messages, set up connections, and exchange any other information relating to communication links and / or data transmission, such as encoded picture data transmission. Decoder 30 is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded picture 31 (for example, by employing the neural network described in the embodiments and examples above).

[0414] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, to obtain post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare the decoded picture data 31 for display by, for example, the display device 34.

[0415] The display device 34 of the destination device 14 is configured to receive post-processed picture data 33 for displaying the picture to, for example, a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, for example, an integrated or external display or monitor, or may include one. The display may be, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a microLED display, an LCoS (liquid crystal on silicon), a DLP (digital light processor), or any other type of display.

[0416] Although Figure 35 shows the source device 12 and the destination device 14 as separate devices, the device embodiment may also include the functionality of both the source device 12 or its corresponding functionality and the destination device 14 or its corresponding functionality. In such embodiments, the source device 12 or its corresponding functionality and the destination device 14 or its corresponding functionality may be implemented using the same hardware and / or software, or by separate hardware and / or software, or by any combination thereof.

[0417] As will become apparent to those skilled in the art based on the description, the presence and (exact) division of functions or functions within different units in the source device 12 and / or destination device 14 as shown in Figure 35 may vary depending on the actual device and application.

[0418] An encoder 20 (e.g., a video encoder 20) or a decoder 30 (e.g., a video decoder 30), or both an encoder 20 and a decoder 30, may be implemented via processing circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated to video coding, or any combination thereof. The encoder 20 may be implemented via processing circuit 46 to embody various modules, including neural networks. The decoder 30 may be implemented via processing circuit 46 and can embody various modules as described in the embodiments and examples above. The processing circuits may be configured to perform various operations, as described later. If the technology is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the disclosure. Either the video encoder 20 or the video decoder 30 may be integrated as part of a combined encoder / decoder (codec) in a single device, for example, as shown in Figure 36.

[0419] The source device 12 and destination device 14 may include any type of handheld or stationary device, including a wide range of devices such as notebook or laptop computers, mobile phones, smartphones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video game consoles, video streaming devices (such as content service servers or content distribution servers), broadcast receiver devices, broadcast transmitter devices, etc., and may or may not use an operating system. In some cases, the source device 12 and destination device 14 may be equipped for wireless communication. Thus, the source device 12 and destination device 14 may be wireless communication devices.

[0420] In some cases, the video coding system 10 shown in Figure 35 is merely an example, and the technology of this application may be applicable to video coding configurations (e.g., video coding or video decoding) that do not necessarily involve data communication between an coding device and a decoding device. In other examples, the data may be retrieved from local memory, streamed over a network, etc. A video coding device may code the data and store it in memory, and / or a video decoding device may retrieve the data from memory and decode it. In some examples, coding and decoding are performed by devices that do not communicate with each other but simply code the data into memory and / or retrieve the data from memory and decode it.

[0421] Figure 37 is a schematic diagram of a video coding device 3700 according to one embodiment of the present disclosure. The video coding device 3700 is suitable for implementing the disclosed embodiments as described herein. In one embodiment, the video coding device 3700 may be a decoder, such as the video decoder 30 in Figure 35, or an encoder, such as the video encoder 20 in Figure 35.

[0422] The video coding device 3700 includes an inlet port 3710 (or input port 3710) and a receiver unit (Rx) 3720 for receiving data, a processor, logic unit, or central processing unit (CPU) 3730 for processing data, a transmitter unit (Tx) 3740 and an exit port 3750 (or output port 3750) for transmitting data, and memory 3760 for storing data. The video coding device 3700 may also include optical-electrical (OE) components and electrical-optical (EO) components coupled to the inlet port 3710, receiver unit 3720, transmitter unit 3740, and exit port 3750 for optical or electrical signal input or output.

[0423] The processor 3730 is implemented by hardware and software. The processor 3730 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 3730 communicates with the input port 3710, the receiver unit 3720, the transmitter unit 3740, the output port 3750, and the memory 3760. The processor 3730 includes a coding module 3770. The coding module 3770 implements the embodiments disclosed above. For example, the coding module 3770 implements, processes, prepares, or provides various coding operations. Thus, including the coding module 3770 provides a substantial improvement to the functionality of the video coding device 3700, resulting in the conversion of the video coding device 3700 to different states. Alternatively, the coding module 3770 is implemented as instructions stored in the memory 3760 and executed by the processor 3730.

[0424] Memory 3760 may include one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device to store such programs when they are selected for execution, and to store instructions and data read during program execution. Memory 3760 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random-access memory (RAM), ternarily content addressable memory (TCAM), and / or static random-access memory (SRAM).

[0425] Figure 38 is a simplified block diagram of a device 3800 that can be used as either or both of the source device 12 and destination device 14 from Figure 35, according to an exemplary embodiment.

[0426] The processor 3802 within the device 3800 may be a central processing unit. Alternatively, the processor 3802 may be any other type of device, or multiple devices, capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed implementation can be practiced using a single processor, e.g., processor 3802, as illustrated, but advantages in speed and efficiency can be achieved using two or more processors.

[0427] The memory 3804 within the device 1100 may, in its implementation, be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as memory 3804. Memory 3804 may include code and data 3806 accessed by the processor 3802 using the bus 3812. Memory 3804 may further include an operating system 3808 and an application program 3810, the application program 3810 including at least one program that enables the processor 3802 to perform the method described herein. For example, the application program 3810 may include applications 1 to N, the applications 1 to N further including picture coding (encoding or decoding) applications that perform the method described herein.

[0428] The device 3800 may also include one or more output devices, such as a display 3818. The display 3818 may, in one example, be a touch-sensitive display that combines a display with a touch-sensitive element capable of operating to sense touch input. The display 3818 may be coupled to the processor 3802 via a bus 3812.

[0429] Although shown here as a single bus, the bus 3812 of device 3800 may consist of multiple buses. Furthermore, secondary storage can be directly coupled to other components of device 3800 or accessed via a network, and may include a single integrated unit such as a memory card, or multiple units such as multiple memory cards. Thus, device 3800 can be implemented in a wide variety of configurations.

[0430] In summary, this disclosure relates to a method and apparatus for encoding data (for static or video processing into a bitstream). In particular, the data is processed by a network comprising multiple cascaded layers. This processing generates a feature map for each layer. Feature maps processed (output) by at least two different layers have different resolutions. In this processing, layers different from the layer that generates the lowest resolution feature map (e.g., latent space) are selected from the cascaded layers. The bitstream contains information related to the selected layer. This approach provides scalable processing that can operate at different resolutions, and as a result, the bitstream can transmit information about such different resolutions. Thus, data can be efficiently coded in a bitstream according to a resolution that may vary depending on the content of the coded picture data.

[0431] This disclosure further relates to a method and apparatus for decoding data (for still or video processing into a bitstream). In particular, two or more sets of feature map elements are obtained from a bitstream. Each set of feature map elements is associated with a feature map. Each of the two or more sets of feature map elements is then input to two or more feature map processing layers of a plurality of cascaded layers, respectively. Decoded data for picture or video processing is then obtained as a result of processing by the plurality of cascaded layers. Thus, data can be decoded from a bitstream in an efficient manner in a hierarchical structure.

[0432] This disclosure further relates to a method and apparatus for decoding data (for still or video processing of a bitstream). Two or more sets of segmentation information elements are obtained from the bitstream. Each of the two or more sets of segmentation information elements is then input to two or more segmentation information processing layers of a plurality of cascaded layers. Each of the two or more segmentation information processing layers processes each set of segmentation information. Decoded data for picture or video processing is obtained based on the segmentation information processed by the plurality of cascaded layers. Thus, data can be decoded from a bitstream in an efficient manner in a hierarchical structure. [Explanation of symbols]

[0433] 10 Video Coding Systems 12 Source Devices 13 Encoded Picture Data 13 Communication Channels 14 Destination device 16 Picture Sources 17 Pictures 18 Preprocessor / Preprocessing Unit 19 Pictures 20 Video Encoders 21 Encoded Picture Data 21-bitstream 22 Communication interface or communication unit 28 Communication interface or communication unit 30 video decoders 31 Decrypted Picture 32 Post-processor / Post-processing Units 33 Post-processing picture data 34 Display Devices 46 Processing Circuit 101 Encoder 102 Quantizer 103 Hyper Encoder / Decoder 104 Decoder 105 Arithmetic Encoder 106 Arithmetic Decoder 107 modules 108 modules 109 modules 110 modules 121 encoders 122 Quantizer 125 Arithmetic Coding Module 123 Hyper Encoder 147 Hyper Decoder 201 Input Interface 203 Blocks 204 Residual Calculation Unit 206 Conversion Processing Unit 208 Quantization Units 210 Inverse Quantization Unit 212 Inverse Transform Processing Unit 214 Reconfiguration Unit 220 Loop Filter Unit 230 Decoded Picture Buffer (DPB) 244 Interpretation Units 254 prediction units 260 Mode Selection Unit 262 divided units 270 Entropy Coding Units 272 Output Interfaces 304 Entropy Decoding Unit 309 Quantization coefficient 310 Inverse Quantization Unit 312 Inverse Transform Processing Unit 313 Residual Block 314 Reconfiguration Unit 315 Reconstructed Blocks 320 Loop Filtering Unit 322 Quantizer 331 Decoded Picture 360 Mode Applicable Unit 365 Prediction Block 401 Downsampling Layer 402 Downsampling Layer 403 Downsampling Layer 404 Downsampling Layer 405 Downsampling Layer 406 Downsampling Layers 413 Output signal 414 Input Images 415 components 407 Upsampling Layer 408 upsampling layers 409 upsampling layers 410 upsampling layers 411 upsampling layers 412 upsampling layers 420 Further layers 430 Convolutional Layers 610 Dense Optical Flow 611 LayerMv Tensor 613 Cost Calculation Unit 614-layer information selection unit 621 LayerMV tensor 622 results 623 Cost Calculation Unit 624-layer information selection unit 625 Pooling operations 631 LayerMv 633 Cost Calculation Unit 634-layer information selection unit 635 MinCost Pooling 705 Indicator 710 Layer 1 Cost Calculation Unit 720 Minimum Cost Pooling 730-layer information selection module 810 Processing 820 Selection 830 Generation Step 900 Network 901 Input 901 Video Processing 911 Cascade layer 912 Cascade Layer 913 Cascade Layer 920 Signal Selection Logic 930 bitstream 940 Signal Supply Logic 951 upsampling layers 952 upsampling layers 953 Upsampling Layers 1010 Array 1020 array 1050 Feature Map 1060 Feature Map 1100 Signal Selection Circuit (Logic) 1120 elements 1110 Dense optical flow 1110 Feature Map 1120 Information 1130 Segmentation Information 1140 Motion Segmentation Network 1150 bitstream 1210 Optical Flow Estimation Module (Unit) 1215 Optical Flow 1220 Motion Specification (or Segmentation) Module 1250 bitstream 1260 Movement Information 1270 Motion generation unit 1275 Motion Vector Field 1280 Motion Compensation Unit 1310 Network 1320 Signal Selection Logic 1350 bitstream 1360 Motion Generation Unit 1400 cost calculation (or estimation) units 1405 Reference Picture 1408 Target Picture 1415 Motion Field Upsampling 1420 Motion compensation 1430 Distortion 1440 Rate Estimation Module 1450-bit normalization 1460 Cost Calculation Module 1470 downsampling 1480 cost tensor 1500 Cost Estimated Units 1501 Downsampling 1505 Reference Picture 1510 Motion compensation 1515 Downsampling 1518 Downsampling 1530 Distortion Evaluation 1540 Estimated Rate 1600 Signal Selection Unit 1700 Signal Selection Logic 1810 Motion Segmentation Unit 1860 Motion Generation Network (Module) 1869 Motion generation unit 1910 Motion Segmentation Network (Module) 1960 Motion Generation Unit 2010 Motion Segmentation Unit 2060 Motion Generation Network (Unit) 2201 block size 2202 block size 2203 block size 2211 Cost Calculation Unit 2212 Cost Calculation Unit 2213 Cost Calculation Unit 2222 Minimum Cost Pooling Operation 2223 MinCost Pooling 2231 Layer Information Selection Unit 2232 Layer Information Selection Unit 2233 Layer Information Selection Unit 230x block size 2300 Cost Calculation Units 2310 units 2320 Reconfiguration Unit 2330 Distortion Calculation Unit 2340 downsampling 2360-bit estimation unit 2410 Cost Calculation Unit 2420 Minimum Cost Pooling 2430 Layer Information Selection Unit 2510 Quadtree splitting 2520 Binary tree splitting 2530 Asymmetric Binary Tree Partitioning 2540 ternary tree splitting 2610 Feature Map Elements 2611 Flag 2620 Feature Map 2621 Flag 2622 Flag 2623 Feature Map Elements 2624 Flag 2630 Feature Map 2800 Signal Supply Logic 2801 Upsampling 2811 Tensor Joining Logic 2812 Tensor Joining Logic 2813 Tensor Joining Logic 2821 Syntax Interpretation Unit 2822 Syntax Interpretation Unit 2823 Syntax Interpretation Unit 2900 Signal Supply Logic 2901 Link 2902 Link 2910 Initialization Unit 2911 Tensor coupling block 2912 Tensor coupling block 2913 Tensor coupling block 2918 Initialization Unit 2919 Initialization Unit 2921 Syntax Interpretation Block 2922 Syntax Interpretation Block 2923 Syntax Interpretation Block 2990 Optical Flow Reconfiguration 3000 Convolutional Filter Units 3030 bitstream 3040 Signal Supply Logic 3100 Upsampling Filter Unit 3700 video coding devices 3710 Entrance Port 3720 Receiver Unit (Rx) 3730 Central Processing Unit (CPU) 3740 Transmitter Unit (Tx) 3750 Exit Port 3760 memory 3770 Coding Modules 3800 equipment 3802 Processor 3804 memory 3806 data 3808 Operating System 3810 Application Program 3812 Bus 3818 Display

Claims

1. A method for decoding data from a bitstream for picture or video processing, wherein the method is The steps include obtaining two or more sets of segmentation information elements from the bitstream, The steps include inputting each of two or more sets of the aforementioned segmentation information elements to two or more segmentation information processing layers among a plurality of cascade layers, Each of the two or more segmentation information processing layers includes the step of processing each set of segmentation information, The step of obtaining decoded data for picture or video processing is based on the segmentation information processed by the plurality of cascade layers, A method comprising: the bitstream comprising one or more feature map elements of a feature map, and the segmentation information contained in the bitstream sets the one or more feature map elements to be transmitted.

2. The method according to claim 1, wherein the step of obtaining the set of segmentation information elements is based on segmentation information processed by at least one segmentation information processing layer among the plurality of cascade layers.

3. The method according to claim 1 or 2, wherein the step of inputting the set of segmentation information elements is based on the processed segmentation information output by at least one of the plurality of cascade layers.

4. The method according to any one of claims 1 to 3, wherein the segmentation information processed in each of the two or more segmentation information processing layers differs in resolution.

5. The method according to claim 4, wherein the step of processing the segmentation information in the two or more segmentation information processing layers includes the step of upsampling.

6. The method according to claim 5, wherein the step of upsampling the segmentation information includes nearest neighbor upsampling.

7. The method according to claim 5 or 6, wherein the step of upsampling the segmentation information includes transposed convolution.

8. For each of the N segmentation information processing layers j among the plurality of cascade layers, - The input step is to input initial segmentation information from the bitstream if j=1, otherwise input segmentation information processed by the (j-1)th segmentation information processing layer, The method according to any one of claims 1 to 7, comprising the step of outputting the processed segmentation information.

9. The step of processing the input segmentation information by each layer j < N of the plurality of segmentation information processing layers is as follows: The method according to claim 8, further comprising the steps of analyzing segmentation information elements from the bitstream and associating the analyzed segmentation information elements with segmentation information output by a preceding layer, wherein the position of the analyzed segmentation information elements within the associated segmentation information is determined based on the segmentation information output by the preceding layer.

10. The method according to claim 9, wherein the amount of segmentation information elements analyzed from the bitstream is determined based on the segmentation information output by the preceding layer.

11. The method according to claim 9 or 10, wherein the analyzed segmentation information elements are represented by a set of binary flags.

12. The step of obtaining decoded data for picture or video processing is based on segmentation information, - Predictive mode within or between pictures, - Picture reference index, - Single-reference or multiple-reference prediction (including dual prediction), - Predicted residual information for existence or absence, - Quantization step size, - Motion information prediction type, - Length of the motion vector, - Motion vector resolution, - Motion vector prediction index, - Motion vector difference size, - Motion vector difference resolution, - Motion interpolation filter, - Loop filter parameters, - Post-filter parameters, The method according to any one of claims 1 to 11, comprising the step of determining at least one of the following.

13. The steps include obtaining a set of feature map elements from the bitstream, and inputting the set of feature map elements into the feature map processing layer among the multiple layers based on the segmentation information processed by the segmentation information processing layer, The method according to any one of claims 1 to 12, further comprising the step of obtaining the decoded data for picture or video processing based on the feature maps processed by the plurality of cascade layers.

14. The method according to claim 13, wherein at least one of the plurality of cascade layers is a segmentation information processing layer and a feature map processing layer.

15. The method according to claim 13, wherein each of the plurality of layers is either a segmentation information processing layer or a feature map processing layer.

16. A computer program stored in a non-temporary medium, which, when executed on one or more processors, performs the method described in any one of claims 1 to 15.

17. A device for decoding an image or video, comprising a processing circuit configured to perform the method described in any one of claims 1 to 15.

18. A device for decoding data from a bitstream for picture or video processing, wherein the device is An acquisition unit configured to acquire two or more sets of segmentation information elements from the bitstream, An input unit configured to input each of the two or more sets of segmentation information elements to two or more segmentation information processing layers among a plurality of cascade layers, Each of the two or more segmentation information processing layers includes a processing unit configured to process each set of segmentation information, A decoding data acquisition unit configured to acquire decoded data for picture or video processing based on the segmentation information processed in the plurality of cascade layers, A device wherein the bitstream includes one or more feature map elements of a feature map, and the segmentation information included in the bitstream sets the one or more feature map elements to be transmitted.

Citation Information

Patent Citations

  • Method for analysing media content

    US20190251360A1