Method and apparatus for encoding and decoding image data

By processing images into overlapping tiles with consistent extent and padding, neural network-based encoding and decoding methods improve the efficiency and accuracy of image and video compression, addressing the limitations of existing codecs.

WO2025209720A1PCT designated stage Publication Date: 2025-10-09HUAWEI TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/054401
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-03
Filing Date
2025-02-19
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing hybrid image and video codecs, such as HEVC, VVC, and EVC, lack efficient integration of neural network architectures for improved encoding and decoding, particularly in handling feature maps across devices, leading to suboptimal performance in compression and reconstruction.

Method used

The use of neural network-based encoding and decoding methods that process images into tiles in a latent space, ensuring each tile overlaps with neighbors and maintains consistent extent, with edge tiles padded if necessary, to facilitate efficient encoding and decoding across devices.

Benefits of technology

This approach enhances the efficiency and accuracy of image and video compression by reducing artifacts and improving the alignment of feature maps, allowing for better scalability and compatibility across different devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025054401_09102025_PF_FP_ABST
    Figure EP2025054401_09102025_PF_FP_ABST
Patent Text Reader

Abstract

For encoding or decoding an image, tiles may be selected so that they overlap each other in a dimension of the image. An image processing device for encoding or decoding the image may select or use tiles such that there is a set of tiles extending fully across the image in the said dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set. The extent of the image in the said dimension uniquely represented by one of the tiles of the set may differ from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension. By using tiles of the same size, the process of encoding or decoding the image can be made simpler or faster.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD AND APPARATUS FOR ENCODING AND DECODING IMAGE DATA

[0002] TECHNICAL FIELD

[0003] Embodiments of the present disclosure generally relate to the field of encoding and decoding data based on a neural network architecture. In particular, some embodiments relate to methods and apparatuses for such encoding and decoding images and / or videos from a bitstream using a plurality of processing layers.

[0004] BACKGROUND

[0005] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, signal is typically encoded block-wisely by predicting a block and by further coding only the difference between the original bock and its prediction. In particular, such coding may include transformation, quantization and generating the bitstream, usually including some entropy coding. Typically, the three components of hybrid coding methods - transformation, quantization, and entropy coding - are separately optimized. Modem video compression standards like High-Efficiency Video Coding (HEVC), Versatile Video Coding (VVC) and Essential Video Coding (EVC) also use transformed representation to code residual signal after prediction.

[0006] Recently, neural network architectures have been applied to image and / or video coding. In general, these neural network (NN) based approaches can be applied in various different ways to the image and video coding. For example, some end-to-end optimized image or video coding frameworks have been discussed. Moreover, deep learning has been used to determine or optimize some parts of the end-to-end coding framework such as selection or compression of prediction parameters or the like. Besides, some neural network based approached have also been discussed for usage in hybrid image and video coding frameworks, e.g. for implementation as a trained deep learning model for intra or inter prediction in image or video coding.

[0007] The end-to-end optimized image or video coding applications discussed above have in common that they produce some feature map data, which is to be conveyed between encoder and decoder.

[0008] Neural networks are machine learning models that employ one or more layers of nonlinear units based on which they can predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. A corresponding feature map may be provided as an output of each hidden layer. Such corresponding feature map of each hidden layer may be used as an input to a subsequent layer in the network, i.e., a subsequent hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters. In a neural network that is split between devices, e.g. between encoder and decoder, a device and a cloud or between different devices, a feature map at the output of the place of splitting (e.g. a first device) is compressed and transmitted to the remaining layers of the neural network (e.g. to a second device).

[0009] Further improvement of encoding and decoding using trained network architectures may be desirable.

[0010] The foregoing and other objects are achieved by the subject matter of the independent claims. Further implementation forms are apparent from the dependent claims, the description and the figures.

[0011] Particular embodiments are outlined in the attached independent claims, with other embodiments in the dependent claims.

[0012] SUMMARY

[0013] According to a first aspect, the present disclosure relates to an image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processors) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0014] The foregoing and other embodiments can each optionally include one or more of the following features, alone or in combination.

[0015] In a possible implementation the processor(s) is / are configured to form the data units such that in the first dimension of the image one tile represents an edge region of the image and padding data. Padding an edge tile can be an effective way to achieve that the tiles are of the same extent.

[0016] Optionally the padding data is of a constant value. This can make it easier or more efficient to populate the padding data.

[0017] Optionally the padding data repeats data of the edge region of the image extending in a second dimension orthogonal to the first dimension. This can help to reduce artefacts when the image is recovered.

[0018] Optionally the processor(s) is / are configured to form the encoded image data to include an indication of the extent of the padding data in the first dimension. This can allow the image to be recovered as intended.

[0019] Optionally the processors) is / are configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. One or more of the overlaps can be adjusted so the tiles fit exactly across the image in the said dimension.

[0020] Optionally one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image. In this way the overlap of an edge tile can be adjusted to fit the image. Alternatively, the overlap(s) of a tile that is not at the edge can be adjusted.

[0021] Optionally the processors) is / are configured to form the encoded image data to include an indication of the width of the overlap between the tiles of the first pair. This can allow an entity that is recovering the image to ascertain the positioning of the tile(s).

[0022] Optionally the processor(s) is / are configured to form the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set. This can help in recovering the image.

[0023] Optionally the processor(s) is / are configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. This can allow the tiles to fit the image in both directions.

[0024] Optionally the processor s) is / are configured to perform an encoding or latent prediction operation on the tiles to form the data units. Using tiles of the same size for such an operation can be efficient.

[0025] Optionally the latent prediction operation comprises one of a hyper-encoding operation and a multistage context modelling operation. Using tiles of the same size for such an operation can be efficient. Optionally the processors) is / are configured to determine whether to form the data units with padding or to align with the edge of the image in dependence on the extent of the image in the first dimension. This can be an efficient way to select the operating mode.

[0026] According to a second aspect, the present disclosure relates to an image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processors) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and one tile of the set represents an edge region of the image and padding data such that all the tiles of the set are of the same extent in the first dimension.

[0027] According to a third aspect, the present disclosure relates to an image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processors) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and the overlap between a first pair of two tiles of the set that are adjacent to each other is larger than the overlap between a second pair of two tiles of the set that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension.

[0028] According to a fourth aspect, the present disclosure relates to a method for encoding an image to form encoded image data, the method comprising processing an image to form a series of data units in a latent space, each data unit representing a respective tile in the image, such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0029] The foregoing and other embodiments can each optionally include one or more of the following features, alone or in combination.

[0030] Optionally the method comprises forming the data units such that in the first dimension of the image one tile represents an edge region of the image and padding data. Padding an edge tile can be an effective way to achieve that the tiles are of the same extent.

[0031] Optionally the padding data is of a constant value. This can make it easier or more efficient to populate the padding data.

[0032] Optionally the padding data repeats data of the edge region of the image extending in a second dimension orthogonal to the first dimension. This can help to reduce artefacts when the image is recovered.

[0033] Optionally the method comprises forming the encoded image data to include an indication of the extent of the padding data in the first dimension. This can allow the image to be recovered as intended.

[0034] Optionally the method comprises forming the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. One or more of the overlaps can be adjusted so the tiles fit exactly across the image in the said dimension. Optionally one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image. In this way the overlap of an edge tile can be adjusted to fit the image. Alternatively, the overlap(s) of a tile that is not at the edge can be adjusted.

[0035] Optionally the comprises forming the encoded image data to include an indication of the width of the overlap between the tiles of the first pair. This can allow an entity that is recovering the image to ascertain the positioning of the tile(s).

[0036] Optionally the method comprises forming the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set. This can help in recovering the image.

[0037] Optionally the method comprises forming the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. This can allow the tiles to fit the image in both directions.

[0038] Optionally the method comprises performing an encoding or latent prediction operation on the tiles to form the data units. Using tiles of the same size for such an operation can be efficient.

[0039] Optionally the latent prediction operation comprises one of a hyper-encoding operation and a multistage context modelling operation. Using tiles of the same size for such an operation can be efficient.

[0040] Optionally the method comprises determining whether to form the data units with padding or to align with the edge of the image in dependence on the extent of the image in the first dimension. This can be an efficient way to select the operating mode.

[0041] According to a fifth aspect, the present disclosure relates to an image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0042] The foregoing and other embodiments can each optionally include one or more of the following features, alone or in combination.

[0043] Optionally the processors) is / are configured to form the data units such that in the first dimension of the image one tile represents an edge region of the image space and padding data. Padding an edge tile can be an effective way to achieve that the tiles are of the same extent. The padding may be performed at the decoder.

[0044] Optionally the padding data is of a constant value. This can make it easier or more efficient to populate the padding data.

[0045] Optionally the padding data repeats data of the edge region extending in a second dimension orthogonal to the first dimension. This can help to reduce artefacts when the image is recovered.

[0046] Optionally the processors) is / are configured to form the encoded image data to determine an extent of the padding data in the first dimension either (i) in accordance with a value signalled in the encoded image data or (ii) in dependence on an extent of the image in the first dimension signalled in the encoded image data. This can allow the size of the padding to be determined effectively.

[0047] Optionally the processor(s) is / are configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. One or more of the overlaps can be adjusted so the tiles fit exactly across the image in the said dimension. The or each adjusted overlap can be decided by the decoder or signalled to the decoder as part of the encoded data.

[0048] Optionally one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image. Alternatively, the overlap(s) of a tile that is not at the edge can be adjusted.

[0049] Optionally the processors) is / are configured to determine the width of the overlap between the tiles of the first pair in accordance with a value signalled in the encoded image data. This can be efficient for the decoder.

[0050] Optionally the processor(s) is / are configured to select as the said one of the tiles of the set a tile whose identity is signalled as such in the encoded image data. This can be efficient for the decoder.

[0051] Optionally the processors) is / are configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. This can allow the tiles to fit the image in both directions.

[0052] Optionally the processor(s) is / are configured to perform a decoding or latent prediction operation on the tiles to form the data units. Using tiles of the same size for such an operation can be efficient.

[0053] Optionally the latent prediction operation comprises one of a hyper-decoding operation and a multistage context modelling operation. Using tiles of the same size for such an operation can be efficient.

[0054] Optionally the processors) is / are configured to determine whether to form the data units with padding or to align with the edge of the image in dependence on the extent of the image in the first dimension. This can be an efficient way to select the operating mode.

[0055] According to a sixth aspect, the present disclosure relates to an image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension of the image one tile representing an edge region of the image space and padding data such that all the tiles of the set are of the same extent in the first dimension.

[0056] According to a seventh aspect the present disclosure relates to an image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension. According to an eighth aspect the present disclosure relates to method for decoding encoded image data to recover an image by processing data units in a latent space, each data unit representing a respective tile in the image, the method being such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0057] According to a ninth aspect the present disclosure relates to a data carrier storing data defining instructions for one or more processors to perform the steps set out above. The data carrier may store the data in non-transient form.

[0058] According to a tenth aspect the present disclosure relates to a computer readable storage medium having stored thereon data defining instructions for one or more processors to perform the steps set out above. The steps may be performed to encode or decode data, as appropriate. The storage medium may store the data in non-transient form.

[0059] According to an eleventh aspect, the present disclosure relates to a computer program product including program code for performing the methods set out above when executed on a computer.

[0060] Each tile may have an equal extent in two orthogonal pixel directions. Thus, the tile may be square. This can efficiently conform to existing decoding algorithms.

[0061] BRIEF DESCRIPTION OF THE DRAWINGS

[0062] In the following embodiments of the present disclosure are described in more detail with reference to the attached figures and drawings, in which:

[0063] Fig. 1 is a schematic drawing illustrating channels processed by layers of a neural network;

[0064] Fig. 2 is a schematic drawing illustrating an autoencoder type of a neural network;

[0065] Fig. 3 A is a schematic drawing illustrating an exemplary network architecture for encoder and decoder side including a hyperprior model;

[0066] Fig. 3B is a schematic drawing illustrating a general network architecture for encoder side including a hyperprior model;

[0067] Fig. 3C is a schematic drawing illustrating a general network architecture for decoder side including a hyperprior model;

[0068] Fig. 4 is a schematic drawing illustrating an exemplary network architecture for encoder and decoder side including a hyperprior model;

[0069] Fig. 5 is a block diagram illustrating a structure of a cloud-based solution for machine-based tasks such as machine vision tasks;

[0070] Fig. 6A is a block diagram illustrating end-to-end video compression framework based on a neural networks;

[0071] Fig. 6B is a block diagram illustrating some exemplary details of application of a neural network for motion field compression;

[0072] Fig. 6C is a block diagram illustrating some exemplary details of application of a neural network for motion compensation;

[0073] Fig. 7 is a block diagram illustrating an example of an encoding apparatus or a decoding apparatus;

[0074] Fig. 8 is a block diagram illustrating another example of an encoding apparatus or a decoding apparatus.

[0075] Fig. 9 is a schematic drawing illustrating an embodiment network architecture for one component of an encoder side including a hyperprior model;

[0076] Fig. 10 is a schematic drawing illustrating an embodiment network architecture for one component of a decoder side including a hyperprior model;

[0077] Fig. 11 is a block diagram of an encoder and decoder with respective modules, including NN-based subnetworks for processing a first and a second plurality of tiles; Fig. 12 shows an example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples in the overlap region, and the concatenation of the cropped regions. Each Li comprises the total receptive field;

[0078] Fig. 13 shows another example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles) similar to Fig. 12, except that cropping is dismissed;

[0079] Fig. 14 shows an example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples in the overlap region and the concatenation of the cropped regions. Each Li comprises a subset of the total receptive field;

[0080] Fig. 15 shows an example of dividing a first and / or second tensor into non-overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples, and the concatenation of the cropped regions. Each Li comprises a subset of the total receptive field;

[0081] Fig. 16 illustrates the various parameters, such as the sizes of regions Li, Ri, and overlap regions etc. of first and / or second tiles, that may be included into (and parsed from) the bitstream. Any of the various parameters may be included as an indication into (and parsed from) the bitstream;

[0082] Fig. 17 shows a schematic diagram of part of the JPEG Al coding scheme and two variations of the tile scheme which may be used in the image coding process;

[0083] Fig . 18 shows the two variations of the tile scheme of Fig . 17 in more detail;

[0084] Fig. 19 shows an example of a tile having padding outside an image area;

[0085] Fig. 20 shows an example of a tile having overlap selected so that the edge of the tile is coincident with the edge of an image;

[0086] Fig. 21 shows examples of decoding artefacts resulting from the approach of figure 19 (in images A and C) and the approach of figure 20 (in images B and D);

[0087] Fig. 22 is a flow diagram illustrating an exemplary method for encoding image data using a set of tiles with a unique extent of the image data in one tile of the set;

[0088] Fig. 23 is a flow diagram illustrating an exemplary method for decoding image data using a set of tiles with a unique extent of the image data in one tile of the set;

[0089] Fig. 24 shows a device for decoding image data for processing by a neural network-based unit;

[0090] Fig. 25 shows a device for encoding image data for processing by a neural network-based unit;

[0091] Fig. 26 is a block diagram showing an example of a video coding system configured to implement embodiments of the present disclosure;

[0092] Fig. 27 is a block diagram showing another example of a video coding system configured to implement embodiments of the present disclosure;

[0093] Fig. 28 is a block diagram illustrating an example of an encoding apparatus or a decoding apparatus;

[0094] Fig. 29 is a block diagram illustrating another example of an encoding apparatus or a decoding apparatus; and

[0095] Fig. 30 is a block diagram illustrating another example of an encoding apparatus or a decoding apparatus.

[0096] Like reference numbers and designations in different drawings may indicate similar elements.

[0097] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0098] In the following description, reference is made to the accompanying figures, which form part of the disclosure, and which show, by way of illustration, specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that embodiments of the present disclosure may be used in other aspects and comprise structural or logical changes not depicted in the figures. The following detailed description, therefore, is not to be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims. For instance, it is understood that a disclosure in connection with a described method may also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or a plurality of specific method steps are described, a corresponding device may include one or a plurality of units, e.g. functional units, to perform the described one or plurality of method steps (e.g. one unit performing the one or plurality of steps, or a plurality of units each performing one or more of the plurality of steps), even if such one or more units are not explicitly described or illustrated in the figures. On the other hand, for example, if a specific apparatus is described based on one or a plurality of units, e.g. functional units, a corresponding method may include one step to perform the functionality of the one or plurality of units (e.g. one step performing the functionality of the one or plurality of units, or a plurality of steps each performing the functionality of one or more of the plurality of units), even if such one or plurality of steps are not explicitly described or illustrated in the figures. Further, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless specifically noted otherwise.

[0099] In the following, an overview over some of the used technical terms and framework within which the embodiments of the present disclosure may be employed is provided.

[0100] Artificial neural networks

[0101] Artificial neural networks (ANN) or connectionist systems are computing systems vaguely inspired by the biological neural networks that constitute animal brains. Such systems "learn" to perform tasks by considering examples, generally without being programmed with task-specific rules. For example, in image recognition, they might learn to identify images that contain cats by analyzing example images that have been manually labelled as "cat" or "no cat" and using the results to identify cats in other images. They do this without any prior knowledge of cats, for example, that they have fur, tails, whiskers and cat-like faces. Instead, they automatically generate identifying characteristics from the examples that they process.

[0102] An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, can transmit a signal to other neurons. An artificial neuron that receives a signal then processes it and can signal neurons connected to it.

[0103] In ANN implementations, the "signal" at a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. The connections are called edges. Neurons and edges typically have a weight that adjusts as learning proceeds. The weight increases or decreases the strength of the signal at a connection. Neurons may have a threshold such that a signal is sent only if the aggregate signal crosses that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (the input layer), to the last layer (the output layer), possibly after traversing the layers multiple times.

[0104] The original goal of the ANN approach was to solve problems in the same way that a human brain would. Over time, attention moved to performing specific tasks, leading to deviations from biology. ANNs have been used on a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even in activities that have traditionally been considered as reserved to humans, like painting.

[0105] The name “convolutional neural network” (CNN) indicates that the network employs a mathematical operation called convolution. Convolution is a specialized kind of linear operation. Convolutional networks are neural networks that use convolution in place of a general matrix multiplication in at least one of their layers.

[0106] Fig. 1 schematically illustrates a general concept of processing by a neural network such as the CNN. A convolutional neural network consists of an input and an output layer, as well as multiple hidden layers. Input layer is the layer to which the input (such as a portion 11 of an input image as shown in Fig. 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that convolve with a multiplication or other dot product. The result of a layer is one or more feature maps (illustrated by empty solid-line rectangles), sometimes also referred to as channels. There may be a resampling (such as subsampling) involved in some or all of the layers. As a consequence, the feature maps may become smaller, as illustrated in Fig. 1. It is noted that a convolution with a stride may also reduce the size (resample) an input feature map. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer or Leaky ReLU, and is subsequently followed by additional convolutions such as pooling layers, fully connected layers and normalization layers, referred to as hidden layers because their inputs and outputs are masked by the activation function and final convolution. Though the layers are colloquially referred to as convolutions, this is only by convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has significance for the indices in the matrix, in that it affects how the weight is determined at a specific index point.

[0107] When programming a CNN for processing images, as shown in Fig. 1, the input is a tensor with shape (number of images) x (image width) x (image height) x (image depth). It should be known that the image depth can be constituted by channels of an image. After passing through a convolutional layer, the image becomes abstracted to a feature map, with shape (number of images) x (feature map width) x (feature map height) x (feature map channels). A convolutional layer within a neural network should have the following attributes. Convolutional kernels defined by a width and height (hyper-parameters). The number of input channels and output channels (hyper-parameter). The depth of the convolution filter (the input channels) should be equal to the number channels (depth) of the input feature map.

[0108] In the past, traditional multilayer perceptron (MLP) models have been used for image recognition. However, due to the full connectivity between nodes, they suffered from high dimensionality, and did not scale well with higher resolution images. A 1000xl000-pixel image with RGB color channels has 3 million weights, which is too high to feasibly process efficiently at scale with full connectivity. Also, such network architecture does not take into account the spatial structure of data, treating input pixels which are far apart in the same way as pixels that are close together. This ignores locality of reference in image data, both computationally and semantically. Thus, full connectivity of neurons is wasteful for purposes such as image recognition that are dominated by spatially local input patterns.

[0109] Convolutional neural networks are biologically inspired variants of multilayer perceptrons that are specifically designed to emulate the behavior of a visual cortex. These models mitigate the challenges posed by the MLP architecture by exploiting the strong spatially local correlation present in natural images. The convolutional layer is the core building block of a CNN. The layer's parameters consist of a set of learnable filters (the above-mentioned kernels), which have a small receptive field, but extend through the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, computing the dot product between the entries of the filter and the input and producing a 2-dimensional activation map of that filter. As a result, the network learns filters that activate when it detects some specific type of feature at some spatial position in the input.

[0110] Stacking the activation maps for all filters along the depth dimension forms the full output volume of the convolution layer. Every entry in the output volume can thus also be interpreted as an output of a neuron that looks at a small region in the input and shares parameters with neurons in the same activation map. A feature map, or activation map, is the output activations for a given filter. Feature map and activation has same meaning. In some papers it is called an activation map because it is a mapping that corresponds to the activation of different parts of the image, and also a feature map because it is also a mapping of where a certain kind of feature is found in the image. A high activation means that a certain feature was found. Another important concept of CNNs is pooling, which is a form of non-linear down- sampling. There are several non-linear functions to implement pooling among which max pooling is the most common. It partitions the input image into a set of nonoverlapping rectangles and, for each such sub-region, outputs the maximum.

[0111] Intuitively, the exact location of a feature is less important than its rough location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. The pooling layer serves to progressively reduce the spatial size of the representation, to reduce the number of parameters, memory footprint and amount of computation in the network, and hence to also control overfitting. It is common to periodically insert a pooling layer between successive convolutional layers in a CNN architecture. The pooling operation provides another form of translation invariance.

[0112] The pooling layer operates independently on every depth slice of the input and resizes it spatially. The most common form is a pooling layer with filters of size 2x2 applied with a stride of 2 at every depth slice in the input by 2 along both width and height, discarding 75% of the activations. In this case, every max operation is over 4 numbers. The depth dimension remains unchanged. In addition to max pooling, pooling units can use other functions, such as average pooling or £2-norm pooling. Average pooling was often used historically but has recently fallen out of favour compared to max pooling, which often performs better in practice. Due to the aggressive reduction in the size of the representation, there is a recent trend towards using smaller filters or discarding pooling layers altogether. "Region of Interest" pooling (also known as ROI pooling) is a variant of max pooling, in which output size is fixed and input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on Fast R-CNN architecture.

[0113] The above-mentioned ReLU is the abbreviation of rectified linear unit, which applies the non- saturating activation function. It effectively removes negative values from an activation map by setting them to zero. It increases the nonlinear properties of the decision function and of the overall network without affecting the receptive fields of the convolution layer. Other functions are also used to increase nonlinearity, for example the saturating hyperbolic tangent and the sigmoid function. ReLU is often preferred to other functions because it trains the neural network several times faster without a significant penalty to generalization accuracy.

[0114] Leaky Rectified Linear Unit, or Leaky ReLU, is a type of activation function based on a ReLU, but it has a small slope for negative values instead of a flat slope. The slope coefficient is determined before training, i.e. it is not learnt during training. This type of activation function is popular in tasks where it suffers from sparse gradients, for example training generative adversarial networks. Leaky ReLU applies the element- wise function:

[0115] LeakyReLU(x)=max(0,x)+negative_slope*min(0,x), or

[0116] Leaky

[0117] JR

[0118] Among them, parameters: negative_slope - Controls the angle of the negative slope. Default: le-2 inplace - can optionally do the operation in-place. Default: False.

[0119] After several convolutional and max pooling layers, the high-level reasoning in the neural network is done via fully connected layers. Neurons in a fully connected layer have connections to all activations in the previous layer, as seen in regular (non- convolutional) artificial neural networks. Their activations can thus be computed as an affine transformation, with matrix multiplication followed by a bias offset (vector addition of a learned or fixed bias term).

[0120] The "loss layer" (including calculating of a loss function) specifies how training penalizes the deviation between the predicted (output) and true labels and is normally the final layer of a neural network. Various loss functions appropriate for different tasks may be used. Softmax loss is used for predicting a single class of K mutually exclusive classes. Sigmoid cross-entropy loss is used for predicting K independent probability values in [0, 1], Euclidean loss is used for regressing to real- valued labels.

[0121] In summary, Fig. 1 shows the data flow in a typical convolutional neural network. First, the input image is passed through convolutional layers and becomes abstracted to a feature map comprising several channels, corresponding to a number of filters in a set of learnable filters of this layer. Then, the feature map is subsampled using e.g. a pooling layer, which reduces the dimension of each channel in the feature map. Next, the data comes to another convolutional layer, which may have different numbers of output channels. As was mentioned above, the number of input channels and output channels are hyper-parameters of the layer. To establish connectivity of the network, those parameters need to be synchronized between two connected layers, such that the number of input channels for the current layers should be equal to the number of output channels of the previous layer. For the first layer which processes input data, e.g. an image, the number of input channels is normally equal to the number of channels of data representation, for instance 3 channels for RGB or YUV representation of images or video, or 1 channel for grayscale image or video representation. The channels obtained by one or more convolutional layers (and possibly resampling layer(s)) may be passed to an output layer. Such output layer may be a convolutional or resampling in some implementations. In an exemplary and non-limiting implementation, the output layer is a fully connected layer.

[0122] Autoencoders and unsupervised learning

[0123] An autoencoder is a type of artificial neural network used to learn efficient data codings in an unsupervised manner. A schematic drawing thereof is shown in Fig. 2. The autoencoder includes an encoder side 210 with an input x inputted into an input layer of an encoder subnetwork 220 and a decoder side 250 with output x’ outputted from a decoder subnetwork 260. The aim of an autoencoder is to learn a representation (encoding) 230 for a set of data x, typically for dimensionality reduction, by training the network 220, 260 to ignore signal “noise”. Along with the reduction (encoder) side subnetwork 220, a reconstructing (decoder) side subnetwork 260 is learnt, where the autoencoder tries to generate from the reduced encoding 230 a representation x’ as close as possible to its original input x, hence its name. In the simplest case, given one hidden layer, the encoder stage of an autoencoder takes the input x and maps it to h: h = ct(Wx + b).

[0124] This image h is usually referred to as code 230, latent variables, or latent representation. Here, cr is an element- wise activation function such as a sigmoid function or a rectified linear unit. W is a weight matrix b is a bias vector. Weights and biases are usually initialized randomly, and then updated iteratively during training through Backpropagation. After that, the decoder stage of the autoencoder maps h to the reconstruction x'of the same shape as x: x' = ct'(W'h' + b') where a’ , W' and b' for the decoder may be unrelated to the corresponding a, W and b for the encoder.

[0125] Variational autoencoder models make strong assumptions concerning the distribution of latent variables. They use a variational approach for latent representation learning, which results in an additional loss component and a specific estimator for the training algorithm called the Stochastic Gradient Variational Bayes (SGVB) estimator. It assumes that the data is generated by a directed graphical model pe(x|h) and that the encoder is learning an approximation q(|)(h|x) to the posterior distribution pe(h| x) where <f> and 0 denote the parameters of the encoder (recognition model) and decoder (generative model) respectively. The probability distribution of the latent vector of a VAE typically matches that of the training data much closer than a standard autoencoder. The objective of VAE has the following form: where p(x) and a>2(x) are the encoder output, while p( / i) and cr2(h) are the decoder outputs.

[0126] Recent progress in artificial neural networks area and especially in convolutional neural networks enables researchers’ interest of applying neural networks-based technologies to the task of image and video compression. For example, End-to-end Optimized Image Compression has been proposed, which uses a network based on a variational autoencoder.

[0127] Accordingly, data compression is considered as a fundamental and well-studied problem in engineering, and is commonly formulated with the goal of designing codes for a given discrete data ensemble with minimal entropy. The solution relies heavily on knowledge of the probabilistic structure of the data, and thus the problem is closely related to probabilistic source modeling. However, since all practical codes must have finite entropy, continuous- valued data (such as vectors of image pixel intensities) must be quantized to a finite set of discrete values, which introduces an error.

[0128] In this context, known as the lossy compression problem, one must trade off two competing costs: the entropy of the discretized representation (rate) and the error arising from the quantization (distortion). Different compression applications, such as data storage or transmission over limited-capacity channels, demand different rate-distortion trade-offs.

[0129] Joint optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional spaces is intractable. For this reason, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous- valued representation, quantizing its elements independently, and then encoding the resulting discrete representation using a lossless entropy code. This scheme is called transform coding due to the central role of the transformation.

[0130] For example, JPEG uses a discrete cosine transform on blocks of pixels, and JPEG 2000 uses a multi-scale orthogonal wavelet decomposition. Typically, the three components of transform coding methods - transform, quantizer, and entropy code - are separately optimized (often through manual parameter adjustment). Modem video compression standards like HEVC, VVC and EVC also use transformed representation to code residual signal after prediction. The several transforms are used for that purpose such as discrete cosine and sine transforms (DCT, DST), as well as low frequency non-separable manually optimized transforms (LFNST). Variational image compression

[0131] Variable Auto-Encoder (VAE) framework can be considered as a nonlinear transforming coding model. The transforming process can be mainly divided into four parts. This is exemplified in Fig. 3A showing a VAE framework.

[0132] The transforming process can be mainly divided into four parts: Fig. 3A exemplifies the VAE framework. In Fig. 3A, the encoder 101 maps an input image x into a latent representation (denoted by y) via the function y = f (x). This latent representation may also be referred to as a part of or a point within a “latent space” in the following. The function f() is a transformation function that converts the input signal x into a more compressible representation y. The quantizer 102 transforms the latent representation y into the quantized latent representation y with (discrete) values by y = Q(y). with Q representing the quantizer function. The entropy model, or the hyper encoder / decoder (also known as hyperprior) 103 estimates the distribution of the quantized latent representation y to get the minimum rate achievable with a lossless entropy source coding.

[0133] The latent space can be understood as a representation of compressed data in which similar data points are closer together in the latent space. Latent space is useful for learning data features and for finding simpler representations of data for analysis. The quantized latent representation T, y and the side information z of the hyperprior 3 are included into a bitstream 2 (are binarized) using arithmetic coding (AE). Furthermore, a decoder 104 is provided that transforms the quantized latent representation to the reconstructed image x, x = g(y). The signal x is the estimation of the input image x. It is desirable that x is as close to x as possible, in other words the reconstruction quality is as high as possible. However, the higher the similarity between x and x, the higher the amount of side information necessary to be transmitted. The side information includes bitstreaml and bitstream2 shown in Fig. 3A, which are generated by the encoder and transmitted to the decoder. Normally, the higher the amount of side information, the higher the reconstruction quality . However, a high amount of side information means that the compression ratio is low. Therefore, one purpose of the system described in Fig. 3A is to balance the reconstruction quality and the amount of side information conveyed in the bitstream.

[0134] In Fig. 3A the component AE 105 is the Arithmetic Encoding module, which converts samples of the quantized latent representation y and the side information z into a binary representation bitstream 1. The samples of y and z might for example comprise integer or floating-point numbers. One purpose of the arithmetic encoding module is to convert (via the process of binarization) the sample values into a string of binary digits (which is then included in the bitstream that may comprise further portions corresponding to the encoded image or further side information).

[0135] The arithmetic decoding (AD) 106 is the process of reverting the binarization process, where binary digits are converted back to sample values. The arithmetic decoding is provided by the arithmetic decoding module 106.

[0136] It is noted that the present disclosure is not limited to this particular framework. Moreover, the present disclosure is not restricted to image or video compression, and can be applied to object detection, image generation, and recognition systems as well.

[0137] In Fig. 3 A there are two sub networks concatenated to each other. A subnetwork in this context is a logical division between the parts of the total network. For example, in Fig. 3A the modules 101, 102, 104, 105 and 106 are called the “Encoder / Decoder” subnetwork. The “Encoder / Decoder” subnetwork is responsible for encoding (generating) and decoding (parsing) of the first bitstream “bitstreaml”. The second network in Fig. 3A comprises modules 103, 108, 109, 110 and 107 and is called “hyper encoder / decoder” subnetwork. The second subnetwork is responsible for generating the second bitstream “bitstream2”. The purposes of the two subnetworks are different. The first subnetwork is responsible for:

[0138] • the transformation 101 of the input image x into its latent representation y (which is easier to compress that x),

[0139] • quantizing 102 the latent representation y into a quantized latent representation y,

[0140] • compressing the quantized latent representation y using the AE by the arithmetic encoding module 105 to obtain bitstream “bitstream 1”,”.

[0141] • parsing the bitstream 1 via AD using the arithmetic decoding module 106, and

[0142] • reconstructing 104 the reconstructed image (x) using the parsed data.

[0143] The purpose of the second subnetwork is to obtain statistical properties (e.g. mean value, variance and correlations between samples of bitstream 1) of the samples of “bitstream 1”, such that the compressing of bitstream 1 by first subnetwork is more efficient. The second subnetwork generates a second bitstream “bitstream2”, which comprises the said information (e.g. mean value, variance and correlations between samples of bitstreaml).

[0144] The second network includes an encoding part which comprises transforming 103 of the quantized latent representation y into side information z, quantizing the side information z into quantized side information z, and encoding (e.g. binarizing) 109 the quantized side information z into bitstream2. In this example, the binarization is performed by an arithmetic encoding (AE). A decoding part of the second network includes arithmetic decoding (AD) 110, which transforms the input bitstream2 into decoded quantized side information z'. The z' might be identical to z, since the arithmetic encoding end decoding operations are lossless compression methods. The decoded quantized side information z' is then transformed 107 into decoded side information y'. y’ represents the statistical properties of y (e.g. mean value of samples of y, or the variance of sample values or like). The decoded latent representation y' is then provided to the above-mentioned Arithmetic Encoder 105 and Arithmetic Decoder 106 to control the probability model of y.

[0145] The Fig. 3 A describes an example of VAE (variational auto encoder), details of which might be different in different implementations. For example, in a specific implementation additional components might be present to more efficiently obtain the statistical properties of the samples of bitstream 1. In one such implementation a context modeler might be present, which targets extracting cross-correlation information of the bitstream 1. The statistical information provided by the second subnetwork might be used by AE (arithmetic encoder) 105 and AD (arithmetic decoder) 106 components.

[0146] Fig. 3A depicts the encoder and decoder in a single figure. As is clear to those skilled in the art, the encoder and the decoder may be, and very often are, embedded in mutually different devices.

[0147] Fig. 3B depicts the encoder and Fig. 3C depicts the decoder components of the VAE framework in isolation. As input, the encoder receives, according to some embodiments, a picture. The input picture may include one or more channels, such as color channels or other kind of channels, e.g. depth channel or motion information channel, or the like. The output of the encoder (as shown in Fig. 3B) is a bitstreaml and a bitstream2. The bitstreaml is the output of the first sub-network of the encoder and the bitstream is the output of the second subnetwork of the encoder.

[0148] Similarly, in Fig. 3C, the two bitstreams, bitstreaml and bitstreaml, are received as input and z, which is the reconstructed (decoded) image, is generated at the output. As indicated above, the VAE can be split into different logical units that perform different actions. This is exemplified in Figs. 3B and 3C so that Fig. 3B depicts components that participate in the encoding of a signal, like a video and provided encoded information. This encoded information is then received by the decoder components depicted in Fig. 3C for encoding, for example. It is noted that the components of the encoder and decoder denoted with numerals 12x and 14x may correspond in their function to the components referred to above in Fig. 3 A and denoted with numerals lOx. Specifically, as is seen in Fig. 3B, the encoder comprises the encoder 121 that transforms an input x into a signal y which is then provided to the quantizer 322. The quantizer 122 provides information to the arithmetic encoding module 125 and the hyper encoder 123. The hyper encoder 123 provides the bitstream2 already discussed above to the hyper decoder 147 that in turn provides the information to the arithmetic encoding module 105 (125).

[0149] The output of the arithmetic encoding module is the bitstreaml. The bitstreaml and bitstream2 are the output of the encoding of the signal, which are then provided (transmitted) to the decoding process. Although the unit 101 (121) is called “encoder”, it is also possible to call the complete subnetwork described in Fig. 3B as “encoder”. The process of encoding in general means the unit (module) that converts an input to an encoded (e.g. compressed) output. It can be seen from Fig. 3B, that the unit 121 can be actually considered as a core of the whole subnetwork, since it performs the conversion of the input x into y, which is the compressed version of the x. The compression in the encoder 121 may be achieved, e.g. by applying a neural network, or in general any processing network with one or more layers. In such network, the compression may be performed by cascaded processing including downsampling which reduces size and / or number of channels of the input. Thus, the encoder may be referred to, e.g. as a neural network (NN) based encoder, or the like.

[0150] The remaining parts in the figure (quantization unit, hyper encoder, hyper decoder, arithmetic encoder / decoder) are all parts that either improve the efficiency of the encoding process or are responsible for converting the compressed output y into a series of bits (bitstream). Quantization may be provided to further compress the output of the NN encoder 121 by a lossy compression. The AE 125 in combination with the hyper encoder 123 and hyper decoder 127 used to configure the AE 125 may perform the binarization which may further compress the quantized signal by a lossless compression. Therefore, it is also possible to call the whole subnetwork in Fig. 3B an “encoder”.

[0151] A majority of Deep Learning (DL) based image / video compression systems reduce dimensionality of the signal before converting the signal into binary digits (bits). In the VAE framework for example, the encoder, which is a non-linear transform, maps the input image x into y, where y has a smaller width and height than x. Since the y has a smaller width and height, hence a smaller size, the (size of the) dimension of the signal is reduced, and, hence, it is easier to compress the signal y. It is noted that in general, the encoder does not necessarily need to reduce the size in both (or in general all) dimensions. Rather, some exemplary implementations may provide an encoder which reduces size only in one (or in general a subset of) dimension.

[0152] In J. Balle, L. Valero Laparra, and E. P. Simoncelli (2015). “Density Modeling of Images Using a Generalized Normalization Transformation”, In: arXiv e-prints, Presented at the 4th Int. Conf, for Learning Representations, 2016 (referred to in the following as “Balle”) the authors proposed a framework for end-to-end optimization of an image compression model based on nonlinear transforms. The authors optimize for Mean Squared Error (MSE), but use a more flexible transforms built from cascades of linear convolutions and nonlinearities. Specifically, authors use a generalized divisive normalization (GDN) joint nonlinearity that is inspired by models of neurons in biological visual systems, and has proven effective in Gaussianizing image densities. This cascaded transformation is followed by uniform scalar quantization (i.e., each element is rounded to the nearest integer), which effectively implements a parametric form of vector quantization on the original image space. The compressed image is reconstructed from these quantized values using an approximate parametric nonlinear inverse transform.

[0153] Such example of the VAE framework is shown in Fig. 4, and it utilizes 6 downsampling layers that are marked with 401 to 406. The network architecture includes a hyperprior model. The left side (ga, gs) shows an image autoencoder architecture, the right side (ha, hs) corresponds to the autoencoder implementing the hyperprior. The factorized-prior model uses the identical architecture for the analysis and synthesis transforms gaand gs. Q represents quantization, and AE, AD represent arithmetic encoder and arithmetic decoder, respectively. The encoder subjects the input image x to ga, yielding the responses y (latent representation) with spatially varying standard deviations. The encoding gaincludes a plurality of convolution layers with subsampling and, as an activation function, generalized divisive normalization (GDN).

[0154] The responses are fed into ha, summarizing the distribution of standard deviations in z. z is then quantized, compressed, and transmitted as side information. The encoder then uses the quantized vector z to estimate a, the spatial distribution of standard deviations which is used for obtaining probability values (or frequency values) for arithmetic coding (AE), and uses it to compress and transmit the quantized image representation y (or latent representation). The decoder first recovers z from the compressed signal. It then uses hsto obtain y, which provides it with the correct probability estimates to successfully recover y as well. It then feeds y into gsto obtain the reconstructed image.

[0155] The layers that include downsampling is indicated with the downward arrow in the layer description. The layer description „Conv N,kl,2]“ means that the layer is a convolution layer, with N channels and the convolution kernel is klxkl in size. For example, kl may be equal to 5 and k2 may be equal to 3. As stated, the 2 (means that a downsampling with a factor of 2 is performed in this layer. Downsampling by a factor of 2 results in one of the dimensions of the input signal being reduced by half at the output. In Fig. 4, the 2 (indicates that both width and height of the input image is reduced by a factor of 2. Since there are 6 downsampling layers, if the width and height of the input image 414 (also denoted with x) is given by w and h, the output signal z '413 is has width and height equal to w / 64 and h / 64 respectively. Modules denoted by AE and AD are arithmetic encoder and arithmetic decoder, which are explained with reference to Figs. 3A to 3C. The arithmetic encoder and decoder are specific implementations of entropy coding. AE and AD can be replaced by other means of entropy coding. In information theory, an entropy encoding is a lossless data compression scheme that is used to convert the values of a symbol into a binary representation which is a revertible process. Also, the “Q” in the figure corresponds to the quantization operation that was also referred to above in relation to Fig. 4 and is further explained above in the section “Quantization”. Also, the quantization operation and a corresponding quantization unit as part of the component 413 or 415 is not necessarily present and / or can be replaced with another unit.

[0156] In Fig. 4, there is also shown the decoder comprising upsampling layers 407 to 412. A further layer 420 is provided between the upsampling layers 411 and 410 in the processing order of an input that is implemented as convolutional layer but does not provide an upsampling to the input received. A corresponding convolutional layer 430 is also shown for the decoder. Such layers can be provided in NNs for performing operations on the input that do not alter the size of the input but change specific characteristics. However, it is not necessary that such a layer is provided.

[0157] When seen in the processing order of bitstream2 through the decoder, the upsampling layers are run through in reverse order, i.e. from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide an upsampling with an upsampling ratio of 2, which is indicated by the f. It is, of course, not necessarily the case that all upsampling layers have the same upsampling ratio and also other upsampling ratios like 3, 4, 8 or the like may be used. The layers 407 to 412 are implemented as convolutional layers (conv). Specifically, as they may be intended to provide an operation on the input that is reverse to that of the encoder, the upsampling layers may apply a deconvolution operation to the input received so that its size is increased by a factor corresponding to the upsampling ratio. However, the present disclosure is not generally limited to deconvolution and the upsampling may be performed in any other manner such as by bilinear interpolation between two neighboring samples, or by nearest neighbor sample copying, or the like.

[0158] In the first subnetwork, some convolutional layers (401 to 403) are followed by generalized divisive normalization (GDN) at the encoder side and by the inverse GDN (IGDN) at the decoder side. In the second subnetwork, the activation function applied is ReLu. It is noted that the present disclosure is not limited to such implementation and in general, other activation functions may be used instead of GDN or ReLu. Cloud solutions for machine tasks

[0159] The Video Coding for Machines (VCM) is another computer science direction being popular nowadays. The main idea behind this approach is to transmit a coded representation of image or video information targeted to further processing by computer vision (CV) algorithms, like object segmentation, detection and recognition. In contrast to traditional image and video coding targeted to human perception the quality characteristic is the performance of computer vision task, e.g. object detection accuracy, rather than reconstructed quality. This is illustrated in Fig. 5.

[0160] Video Coding for Machines is also referred to as collaborative intelligence and it is a relatively new paradigm for efficient deployment of deep neural networks across the mobile-cloud infrastructure. By dividing the network between the mobile side 510 and the cloud side 590 (e.g. a cloud server), it is possible to distribute the computational workload such that the overall energy and / or latency of the system is minimized. In general, the collaborative intelligence is a paradigm where processing of a neural network is distributed between two or more different computation nodes; for example, devices, but in general, any functionally defined nodes. Here, the term “node” does not refer to the above-mentioned neural network nodes. Rather the (computation) nodes here refer to (physically or at least logically) separate devices / modules, which implement parts of the neural network. Such devices may be different servers, different end user devices, a mixture of servers and / or user devices and / or cloud and / or processor or the like. In other words, the computation nodes may be considered as nodes belonging to the same neural network and communicating with each other to convey coded data within / for the neural network. For example, in order to be able to perform complex computations, one or more layers may be executed on a first device (such as a device on mobile side 510) and one or more layers may be executed in another device (such as a cloud server on cloud side 590). However, the distribution may also be finer and a single layer may be executed on a plurality of devices. In this disclosure, the term “plurality” refers to two or more. In some existing solution, a part of a neural network functionality is executed in a device (user device or edge device or the like) or a plurality of such devices and then the output (feature map) is passed to a cloud. A cloud is a collection of processing or computing systems that are located outside the device, which is operating the part of the neural network. The notion of collaborative intelligence has been extended to model training as well. In this case, data flows both ways: from the cloud to the mobile during back-propagation in training, and from the mobile to the cloud (illustrated in Fig. 5) during forward passes in training, as well as inference.

[0161] Some works presented semantic image compression by encoding deep features and then reconstructing the input image from them. The compression based on uniform quantization was shown, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient, to transmit from the mobile part 510 to the cloud 590 an output of a hidden layer (a deep feature map) 550, rather than sending compressed natural image data to the cloud and perform the object detection using reconstructed images. It may thus be advantageous to compress the data (features) generated by the mobile side 510, which may include a quantization layer 520 for this purpose. Correspondingly, the cloud side 590 may include an inverse quantization layer 560. The efficient compression of feature maps benefits the image and video compression and reconstruction both for human perception and for machine vision. Entropy coding methods, e.g. arithmetic coding is a popular approach to compression of deep features (i.e. feature maps).

[0162] Nowadays, video content contributes to more than 80% internet traffic, and the percentage is expected to increase even further. Therefore, it is critical to build an efficient video compression system and generate higher quality frames at given bandwidth budget. In addition, most video related computer vision tasks such as video object detection or video object tracking are sensitive to the quality of compressed videos, and efficient video compression may bring benefits for other computer vision tasks. Meanwhile, the techniques in video compression are also helpful for action recognition and model compression. However, in the past decades, video compression algorithms rely on hand-crafted modules, e.g., block-based motion estimation and Discrete Cosine Transform (DCT), to reduce the redundancies in the video sequences, as mentioned above. Although each module is well designed, the whole compression system is not end-to-end optimized. It is desirable to further improve video compression performance by jointly optimizing the whole compression system.

[0163] End-to-end image or video compression

[0164] DNN based image compression methods can exploit large scale end-to-end training and highly non-linear transform, which are not used in the traditional approaches. However, it is non-trivial to directly apply these techniques to build an end-to-end learning system for video compression. First, it remains an open problem to learn how to generate and compress the motion information tailored for video compression. Video compression methods heavily rely on motion information to reduce temporal redundancy in video sequences.

[0165] A straightforward solution is to use the learning based optical flow to represent motion information. However, current learning based optical flow approaches aim at generating flow fields as accurate as possible. The precise optical flow is often not optimal for a particular video task. In addition, the data volume of optical flow increases significantly when compared with motion information in the traditional compression systems and directly applying the existing compression approaches to compress optical flow values will significantly increase the number of bits required for storing motion information. Second, it is unclear how to build a DNN based video compression system by minimizing the rate-distortion based objective for both residual and motion information. Rate-distortion optimization (RDO) aims at achieving higher quality of reconstructed frame (i.e., less distortion) when the number of bits (or bit rate) for compression is given. RDO is important for video compression performance. In order to exploit the power of end-to-end training for learning based compression system, the RDO strategy is required to optimize the whole system.

[0166] In Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, Zhiyong Gao; „DVC: An End-to-end Deep Video Compression Framework". Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015, authors proposed the end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual coding.

[0167] Such encoder is illustrated in Figure 6A. In particular, Figure 6A shows an overall structure of end-to-end trainable video compression framework. In order to compress motion information, a CNN was designated to transform the optical flow vLto the corresponding representations mtsuitable for better compression. Specifically, an auto-encoder style network is used to compress the optical flow. The motion vectors (MV) compression network is shown in Figure 6B. The network architecture is somewhat similar to the ga / gs of Figure 4. In particular, the optical flow vLis fed into a series of convolution operation and nonlinear transform including GDN and IGDN. The number of output channels c for convolution (deconvolution) is here exemplarily 128 except for the last deconvolution layer, which is equal to 2 in this example. The kernel size is k, e.g. k=3. Given optical flow with the size of M x N x 2, the MV encoder will generate the motion representation mtwith the size of M / 16xN / 16xl28. Then motion representation is quantized (Q), entropy coded and sent to bitstream as mt. The MV decoder receives the quantized representation mtand reconstruct motion information vLusing MV encoder. In general, the values for k and c may differ from the abovementioned examples as is known from the art.

[0168] Figure 6C shows a structure of the motion compensation part. Here, using previous reconstructed frame xt-i and reconstructed motion information, the warping unit generates the warped frame (normally, with help of interpolation filter such as bi-linear interpolation filter). Then a separate CNN with three inputs generates the predicted picture. The architecture of the motion compensation CNN is also shown in Figure 6C.

[0169] The residual information between the original frame and the predicted frame is encoded by the residual encoder network. A highly non-linear neural network is used to transform the residuals to the corresponding latent representation. Compared with discrete cosine transform in the traditional video compression system, this approach can better exploit the power of non-linear transform and achieve higher compression efficiency.

[0170] From above overview it can be seen that CNN based architecture can be applied both for image and video compression, considering different parts of video framework including motion estimation, motion compensation and residual coding. Entropy coding is popular method used for data compression, which is widely adopted by the industry and is also applicable for feature map compression either for human perception or for computer vision tasks.

[0171] Video Coding for Machines

[0172] The Video Coding for Machines (VCM) is another computer science direction being popular nowadays. The main idea behind this approach is to transmit the coded representation of image or video information targeted to further processing by computer vision (CV) algorithms, like object segmentation, detection and recognition. In contrast to traditional image and video coding targeted to human perception the quality characteristic is the performance of computer vision task, e.g. object detection accuracy, rather than reconstructed quality.

[0173] A recent study proposed a new deployment paradigm called collaborative intelligence, whereby a deep model is split between the mobile and the cloud. Extensive experiments under various hardware configurations and wireless connectivity modes revealed that the optimal operating point in terms of energy consumption and / or computational latency involves splitting the model, usually at a point deep in the network. Today’s common solutions, where the model sits fully in the cloud or fully at the mobile, were found to be rarely (if ever) optimal. The notion of collaborative intelligence has been extended to model training as well. In this case, data flows both ways: from the cloud to the mobile during back-propagation in training, and from the mobile to the cloud during forward passes in training, as well as inference.

[0174] Lossy compression of deep feature data has been studied based on HEVC intra coding, in the context of a recent deep model for object detection. It was noted the degradation of detection performance with increased compression levels and proposed compression-augmented training to minimize this loss by producing a model that is more robust to quantization noise in feature values. However, this is still a sub-optimal solution, because the codec employed is highly complex and optimized for natural scene compression rather than deep feature compression.

[0175] The problem of deep feature compression for the collaborative intelligence has been addressed by an approach for object detection task using popular YOLOv2 network for the study of compression efficiency and recognition accuracy trade-off. Here the term deep feature has the same meaning as feature map. The word ‘deep’ comes from the collaborative intelligence idea when the output feature map of some hidden (deep) layer is captured and transferred to the cloud to perform inference. That appears to be more efficient rather than sending compressed natural image data to the cloud and perform the object detection using reconstructed images.

[0176] The efficient compression of feature maps benefits the image and video compression and reconstruction both for human perception and for machine vision. Said about disadvantages of state-of-the art autoencoder based approach to compression are also valid for machine vision tasks.

[0177] Functional modules

[0178] Variable bitrate module

[0179] An encoder can output bitstreams at different bit rates. Therefore, in some methods, an output of an encoding network is scaled (for example, each channel is multiplied by a corresponding scaling factor that is also referred to as a target gain value), and an input of a decoding network is inversely scaled (for example, each channel is multiplied by a corresponding scaling factor reciprocal that is also referred to as a target inverse gain value), as shown in FIG. 7. The scaling factor may be preset. Different quality levels or quantization parameters correspond to different target gain values. If the output of the encoding network is scaled to a smaller value, a bitstream size may be decreased. Otherwise, the bitstream size may be increased.

[0180] Color format transform

[0181] RGB and YUV are common color spaces. Conversion between RGB and YUV may be performed according to an equation specified in standards such as CCIR 601 and BT.709.

[0182] Separate structure for luma and chroma

[0183] Some VAE-based codecs use the YUV color space as an input of an encoder and an output of a decoder, as shown in FIG. 8. A Y component indicates luma, and a UV component indicates chroma. Resolution of the UV component may be the same as or lower than that of the Y component. Typical formats include YUV4:4:4, YUV4:2:2, and YUV4:2:0. The Y component is converted into a feature map F_Y through a network, and an entropy encoding module generates a bitstream of the Y component based on the feature map F_Y. The UV component is converted into a feature map F_UV through another network, and the entropy encoding module generates a bitstream of the UV component based on the feature map F_UV. Under this structure, the feature map of the Y component and the feature map of the UV component may be independently quantized, so that bits are flexibly allocated for luma and chroma. For example, for a color-sensitive image, a feature map of a UV component may be less quantized, and a quantity of bitstream bits for a UV component may be increased, to improve reconstruction quality of the UV component and achieve better visual effect.

[0184] In some other methods, an encoder concatenates (concatenate) a Y component and a UV component and then sends to a UV component processing module (for converting image information into a feature map). In addition, a decoder concatenates a reconstructed feature map of the Y component and a reconstructed feature map of the UV component and then sends to a UV component processing module 2 (for converting a feature map into image information). In this method, a correlation between the Y component and the UV component may be used to reduce a bitstream of the UV component.

[0185] In the present specification, a ‘parameter’ is a value used in an operation process of each layer forming a neural network, and for example, may include a weight used when an input value is applied to a certain operation expression. Here, the parameter may be expressed in a matrix form. The parameter is a value set as a result of training, and may be updated through separate training data when necessary.

[0186] Fig. 9 depicts one particular example of an encoder according to presently-disclosed embodiments. In particular, this figure depicts the encoding process for a single component (e.g., either a Y or UV component of an image, representing luma and chroma respectively). The per component encoder is a multi-step process. The first step is analysis transform. Analysis transform receives two inputs x[Cin, Hin, Wln] - the signal to be encoded and x[Ce,Hln, Wln] - auxiliary information to help encoding. Analysis transform outputs latent tensor y[C, h4,w4], Sizes oftensors are determined depending on input picture height H and width W and scaling factors for primary (sy) and secondary (suv) components. For primary component the parameter Ce=0, this means that primary component’s Analysis transform receives no auxiliary information (encoded independently). Forthe secondary componentthe parameter Ce=l . For secondary component’s Analysis transform the auxiliary information is tensor xY, which is re-sampled by factor suvprimary component tensor xY.

[0187] The output of analysis transform goes to the Hyper Encoder. It generates hyper-parameters tensor z[C, h6,w6], which are rounded to z , re-shaped to ID array and encoded by loss-less coder (me — tANS encoder) forming stream z. The probability distribution for loss-less coding of z is assumed to be defined by the pre-trained parameters (part of the trained model), Commulative Distribution Function (denoted as CDF(z)) computed based on those pre-trained parameters is used in loss-less entropy encoder (me-tANS encoder).

[0188] Then several steps identical to decoder operations (shown in Fig. 10) are performed on encoder side to produce entropy parameters for r encoding. 3D tensor r is converted to set of symbols } to be encoded inside Encoder SKIP module, values of r may be skipped for encoding.

[0189] Fig. 10 depicts one particular example of a decoder according to presently-disclosed embodiments. In particular this figure depicts the decoding process for a single component (e.g., either a Y or UV component of an image, representing luma and chroma respectively). The codestream is composed from bit-streams stream z and stream y for primary and secondary components. For primary and secondary colour components code streams can be parsed independently and reconstructed using modules consisting of same sequence of same neural-network layers, with the only difference in sizes on input tensors and number of tensor channels. It is this separable decoding which is depicted in Fig. 10.

[0190] First, stream z shall be parsed by loss-less entropy decoder (me — tANS decoder). The probability distribution for loss-less coding of z is assumed to be defined by the pre-trained parameters (part of the trained model), Commulative Distribution Function (denoted on a Fig. 10 as CDF(z)) computed based on those pre-trained parameters is used in loss-less entropy decoder.

[0191] Decoded hyper-prior tensors z is used as an input for two different processes: Hyper Decoder (section 11.2) and Hyper Scale Decoder (section 10.3).

[0192] Then stream y shall be parsed by loss-less decoder (me — tANS decoder). The probability distribution for parsing r is assumed to be Gaussian with zero mean value and standard deviation given as an output if following steps: Hyper Scale Decoder outputs tensors of standard deviation in log-domain la[C, h4,w4], then it is scaled according to the rate control parameter l inside Sigma Scale to produce as and then masked and scaled according to RVS parameters section inside Adaptive Sigma Scale producing I"a. Finally, tensor I"avalues are quantized (converted to the index of probability distribution table) in. Some elements of residual tensor may be skipped (not encoded / decoded) and replaced by zeros in Decoder SKIP module, which receives parsed set of syntax elements } from tANS Decoder (section 9.5), mask_sigma from SKIP Mask generation module and outputs re- shaped to 3D shape reconstructed residual tensor r [C, h4,w4].

[0193] At the decoder side, the residual r is scaled by Inverse Gain Unit according to the parameter , producing r'. Then residual tensor is scaled in invRVS (Inverse Residual and Variance Scale) module (sectionl3.2.3) forming residual tensor r". This is used for reconstructed latent tensor y. Hyper decoder generates explicit_prediction input to Multi-stage Context Model - MCM, which which is eight stages neural network process, which also takes reconstructed residual r" as an input and outputs latent space tensors y'. After Latent Scaling Before Synthesis-LSBS reconstructed latent space tensor y is ready for signal reconstruction. Latent tensors reconstructions for primary and secondary components are independent from each other.

[0194] In the context of Figs. 9 and 10, it is noted that the rate control parameter, controls compression ratio and defines operations in each of the Gain Unit, Sigma Scale, and the Inverse Gain Units shown in these figures. The forward gain tensor m is used at encoder side in Gain Unit. Forward gain tensor in logarithmic scale miogis used in Sigma scale. The inverse gain tensor 24 is used at decoder side in Inverse Gain Unit. All three forward m, inverse 24 and logarithmic domain mioggain tensors have size[C, 114, W4 ] equal to the size of residual tensor. With the sizes of the tiles being determined as discussed herein, the generating of the bitstream further comprises including into a bitstream an indication of size of the tiles in the first plurality of tiles and / or an indication of size of the tiles of the second plurality of tiles.

[0195] The encoding processing of the input tensor has its decoding counterpart, sharing functional correspondence in the processing. In this exemplary and non-limiting embodiment, a method is provided for decoding a tensor representing picture data. The tensor can have a matrix form with width=w, height=h in two spatial dimensions, and a third dimension (e.g. depth or number of channels) whose size is equal to D. The method comprises processing an input tensor representing the picture data by a neural network that includes at least a first subnetwork and a second subnetwork. It is noted that the width and height of the input tensor of the decoder may be different from the width and height of the input tensor processed by the encoder. Moreover, as is clear for those skilled in the art, the first subnetwork and second subnetwork of the decoder may perform entirely or partly (e.g. functionally) inverse functions as the first subnetwork and second subnetwork of the encoder. However, inverse function may not be interpreted in a strict mathematical manner. Rather the term “inverse” refers to processing for the purpose of decoding a tensor, so as to reconstruct original picture data. It is understood by those skilled in the art that encoding compression and decoding decompression may include further processing which may not be needed for decoding and / or encoding. For example, the RDOQ shown in Fig. 11 is an encoder-only processing. It is noted further that the terms “first” and “second” subnetwork are mere labels to distinguish between the subnetworks of the decoder (and for that purpose also of the encoder discussed above).

[0196] Examples of the first and / or second subnetwork for the encoding branch are illustrated in Fig. 11, including decoder 1604 and post filter 1611. In the method, the processing comprises: applying the first subnetwork to a first tensor including dividing the first tensor in spatial dimensions into a first plurality of tiles and processing the first plurality of tiles by the first subnetwork; after applying the first subnetwork, applying the second subnetwork to a second tensor including dividing the second tensor in the spatial dimensions into a second plurality of tiles and processing the second plurality of tiles by the second subnetwork. In the example of Fig. 11, the first subnetwork is the decoder 1604, whose output is provided as input to the second subnetwork, which is the post filter 1611. Here, the first tensor is the quantized feature tensor y in latent space, which is decoded beforehand from a bitstream Bitstream 1 by arithmetic decoder 1606. In turn, the second tensor x' input to the second subnetwork 1611 for post filtering is a feature tensor, such as a feature a feature map or a feature map in latent space. Hence, similar to the processing of the encoder, the type of input tensor (e.g. first and second tensor) may depend on the processing performed by the preceding subnetwork. In the example of Fig. 11, said preceding subnetwork is the encoder 1604. The above term “after” does not limit above processing of the decoding to immediately after in terms of an output of the first subnetwork being directly input to the second subnetwork. Rather, “after” means that the first and second plurality of tiles are processed, for example, within a same pipe in a certain temporal order, which may not be immediate in time.

[0197] Moreover, the type of input may also depend on the layer (e.g. a layer of a neural network NN or a layer of a non- trained network) at which processed input data may be branched to be used as input to another (e.g. subsequent) subnetwork. Such a layer may, for example, be an output of a hidden layer. Similar to the encoding processing, the first and second tensor are divided into so-called tiles in the decoding processing, which have been defined already above.

[0198] In Fig . 11 , the decoder 1604 may be the first subnetwork with processing 1 to processing N (processing pipelines 1 to N), As Fig. 11 depicts, the decoder 1604 takes the input feature tensor y is divided into N tiles to yN(first plurality of tiles). The respective tensor tiles are then processed by the respective blocks, i.e. processing 1 to processing N which do not have to interact with each other (e.g. wait for each other during the processing). The result of each processing provides N tiles x to xN’ (tensors). After the processing of the first plurality of tiles, the decoder subnetwork 1604 may merge the processed tiles into a first output tensor x'. In Fig. 11, said first output tensor may be the second tensor used by post-filter 1611 as input. The post-filter 1611 may be the second subnetwork with processing 1 to processing N (processing pipelines 1 to N), Similar to decoder 1604, the post-filter 1611 divides input tensor x' into a second plurality of tiles x to xN', which are processed by the respective processing 1 to processing N of post filter 1611. Processing 1 to processing N provides as output a respective tile x, to xN, which may be merged into tile x. In the example of Fig. 11, the merged x refers to the decoded tensor representing the reconstructed picture data. The merging may (but does not have to) involve cropping as illustrated in Figs. 12 to 16. It is noted that the merge / combination into tensor x' or x does not need to be performed. It is conceivable that a second subnetwork reuses the tiling of the first subnetwork and merely modifies it (refines the tiling by further division of tiles or coarsens the tiling by joining a plurality of tiles into one). In the example of Fig. 11 , the processing 1 to N may perform the processing on a tile basis, i.e. processing i processes a tile i. Alternatively, processing i may process a component i among multiple components of an input tensor. In this case, processing i divides component i into a plurality of tiles and processes the tiles separately or in parallel.

[0199] In some exemplary implementation the first subnetwork can be Hyper Decoder or the combination of Hyper Decoded and a context modelling network (e.g. MCM, Multistage Context Modelling) and the second subnetwork can be combination of a context modelling network (e.g. MCM, Multistage Context Modelling) and a synthesis transform or just a synthesis transform or just the context modelling network.

[0200] Further, at least two respective collocated tiles of the first plurality of tiles and the second plurality of tiles differ in size. In other words, the subdivision into tiles can differ for each subnetwork, and hence each subnetwork can use different tile sizes. However, the input of each subnetwork (i.e. the first and second tensor) is split into a grid of tiles of the same size, except for tiles at bottom and right image boundary, which can have a smaller size.

[0201] Otherwise, properties and / or characteristics of the first and second plurality of tiles used in the decoding processing are similar to the one of the encoding processing discussed before. In particular, in the above exemplary implementation, tiles of the first plurality of tiles that are adjacent in at least one dimension of the spatial dimensions partly overlap; and / or tiles of the second plurality of tiles that are adjacent in at least one dimension of the spatial dimensions partly overlap. Examples for adjacent tiles along with partly overlap are shown in Figs. 12, 13, 14, 15, and 16. Regions of the image where no tile overlaps another are uniquely represented by a single tile.

[0202] Further, in some exemplary implementations, tiles of the first plurality of tiles are processed independently by the first subnetwork; and / or tiles of the second plurality of tiles are processed independently by the second subnetwork. In other words, the processing of the tiles is independent from each other and thus, parallelizable. In an example, at least two tiles of the first plurality of tiles are processed in parallel by the first subnetwork; and / or at least two tiles of the second plurality of tiles are processed in parallel by the second subnetwork. Fig. 11 illustrates the parallel processing on the decoder side, where components (e.g. tiles and / or spatial components) to yNof the first input tensor are processed by the decoder 1604 without interaction among the processing for components 1 to N. The result of each processing is output tensor components x to xNThey may be further combined to an output tensor x’. The second subnetwork may be post-filter 1611 in Fig. 11, which takes as input the tensor x' from decoder 1604 (first subnetwork). In this example, the input of post-filer 1611 is directly connected to the output of decoder 1604. Alternatively, there may be further processing between the post-filter and decoder in Fig. 11. The input tensor x' is divided into N tiles x, 'to xN' by the post-filter, which then processes the respective tiles independently or in parallel. The parallel processing is reflected in Fig. 11 in that the processing of component 1 to N may not interact among each other. The result of the parallel processing of the post filter are reconstructed tiles x, to xN, which may be merged into one tensor x representing reconstructed picture data. In addition, said dividing of the first tensor includes determining sizes of tiles in the first plurality of tiles based on a first predefined condition; and / or said dividing of the second tensor includes determining sizes of tiles in the second plurality of tiles based on a second predefined condition. F or example, the first predefined condition and / or the second predefined condition is based on available decoder hardware resources and / or motion present in the picture data or other features as already described above with reference to the encoding.

[0203] The first and second subnetwork for the decoding processing may have a similar configuration as the subnetwork(s) for the encoding processing. Specifically, the first subnetwork performs processing by one or more layers including at least one convolutional layer or at least one pooling layer; and / or the second subnetwork performs processing by one or more layers including at least one convolutional layer or at least one pooling layer. Fig. 6A and 6B show a decoder within the NN-based VAE framework, with respective convolutional layers, where the respective size of the input tensor (feature map) y is subsequently enlarged by 2 (upsampling). Moreover, the first subnetwork and the second subnetwork perform respective processing that is a part of picture or moving picture decompression. Such a processing may be provided by the VAE encoder shown in Fig. 6A and 6B, taking as input feature tensor y so as to decompress and reconstruct an output image as output of conv layer, representing the reconstructed picture data (i.e. decoded tensor). For example, the first subnetwork and / or the second subnetwork perform one of: picture decoding by a convolutional subnetwork; and picture filtering.

[0204] In an implementation, the input tensor is a picture or a sequence of pictures including one or more components of which at least one is a color component. Alternatively, the input tensor may be a latent space representation of a picture, which may be an output (e.g. output tensor) of a pre-processing. The one or more components are color components and / or depth and / or motion map and / or other feature map(s) associated with picture samples. The input tensor has at least two components, namely a first component and a second component; and the first subnetwork divides the first component into a plurality of tiles and divides the second component into a further plurality of tiles, wherein at least two respective collocated tiles of the plurality of tiles and the further plurality of tiles differ in size; and / or the second subnetwork divides the first component into a plurality of tiles and divides the second component into a further plurality of tiles, wherein at least two respective collocated tiles of the plurality of tiles and the further plurality of tiles differ in size.

[0205] Fig. 12 shows an example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples in the overlap region, and the concatenation of the cropped regions. Each Li comprises the total receptive field.

[0206] Fig. 13 shows another example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles) similar to Fig. 12, except that cropping is dismissed.

[0207] Fig. 14 shows an example of dividing a first and / or second tensor into overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples in the overlap region and the concatenation of the cropped regions. Each Li comprises a subset of the total receptive field.

[0208] Fig. 15 shows an example of dividing a first and / or second tensor into non-overlapping regions Li (i.e. first and second tiles), the subsequent cropping of samples, and the concatenation of the cropped regions. Each Li comprises a subset of the total receptive field.

[0209] Fig. 16 illustrates the various parameters, such as the sizes of regions Li, Ri, and overlap regions etc. of first and / or second tiles, that may be included into (and parsed from) the bitstream. Any of the various parameters may be included as an indication into (and parsed from) the bitstream. Although Fig. 16 shows an image split only into quarters, it should be understood that an image may be split into many tiles. As such, the dimensions may be mapped as W1 + x is the tile_size signaled when using more than four tiles and the signalled overlap may be x + x. Half of the overlap (one x) is then discarded on each side of a tile in order to remove or avoid artefacts.

[0210] Fig 17 shows a schematic diagram of part of the JPEG Al coding scheme. The whole scheme comprises analysis on the encoder side and synthesis on the decoder side. In the JPEG standard there exist predictions and residuals. There is a prediction module on both of the encoder and decoder sides. On the decoder side there is also a synthesis module. In the JPEG standard the method of tiling for processing is only found in the synthesis module of the decoder side. This process treats each tile separately, where after synthesis the tiles are combined to recreate the final image.

[0211] Currently, tiling is only used in the synthesis module. In the synthesis module the y tensor is split into a few subtensors forming tiles and then the synthesis network is applied tile by tile. This is done primarily to save memory. The synthesis network is quite complex and quite a lot of memory is needed for this inference. However, on some devices there is not a lot of memory available. Therefore, tensor y may be split into a few smaller tensors called tiles. Inference may be then performed tile by tile. After synthesis the results of the inference of each subtensor may be combined to form the output (or the input for post-filters).

[0212] Currently in the typical synthesis occurs in only one module, which is the synthesis module. There is typically no tiling used in other modules. Therefore, there is only a need to signal one size of tile and only one tile overlap.

[0213] Specifically, in the present context of still image coding, the tiles are implemented in latent space y. Latent space may also be referred to as feature space. In JPEG Al coding there exists two feature spaces, latent space yAand hyper latent space zA. The input to the network can be split into tiles and combined using an overlap. The overlap is used because inference in a neural network close to boundaries of the tiles usually creates artefacts. Each tile goes through inference. By removing the overlap amount the artefacts can be removed.

[0214] Two tiles are recombined with an overlapping area. When the image is reconstructed the tiles are cropped such that the areas between the midpoint of the overlap and the edge of each tile are removed. That is, the cropping is carried out from the middle of the overlapping area up to the nearest tile edge. As a result, part of the image may be reconstructed without the artefacts on the boundary caused by inference. The action of overlapping and cropping the tiles also has a smoothing effect. Therefore, if the overlap is big enough then there are no artefacts at the boundaries between tiles. The boarders between the tiles are therefore not visible when the overlap is big enough.

[0215] The overlap needed may be different for different images. For some images, a big overlapping area is not needed. For some images, even some small overlap may be enough to remove all boundary artefacts. It is preferable to use the smallest overlap possible while removing as many boundary artefacts as possible. This is because the overlap causes additional complexity to the process due to processing more samples (processing some samples of the input tensor twice). However, for some images a larger overlap is needed and therefore the additional complexity is unavoidable. Thus, it has been recognised that it would be beneficial not to fix the size of the overlap, but have it signalled instead.

[0216] The required size of the overlapping area may be selected based on a metric (e.g. PSNR, MS-SSIM, VMAF, etc.) evaluation for an overlap which can be measured against a predetermined threshold. If the threshold is exceeded, then the overlap may be increased in size or vice versa.

[0217] Fig 17 also shows two variations of the tile scheme which may be used in the image coding process. In this example and others described herein, the tiles are used in the decoder side of the image processing. However, the tile arrangement may also be used on the encoder side and in other modules than those discussed herein. The two types of tile schemes shown are collocated tiles and hierarchical tiles. It can be seen that in these scheme types tiles in different modules are needed. Specifically, in this example, tiles are used in the latent prediction module and the synthesis module. The first example variant scheme is a Collocated tile and starts with a tensor zAof a particular size. Then a latent prediction is made using zAand resulting in yAof a particular size. Through a following step of synthesis on yAthe image data is recombined. That is, each zAtile is processed through the synthesis module one at a time. Latent space yAinformation is not used for the next zAtile.

[0218] The second example variant scheme is a Hierarchical tile scheme. In this scheme zAconsists of larger tiles enabling one prediction to be made at the prediction module for the entire tile and thus a bigger region. The larger region is then split into smaller yAregions and synthesis is performed. YAtiles are smaller because the synthesis network is much more complex and inference requires much more memory, requiring smaller tiles. However, the latent prediction is a comparatively shallow metric so the memory required is not as large and bigger tiles can be handled by the device during processing.

[0219] Fig. 18 shows the two variants of tile schemes in more detail. In this example, the two modules across which the tiles are processed are the prediction module and the synthesis module. Here you can see that tiles have corresponding sizes. However, zAhas a smaller resolution. For example, an area of 16 by 16 in zAcorresponds to an area of 64 by 64 in synthesis (yA). There are no hierarchical tiles, for example where one big tile in yAis split into multiple smaller tiles. This is why they are said to be collocated tiles, because the sizes of the tiles correspond to each other. Even though the different modules have different resolutions, the size of the tile in zAcan be derived from the size of the tile in yAjust with division. The sizes of tiles in yAare multiples of this constant representing the relationship of the resolution differences between modules, not the number of tiles.

[0220] The second tile scheme shows hierarchical tiles. In this example, the prediction module uses bigger tiles. Each of the bigger tiles may then be split into smaller tiles for use in the next module, for example, the synthesis module. The advantage of using hierarchical tiles is that bigger tiles are able to be used by subnetworks with lower memory requirements. That is, some modules running a less complex process can handle applying that process to larger amounts of data in one go. For example, the prediction module is able to handle more data and therefore can use bigger tiles. For other modules applying different subnetworks, for example the synthesis module, more memory is needed to process the same area. Accordingly, tiles for this operation will need to be smaller. Where collocated tiles are used, the tiles for prediction are required to be the same size as those used for synthesis. The synthesis module provides the limiting factor on tile size. However, this means that the prediction module is required to process more tiles. This introduces more complexity for the prediction module and also requires more processing for the increased number of overlapping areas. If the resolution in zAis four times more in each dimension, and the tile overlap is required to be at least, for example, three or four samples for normal reconstruction, then using collocated tiles with a four sample overlap in zAleads to an overlap of 16 samples in the synthesis domain. This is more overlap than is needed and leads to a significant additional complexity in the synthesis module.

[0221] Therefore, by specifying a tile overlap for one module separate to the overlap for a second module, the same size overlap in samples can be achieved for both modules even if they have different resolutions and tiles sizes.

[0222] The only drawback for using hierarchical tiles is the introduction of an additional latency. This latency occurs when the second module (e.g. the synthesis module), has to wait for the whole of the larger tile being processed by the first module (e.g. the prediction module), to be output by that first module. This small amount of time delay is the trade-off for decreasing the processing memory required to implement tiles in two consecutive modules.

[0223] In existing tile schemes only one module in the neural network uses tiles. This is typically the synthesis module. Therefore, currently it is not necessary to specify whether the tile size or tile overlap indicated is for any particular module. However, it would be possible to use tiles in multiple different modules or subnetworks of the neural network and have different tile sizes used in each module. For zAin hyper latent space, the input of the prediction module, the resolution is four times smaller than for yAin latent space, which is the input of the synthesis module.

[0224] Tiling may happen in the synthesis transform, also in prediction. In prediction there are two modules, hyperencoder / decoder, MCM, where tiling may also be used. Generally, when defining tile size for operations comprising tiling, all but the last tile have the same size. Usually the last tile is smaller than the preceding tiles. This is because there is usually no way to guarantee that the image data provided with be of a size which can be divided into an integer number of tiles. That is, the size of the image is not always an integer multiple of the size of the tile including overlap.

[0225] When neural networks are used for image / video compression, the size of the content is not known in advance, it varies from bitstream to bitstream. Reinitializing a neural network (changing the input tensor size) for every new image / video brings a significant inference speed penalty.

[0226] There is a desire to run the JPEG Al codec for images on smaller devices such as smartphones and other portable devices. Smartphones are equipped with what is called a Neural Processing Unit (NPU). The NPU is needed for inference of a neural network. The NPU forms a specific part of the smartphone and is embodied on a specific dedicated part of the processor chip. The NPU enables a much faster deployment of the Neural Network than is possible by using just the standard CPU. The NPU can be considered as separate hardware. However, the hardware thus has some restriction. One of these restrictions is the size of an input image tensor it is able to process. Neural network inference implemented on NPU works most efficiently when the sizes of the input tensors are fixed. In this case the neural network can be optimized in advance for the particular input size and considering specifics of the target hardware, and so can be inferenced faster on the end-user device (e.g. on a smartphone).

[0227] Therefore, a convenient situation for implementing the NPU is that the size of the tensor is known in advance. In this case, in practise there exists specific software to optimise the model for a particular tensor size for a particular device. After optimisation the run time is much smaller than it is without the optimisation. Although this is good from the runtime point of view, each optimisation requires the input tensor size for the neural network to be redefined. This is fine for some applications such as object detection or facial recognition because the interest is only in a part of the image. Thus, it is possible to crop around the area of interest to re-size the input tensor and then provide this reduced version of the image with a tensor size corresponding to optimisation to the neural network.

[0228] However, for other applications where the resolution of the incoming image is unknown and the entire image is needed, cropping the image is not possible. Additionally, it is not possible to prepare a neural network for a portable device to handle all the possible types of image resolution and size that exist. For example, one bit stream may comprise a 4K image and another bit stream may comprise an HD image. It may be possible to have multiple models which are optimised to handle a different size each, such as two or three, on a single device. However, it is a limitation of the NPU that it cannot handle the number of multiple different models that would be required such that each was optimised for a different possible image type.

[0229] Therefore, in the presently described scenario, it is not possible to focus the neural network on one specific tensor size. This presents a problem that needs to be solved. Simply resizing the incoming picture will significantly impact the speed of image processing and also quality of the image, which is especially critical for the applications where, so-called, visually lossless quality / compression is required. Not resizing the incoming picture will limit the possible application of the neural network to tensors only of a specific size and thus received images of only the corresponding specific size. This may be adequate for applying to some messenger processes, but not for applying generally. The first step of processing an image is to split the input tensor into tiles of a predefined size. The predefined tile size is known in advance and the neural network model is prepared and optimised for this particular size on this particular device. Once the model is prepared the input tensor is simply split into the tiles of the predefined size and then inferenced tile by tile. The process of inference at the neural network is comparatively fast when the tensor is processed tile by tile. The only problem is the last tile. The last tile often is of a different size to the other tiles in order to account for the last part of the tensor which often does not fit completely within a tile to fill it. Processing a differently sized tile to the optimised tile size causes a problem for the hardware similar to that described above - it cannot be inferenced well by the neural network.

[0230] In a first embodiment, there is proposed a method of generating a set of tiles extending fully across one dimension of the image in which one tile uniquely represents an extend of the image in the first dimension compared to that represented by any of the other tiles in the set by representing an edge region of the image and padding data. Each of the tiles of the set overlapping at least one neighbouring tile and all the tiles of the set are of the same extent in the first dimension. That is, the tile size corresponds to the predefined tile size of the optimised neural network. The first embodiment will now be described in terms of the last tile in a row. Where this last tile is prepared for processing by padding the contents of the tile. The padding is added to the part of the tile which extends beyond the edge of the image. The tile may be padded with samples of a constant value. Ideally the constant value is chosen to be zero. This is because the bits needed to signal the value of the samples used to create the padding is smallest when signalling zero. In case of applying the method on the decoder side, using a constant value for the padding can help to reduce the influence of the padded area on neighbouring samples, especially when the entirety of the padding is with the value zero. Such influence can occur when the convolutional layers are applied. However even with the constant value equal to zero this impact still can be observed without retraining of the neural network.

[0231] As described above, the first embodiment comprises padding a tile, e.g. the last tile in a row, to achieve the predefined tile size for all tiles and to include the edge of the image. Fig. 19 illustrates the first embodiment comprising padding a tile in a first dimension of the image tensor rather than providing a smaller tile. Fig. 19 shows a first tile 1910 and a second tile 1920 of the predefined size. The size of the image remaining to be tiled is smaller than double the width of the predefined tile size minus a predefined overlap size. Therefore, a padded area 1930 is added where the second tile 1920 of predefined size does not comprise part of the image.

[0232] The first embodiment comprising padding a tile is simple to implement. However, even when the padding comprises zero values, after a few convolution layers of the neural network these zeros are no longer zeros. This is because in convolutional networks some convolutions introduce biases. These bias values are constants which are added after the convolution layer. But because of these added constants, even though the introduced padding may comprise only zeros, after the first convolution layer with a bias there will no longer be zeros in the padded area but a different constant instead. The introduced constants are then present for the next layer of the convolutional network. The input to the following convolution then comprises non-zero values in the boundary samples. Such values can create artefacts at the edge of the image which result in a haze or shadow that is visible when looking at the final image. It is not possible to add the padding after each layer due to the size of the padding required for the image not being known in advance. This is because the model being used is always doing the same process for each input image, and when the padding area changes for each individual image there is no way to set such a layer by layer zero padding process within the model.

[0233] There is proposed herein a second embodiment which can be implemented such that no artefacts are introduced, unlike for the first embodiment. The second embodiment comprises increasing an overlap between two tiles in a first dimension of the image such that the second tile of the pair is still filled by the image but is also of the predefined tile size. Fig. 20 illustrates the second embodiment comprising shifting a tile to change the amount of overlap between a pair of tiles in a first dimension of the image. Fig. 20 shows a first tile 2010 and a second tile 2020 of the predefined size. Again, the size of the image remaining to be tiled is smaller than double the width of the predefined tile size minus a predefined overlap size. Therefore, in this example, there is a bigger overlapping area 2030 between the last two tiles of a row such that the edge of the image coincides with the edge of the last tile 2020 in the row. The increased overlap means that the number of samples which are processed twice is increased, but the amount of processing time based on the number of samples to be processed is not increased in comparison with the first embodiment.

[0234] Further, if, for example, the size of the overlapping area is 500 samples wide and there are only 5 tiles, then the overlapping area may be divided widthways by 5 and the overlap may be set to 100 samples for every tile-to-tile boundary. This may be easily implemented if the size of the overlapping area is big but the number of tiles is comparatively small. Such a distribution of the additional overlapping area may also be beneficial for reducing artefacts at tile boundaries by increasing the overlap by more than the minimum amount required to avoid such artefacts.

[0235] Figs. 19 and 20 depict the scenario for the horizonal or x dimension of the image where the width of the image is not the width covered by an integer number of tiles. The same method can be applied in the vertical or y dimension of the image as shown in the figures. It is noted that the same restrictions are not placed on operations which use tiling outside of neural networks. This is because processing an image of any size comprising a number of samples is a straightforward adaptation for such operations. Whereas for a neural network which requires training within certain parameters in order to be optimised, changing these parameters such as tile size for a single inference of a trained network on one tile will not lead to an optimum output for the data in that tile.

[0236] The two proposed approaches are approximately the same with regard to processing time. This is because the time it takes to process the zero samples in the padding embodiment is similar to the time it takes to process the samples in the overlapping area twice.

[0237] Specifically, there are proposed herein two different embodiments for providing a last tile of the same size as the other tiles. However, it should be understood that the two embodiments may be applied exclusively of the other or in a combined manner to the same image. Further, either or both embodiments may be applied to a different single tile or pair of tiles in a row or column of the image tensor tiles. For example, the first tile instead of the last tile may be padded or shifted in each row or column. Alternatively, a central tile or one of two central tiles may be shifted accordingly such that the larger overlap exists in the middle of the image rather than at the leading or end edge of each row or column. Similarly, although examples will be discussed herein in reference to the tiles of one or more rows, the proposed embodiments may be applied to the tiles of one or more columns. Thus, the proposed embodiments may also be applied to both the multiple rows and columns of the tiles of the input tensor as needed. It should therefore be understood that the described adjustment to be made to a tile of the input tensor may be applied to any orthogonal extent of the tensor as it aligns to the edge of the image as required.

[0238] Fig. 21 shows two pairs of side-by-side image examples, one of each pair is generated by the first embodiment and the second embodiment. The first pair comprises images A and B. The second pair comprises images C and D. The left-hand images, A and C, are generated with tiling according to the first embodiment. The right-hand images, C and D, are generated with tiling according to the second embodiment.

[0239] The first pair of images, A and B, illustrate a portion of an image processed using tiles. Some of which are located at the bottom of a plurality of columns within which a horizonal image edge exists. The image on the left, A, is processed using the first embodiment, where the image in the last tiles or the columns has been padded before inference by the neural network. The image on the right, B, has been processed using the second embodiment, where the image in the last tiles of the columns has been shifted such that it fills the tile with a larger overlap. It can be seen that there are artefacts present along the bottom of the picture in image A. The artefacts appear as a slight lightening or cloudy type of effect on the image, reducing the clarity in this region. These artefacts are not visible in the same picture in image B, which has no such visible artefacts and as such shows a clearly defined line between one side of the building and the other side of the building in the picture.

[0240] The second pair of images, C and D, also illustrate a portion of an image processed using tiles. Some of these tiles are located at the right of a plurality of rows within which a vertical image edge exists. The image on the left, C, is processed using the first embodiment, where the image in the last tiles of the rows has been padded before inference by the neural network. The image on the right, D, has been processed using the second embodiment, where the image in the last tiles of the rows has been shifted such that it fills the tile with a larger overlap. It can be seen that there are artefacts present along the right-hand edge of the picture in image C. The artefacts appear as a slight darkening of the image, reducing the clarity in this region and darkening the brightest area compared to the same feature elsewhere in the image. These artefacts are not visible in the same picture in image D, which has no such visible artefacts and as such shows more clearly defined edges between different features in the picture.

[0241] The decision of whether to use the first embodiment and padding versus the second embodiment and shifting depends on the number of samples. In principle, it is also possible to minimise the artefacts seen in Fig. 21 by increasing the overlap and as such reducing the padding required. Therefore, the boundary artefacts would not be so strong and padding could still be used. However, as can be seen from Fig. 21, for some sizes of padding, boundary artefacts are introduced. However, if a single solution is preferred, for example at an encoder side, then shifting may be chosen to avoid such significant artefacts.

[0242] There are advantages associated with each embodiment over the other. For example, padding a tile at the edge of an image according to the first embodiment is straightforward to implement and signal. The remaining area within the tile which is not filled by the image may be padded by zeros. The reduced amount of image data in the last tile also means that the last tile is relatively fast to process by the neural network compared to other tiles of the image. Another benefit is that padded tiles can be used with small images, i.e. with around a 1024 by 1024 resolution. That is, padding can be used for images with low resolution where not many tiles are needed. For example, if the tile size is 1000 by 1000 samples but the image resolution is 1256 by 1256, then it is not possible to efficiently shift two tiles according to the second embodiment so as to cover this image. Using four tiles of 1000 by 1000 samples in this way would result in more processing than is practical for this image resolution. However, using padding to process such an image in four tiles is simple.

[0243] A combined strategy may also be proposed, where for all images significantly larger than a single tile size, for example at least twice the size of a single tile, embodiment two and shifting may be used. Whereas for smaller images, for example which are no more than 50% larger than a single tile in any one dimension, then the first embodiment and padding may be used. Additionally or alternatively, for images where the tile size is close to the size of padding needed in the final tile, a combination of both embodiments may be used, as described above.

[0244] The second embodiment has the advantages of not requiring retraining of the neural network to implement. This is because the tiles remain the same size and the processing used on data being handled is the same as for any other tile of the image. There are no visible artefacts introduced at the boundary because zero value padding is not being included over multiple convolution layers. And finally, as the previous tile is normally still in a memory cache, the runtime may be smaller than for implementing padding according to the first embodiment in the same scenario. This is because the additional data which requires processing is always data which has already been processed in the previous tile. It may be possible to train the neural network to identify areas of padding and reduce artefacts incurred as a result of processing such images.

[0245] Embodiment 1 can be applied for the input image on the encoder side, e.g. the image can be padded in order to be covered by the tiles of the same size. However, in this case padded samples will be signalled, which may cause the bitrate increase. When the method (embodiment 1 or embodiment 2) is applied only for the decoder sub-networks (e.g. for prediction sub-network and / or synthesis sub-network) additional samples are not signalled. For some application scenarios where the encoder should know the reconstructed image (or some features of the reconstructed image), e.g. video compression, the described method should be applied also on the encoder-side for the subnetworks used to obtain the reconstructed image (or some features of the reconstructed image), e.g. prediction sub-network and / or synthesis sub-network.

[0246] There is proposed herein a method wherein all the tiles of an image may be maintained at the same predefined size in order to be successfully processed by the optimised neural network while still covering the full extent of the image without pre-existing knowledge of the image size. Thus, the proposed embodiments provide a method for efficient use of an NPU for neural network inference within various application scenarios, where the input tensor size is unknown at the moment of the model conversion (i.e. during optimization for the particular device).

[0247] FIG. 22 is a flow diagram illustrating an exemplary method for encoding an image based on a neural network architecture. The bitstream includes information comprising image data having been divided up into tiles to be transmitted. The method described herein is for encoding an image to form encoded image data. Specifically, wherein the outgoing or incoming bitstream comprises images of a resolution for which the neural network has not been optimised. The embodiment according to FIG.22 may be configured to provide an output readily decoded by the decoding method described with reference to FIG.25. The method 2200 of FIG. 22 will be described as being performed by a neural network system of one or more computers located in one or more locations. For example, a system configured to perform image compression, e.g., the neural network of FIG. 1 can perform the method 2200.

[0248] The first step S2210 of the method comprises processing an image to form a series of data units in a latent space, each data unit representing a respective tile in the image, such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension. That is, tiles are formed such that in a first dimension (such as width of an image) the tiles have a regular size across the entire dimension of the image and each comprise the same proportion of the image samples in latent space except for one tile which comprises image samples not included in any other tile and for a different proportion of the image than the other tiles. In the embodiments described above, that may be implemented by containing a smaller proportion of the image in one tile with the rest of the tile padded, or by containing the same proportion of the image in one tile but with a larger overlap such that the proportion of the image covered only by that tile is smaller than for the other tiles.

[0249] The method may comprise the additional step of S2220 forming the data units such that in the first dimension of the image one tile represents an edge region of the image and padding data. The padding data may be of a constant value. The padding data may repeat data of the edge region of the image extending in a second dimension orthogonal to the first dimension. The method may comprise forming the encoded image data to include an indication of the extent of the padding data in the first dimension.

[0250] The method may comprise the additional step of S2230 forming the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. One tile of the first pair may have an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image.

[0251] The method may comprise forming the encoded image data to include an indication of the width of the overlap between the tiles of the first pair. The method may comprise forming the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set. The method may comprise forming the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. The method may comprises performing an encoding or latent prediction operation on the tiles to form the data units. The latent prediction operation may comprise one of a hyper-encoding operation and a multistage context modelling operation. The method may comprise determining whether to form the data units using padding data or increasing the overlap between two tiles in dependence on the extent of the image in the first dimension.

[0252] Fig. 23 is a flow diagram illustrating an exemplary method for decoding an image based on a neural network architecture. The bitstream includes information comprising an image tensor to be processed using tiling to reconstruct the image from received tiles. The image tensor may be used by one or more subnetworks of the neural network. The method described herein below is for decoding an image to form decoded image data. Specifically, wherein the incoming bitstream comprises image data of a resolution for which the neural network has not been optimised.

[0253] The method described is for decoding encoded image data to recover an image. The method comprises processing data units in a latent space, each data unit representing a respective tile in the image, the method being such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0254] It should be understood that the process for decoding image tensor data comprises the complimentary steps to those according to the method for encoding described herein above.

[0255] There is also proposed herein a data carrier storing data defining instructions for one or more processors to perform the steps of any of any of the above-described methods.

[0256] While operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0257] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

[0258] In this embodiment, there is provided a computer program stored on a non-transitory medium comprising code which when executed on one or more processors performs steps of any of the methods discussed above. The respective flowcharts of the processing are shown in Figs. 22 and 23. Moreover, as already mentioned, the present disclosure also provides devices (apparatus(es)) which are configured to perform the steps of the methods described above.

[0259] In this exemplary and non-limiting embodiment, a processing apparatus is provided for encoding a tensor representing picture data. Fig. 24 shows the processing apparatus 2400 having respective modules to perform the discussed steps of the encoding processing, comprising processing circuitry 2410. The processing circuitry is configured to: process input picture data by a neural network that includes at least a first subnetwork realized by NN processing module - subnetwork 2411. NN processing module - subnetwork 2411 may have a separate dividing module 2413 which divides the picture data in spatial dimensions into a first plurality of tiles and processes the first plurality of tiles by the subnetwork. Alternatively, the respective modules processing the picture data may be implemented in one single module, which may be included in a single circuitry or separate circuitry.

[0260] The processing apparatus may have further a bitstream 2415, providing a function of signalling in a bitstream an indication of size of the tiles.

[0261] In an exemplary and non-limiting embodiment, a processing apparatus for encoding an image to form encoded image data, the processing apparatus comprising: one or more processors; and a non- transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, wherein the programming, when executed by the one or more processors, configures the encoder to carry out the method according to the encoding method described above.

[0262] The image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processors) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0263] The padding data may repeat data of the edge region of the image extending in a second dimension orthogonal to the first dimension. That is, the padding data may comprise image data near the edge of the image as covered by the respective tile repeated to the edge of the tile comprising the edge. The image data near the edge may comprise the last sample of the image at the edge of each row or column of data to be extended to the edge of the tile, regardless of value. The processor(s) may be configured to form the encoded image data to include an indication of the extent of the padding data in the first dimension.

[0264] The processors) may be configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. Thus, the second embodiment as described above is implemented. One tile of the first pair may have an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image. The processor(s) may be configured to form the encoded image data to include an indication of the width of the overlap between the tiles of the first pair. The processors) is / are configured to form the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set.

[0265] The processor(s) may be configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. The processors) may be configured to perform an encoding or latent prediction operation on the tiles to form the data units. The latent prediction operation may comprise one of a hyper-encoding operation and a multistage context modelling operation.

[0266] Each tile may have equal height and width. Conveniently, all the tiles have the same height. Conveniently, all the tiles have the same width.

[0267] The processor(s) may be configured to determine whether to form the data units one tile comprises an edge region of the image and padding data or where the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adj acent to each other in dependence on the extent of the image in the first dimension.

[0268] There is proposed herein an image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processor(s) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and one tile of the set represents an edge region of the image and padding data such that all the tiles of the set are of the same extent in the first dimension.

[0269] Further, there is proposed herein an image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processor(s) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and the overlap between a first pair of two tiles of the set that are adjacent to each other is larger than the overlap between a second pair of two tiles of the set that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension.

[0270] In this exemplary and non-limiting embodiment, a processing apparatus is provided for decoding a tensor representing picture data. Fig. 25 shows the processing apparatus 2500 having respective modules to perform the discussed steps of the decoding processing, comprising processing circuitry 2510. The processing circuitry is configured to: process an input tensor by a neural network that includes at least a first subnetwork realized by NN processing module - subnetwork 2511. NN processing module - subnetwork 2511 may have a separate dividing module 2513 which divides the first tensor in spatial dimensions into a first plurality of tiles and processes the first plurality of tiles by the subnetwork. Alternatively, the respective modules processing the first input tensor and / or the first plurality of tiles may be implemented in one single module, which may be included in a single circuitry or separate circuitry.

[0271] The processing apparatus may have further a parsing module 2515, providing a function of parsing from a bitstream an indication of size of the tiles of the first and / or second plurality of tiles. Module 2515 may further provide a function of extracting from the bitstream the input tensor. In this exemplary and non-limiting embodiment, a processing apparatus for decoding a tensor representing picture data, the processing apparatus comprising: one or more processors; and a non- transitory computer-readable storage medium coupled to the one or more processors and storing programming for execution by the one or more processors, wherein the programming, when executed by the one or more processors, configures the encoder to carry out the method according to the decoding method described above.

[0272] An image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

[0273] The processors) may be configured to form the data units such that in the first dimension of the image one tile represents an edge region of the image space and padding data. That is, one tile of the set of tiles in the first direction includes the edge of the image and some padding data. The padding data may be of a constant value. For example, the padding data may comprise Os or Is or another low bit number. The padding data may repeat data of the edge region extending in a second dimension orthogonal to the first dimension. That is, the padding data may comprise one or more samples repeated from the edge of the image up to the edge of the tile.

[0274] The processor(s) may be configured to form the encoded image data to determine an extent of the padding data in the first dimension either (i) in accordance with a value signalled in the encoded image data or (ii) in dependence on an extent of the image in the first dimension signalled in the encoded image data. The processor(s) may be configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other. One tile of the first pair may have an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image. The processor(s) may be configured to determine the width of the overlap between the tiles of the first pair in accordance with a value signalled in the encoded image data.

[0275] The processors) may be configured to select as the said one of the tiles of the set a tile whose identity is signalled as such in the encoded image data. The processor(s) may be configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent. The processors) may be configured to perform a decoding or latent prediction operation on the tiles to form the data units. The latent prediction operation may comprise one of a hyper-decoding operation and a multistage context modelling operation.

[0276] The processor(s) may be configured to determine whether to form the data units where one tile comprises an edge region of the image and padding data or where the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other, in dependence on one of (i) the extent of the image in the first dimension and (ii) a value signalled in the encoded image data.

[0277] There is also proposed herein an image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension of the image one tile representing an edge region of the image space and padding data such that all the tiles of the set are of the same extent in the first dimension.

[0278] Further, there is also proposed herein an image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension.

[0279] Some exemplary implementations in hardware and software

[0280] The corresponding system which may deploy the above-mentioned encoder-decoder processing chain is illustrated in Fig. 26. Fig. 26 is a schematic block diagram illustrating an example coding system, e.g. a video, image, audio, and / or other coding system (or short coding system) that may utilize techniques of this present application. Video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of video coding system 10 represent examples of devices that may be configured to perform techniques in accordance with various examples described in the present application. For example, the video coding and decoding may employ neural network such which may be distributed and which may apply the above-mentioned bitstream parsing and / or bitstream generation to convey feature maps between the distributed computation nodes (two or more).

[0281] As shown in Fig. 26, the coding system 10 comprises a source device 12 configured to provide encoded picture data 21 e.g. to a destination device 14 for decoding the encoded picture data 13.

[0282] The source device 12 comprises an encoder 20, and may additionally, i.e. optionally, comprise a picture source 16, a preprocessor (or pre-processing unit) 18, e.g. a picture pre-processor 18, and a communication interface or communication unit 22.

[0283] The picture source 16 may comprise or be any kind of picture capturing device, for example a camera for capturing a real- world picture, and / or any kind of a picture generating device, for example a computer-graphics processor for generating a computer animated picture, or any kind of other device for obtaining and / or providing a real-world picture, a computer generated picture (e.g. a screen content, a virtual reality (VR) picture) and / or any combination thereof (e.g. an augmented reality (AR) picture). The picture source may be any kind of memory or storage storing any of the aforementioned pictures.

[0284] In distinction to the pre-processor 18 and the processing performed by the pre-processing unit 18, the picture or picture data 17 may also be referred to as raw picture or raw picture data 17.

[0285] Pre-processor 18 is configured to receive the (raw) picture data 17 and to perform pre-processing on the picture data 17 to obtain a pre-processed picture 19 or pre-processed picture data 19. Pre-processing performed by the pre-processor 18 may, e.g., comprise trimming, color format conversion (e.g. from RGB to YCbCr), color correction, or de-noising. It can be understood that the pre-processing unit 18 may be optional component. It is noted that the pre-processing may also employ a neural network (such as in any of Figs. 1 to 7) which uses the presence indicator signalling.

[0286] The video encoder 20 is configured to receive the pre-processed picture data 19 and provide encoded picture data 21. Communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and to transmit the encoded picture data 21 (or any further processed version thereof) over communication channel 13 to another device, e.g. the destination device 14 or any other device, for storage or direct reconstruction.

[0287] The destination device 14 comprises a decoder 30 (e.g. a video decoder 30), and may additionally, i.e. optionally, comprise a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32) and a display device 34.

[0288] The communication interface 28 of the destination device 14 is configured receive the encoded picture data 21 (or any further processed version thereof), e.g. directly from the source device 12 or from any other source, e.g. a storage device, e.g. an encoded picture data storage device, and provide the encoded picture data 21 to the decoder 30.

[0289] The communication interface 22 and the communication interface 28 may be configured to transmit or receive the encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g. a direct wired or wireless connection, or via any kind of network, e.g. a wired or wireless network or any combination thereof, or any kind of private and public network, or any kind of combination thereof.

[0290] The communication interface 22 may be, e.g., configured to package the encoded picture data 21 into an appropriate format, e.g. packets, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.

[0291] The communication interface 28, forming the counterpart of the communication interface 22, may be, e.g., configured to receive the transmitted data and process the transmission data using any kind of corresponding transmission decoding or processing and / or de-packaging to obtain the encoded picture data 21.

[0292] Both, communication interface 22 and communication interface 28 may be configured as unidirectional communication interfaces as indicated by the arrow for the communication channel 13 in Fig. 26 pointing from the source device 12 to the destination device 14, or bi-directional communication interfaces, and may be configured, e.g. to send and receive messages, e.g. to set up a connection, to acknowledge and exchange any other information related to the communication link and / or data transmission, e.g. encoded picture data transmission. The decoder 30 is configured to receive the encoded picture data 21 and provide decoded picture data 31 or a decoded picture 31.

[0293] The post-processor 32 of destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), e.g. the decoded picture 31, to obtain post-processed picture data 33, e.g. a post-processed picture 33. The post-processing performed by the post-processing unit 32 may comprise, e.g. color format conversion (e.g. from YCbCr to RGB), color correction, trimming, or re-sampling, or any other processing, e.g. for preparing the decoded picture data 31 for display, e.g. by display device 34.

[0294] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 for displaying the picture, e.g. to a user or viewer. The display device 34 may be or comprise any kind of display for representing the reconstructed picture, e.g. an integrated or external display or monitor. The displays may, e.g. comprise liquid crystal displays (LCD), organic light emitting diodes (OLED) displays, plasma displays, projectors, micro LED displays, liquid crystal on silicon (LCoS), digital light processor (DLP) or any kind of other display.

[0295] Although Fig. 26 depicts the source device 12 and the destination device 14 as separate devices, embodiments of devices may also comprise both or both functionalities, the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality. In such embodiments the source device 12 or corresponding functionality and the destination device 14 or corresponding functionality may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.

[0296] As will be apparent for the skilled person based on the description, the existence and (exact) split of functionalities of the different units or functionalities within the source device 12 and / or destination device 14 as shown in Fig. 26 may vary depending on the actual device and application.

[0297] The encoder 20 (e.g. a video encoder 20) or the decoder 30 (e.g. a video decoder 30) or both encoder 20 and decoder 30 may be implemented via processing circuitry, such as one or more microprocessors, digital signal processors (DSPs), applicationspecific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video coding dedicated or any combinations thereof. The encoder 20 may be implemented via processing circuitry 46 to embody the various modules including the neural network or its parts. The decoder 30 may be implemented via processing circuitry 46 to embody any coding system or subsystem described herein. The processing circuitry may be configured to perform the various operations as discussed later. If the techniques are implemented partially in software, a device may store instructions for the software in a suitable, non-transitory computer-readable storage medium and may execute the instructions in hardware using one or more processors to perform the techniques of this disclosure. Either of video encoder 20 and video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, for example, as shown in Fig. 27.

[0298] Source device 12 and destination device 14 may comprise any of a wide range of devices, including any kind of handheld or stationary devices, e.g. notebook or laptop computers, mobile phones, smart phones, tablets or tablet computers, cameras, desktop computers, set-top boxes, televisions, display devices, digital media players, video gaming consoles, video streaming devices(such as content services servers or content delivery servers), broadcast receiver device, broadcast transmitter device, or the like and may use no or any kind of operating system. In some cases, the source device 12 and the destination device 14 may be equipped for wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.

[0299] In some cases, video coding system 10 illustrated in Fig. 26 is merely an example and the techniques of the present application may apply to video coding settings (e.g., video encoding or video decoding) that do not necessarily include any data communication between the encoding and decoding devices. In other examples, data is retrieved from a local memory, streamed over a network, or the like. A video encoding device may encode and store data to memory, and / or a video decoding device may retrieve and decode data from memory. In some examples, the encoding and decoding is performed by devices that do not communicate with one another, but simply encode data to memory and / or retrieve and decode data from memory.

[0300] Fig. 28 is a schematic diagram of a video coding device 8000 according to an embodiment of the disclosure. The video coding device 8000 is suitable for implementing the disclosed embodiments as described herein. In an embodiment, the video coding device 8000 may be a decoder such as video decoder 30 of Fig. 26 or an encoder such as video encoder 20 of Fig. 26.

[0301] The video coding device 8000 comprises ingress ports 8010 (or input ports 8010) and receiver units (Rx) 8020 for receiving data; a processor, logic unit, or central processing unit (CPU) 8030 to process the data; transmitter units (Tx) 8040 and egress ports 8050 (or output ports 8050) for transmitting the data; and a memory 8060 for storing the data. The video coding device 8000 may also comprise optical-to-electrical (OE) components and electrical- to-optical (EO) components coupled to the ingress ports 8010, the receiver units 8020, the transmitter units 8040, and the egress ports 8050 for egress or ingress of optical or electrical signals. The processor 8030 is implemented by hardware and software. The processor 8030 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 8030 is in communication with the ingress ports 8010, receiver units 8020, transmitter units 8040, egress ports 8050, and memory 8060. The processor 8030 comprises a neural network-based codec 8070. The neural network-based codec 8070 implements the disclosed embodiments described above. For instance, the neural network-based codec 8070 implements, processes, prepares, or provides the various coding operations. The inclusion of the neural network-based codec 8070 therefore provides a substantial improvement to the functionality of the video coding device 8000 and effects a transformation of the video coding device 8000 to a different state. Alternatively, the neural network-based codec 8070 is implemented as instructions stored in the memory 8060 and executed by the processor 8030.

[0302] The memory 8060 may comprise one or more disks, tape drives, and solid-state drives and may be used as an over-flow data storage device, to store programs when such programs are selected for execution, and to store instructions and data that are read during program execution. The memory 8060 may be, for example, volatile and / or non-volatile and may be a read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).

[0303] Fig. 29 is a simplified block diagram of an apparatus that may be used as either or both of the source device 12 and the destination device 14 from Fig. 26 according to an exemplary embodiment.

[0304] A processor 9002 in the apparatus 9000 can be a central processing unit. Alternatively, the processor 9002 can be any other type of device, or multiple devices, capable of manipulating or processing information now-existing or hereafter developed. Although the disclosed implementations can be practiced with a single processor as shown, e.g., the processor 9002, advantages in speed and efficiency can be achieved using more than one processor.

[0305] A memory 9004 in the apparatus 9000 can be a read only memory (ROM) device or a random access memory (RAM) device in an implementation. Any other suitable type of storage device can be used as the memory 9004. The memory 9004 can include code and data 9006 that is accessed by the processor 9002 using a bus 9012. The memory 9004 can further include an operating system 9008 and application programs 9010, the application programs 9010 including at least one program that permits the processor 9002 to perform the methods described here. For example, the application programs 9010 can include applications 1 through N, which further include a video coding application that performs the methods described here.

[0306] The apparatus 9000 can also include one or more output devices, such as a display 9018. The display 9018 may be, in one example, a touch sensitive display that combines a display with a touch sensitive element that is operable to sense touch inputs. The display 9018 can be coupled to the processor 9002 via the bus 9012.

[0307] Although depicted here as a single bus, the bus 9012 of the apparatus 9000 can be composed of multiple buses. Further, a secondary storage can be directly coupled to the other components of the apparatus 9000 or can be accessed via a network and can comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. The apparatus 9000 can thus be implemented in a wide variety of configurations.

[0308] Fig. 30 is a block diagram of a video coding system 10000 according to an embodiment of the disclosure.

[0309] A platform 10002 in the system 10000 can be could a server or local server. Alternatively, the platform 10002 can be any other type of device, or multiple devices, capable of calculation, storing, transcoding, encryption, rendering, decoding or encoding. Although the disclosed implementations can be practiced with a single platform as shown, e.g., the platform 10002, advantages in speed and efficiency can be achieved using more than one platform.

[0310] A content delivery network (CDN) 10004 in the system 10000 can be a group of geographically distributed servers. Alternatively, the CDN 10004 can be any other type of device, or multiple devices, capable of data buffering, scheduling, dissemination or speed up the delivery of web content by bringing it closer to where users are. Although the disclosed implementations can be practiced with a single CDN as shown, e.g., the CDN 10004, advantages in speed and efficiency can be achieved using more than one CDN. A terminal 10006 in the apparatus 10000 can be a mobile phone, computer, television, laptop, camera. Alternatively, the terminal 10006 can be any other type of device, or multiple devices, capable of displaying video or image.

Claims

CLAIMS1. An image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processor(s) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

2. An image processing device as claimed in claim 1, wherein the processors) is / are configured to form the data units such that in the first dimension of the image one tile represents an edge region of the image and padding data.

3. An image processing device as claimed in claim 2, wherein the padding data is of a constant value.

4. An image processing device as claimed in claim 2, wherein the padding data repeats data of the edge region of the image extending in a second dimension orthogonal to the first dimension.

5. An image processing device as claimed in any of claims 2 to 4, wherein the processor(s) is / are configured to form the encoded image data to include an indication of the extent of the padding data in the first dimension.

6. An image processing device as claimed in claim 1, wherein the processors) is / are configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other.

7. An image processing device as claimed in claim 6, wherein one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image.

8. An image processing device as claimed in claim 6 or 7, wherein the processor(s) is / are configured to form the encoded image data to include an indication of the width of the overlap between the tiles of the first pair.

9. An image processing device as claimed in any preceding claim, wherein the processor(s) is / are configured to form the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set.

10. An image processing device as claimed in any preceding claim, wherein the processors) is / are configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent.

11. An image processing device as claimed in any preceding claim, wherein the processor(s) is / are configured to perform an encoding or latent prediction operation on the tiles to form the data units.

12. An image processing device as claimed in claim 11, wherein the latent prediction operation comprises one of a hyperencoding operation and a multistage context modelling operation.

13. An image processing device as claimed in any preceding claim, wherein the processor(s) is / are configured to determine whether to form the data units as claimed in claim 2 or as claimed in claim 6 in dependence on the extent of the image in the first dimension.

14. An image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processor(s) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and one tile of the set represents an edge region of the image and padding data such that all the tiles of the set are of the same extent in the first dimension.

15. An image processing device for encoding an image to form encoded image data, the image processing device comprising one or more processors configured to process an image to form a series of data units in a latent space, each data unit representing a respective tile in the image; the processors) being configured to form the data units such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and the overlap between a first pair of two tiles of the set that are adjacent to each other is larger than the overlap between a second pair of two tiles of the set that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension.

16. A method for encoding an image to form encoded image data, the method comprising processing an image to form a series of data units in a latent space, each data unit representing a respective tile in the image, such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

17. A method as claimed in claim 16, comprising forming the data units such that in the first dimension of the image one tile represents an edge region of the image and padding data.

18. A method as claimed in claim 17, wherein the padding data is of a constant value.

19. A method as claimed in claim 17, wherein the padding data repeats data of the edge region of the image extending in a second dimension orthogonal to the first dimension.

20. A method as claimed in any of claims 17 to 19, comprising forming the encoded image data to include an indication of the extent of the padding data in the first dimension.

21. A method as claimed in claim 16 comprising forming the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other.

22. A method as claimed in claim 21 , wherein one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image.

23. A method as claimed in claim 21 or 22, comprising forming the encoded image data to include an indication of the width of the overlap between the tiles of the first pair.

24. A method as claimed in any of claims 16 to 23, comprising forming the encoded image data to include an indication of which tile of the set is the said one of the tiles of the set.

25. A method as claimed in any of claims 16 to 24, comprising forming the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent.

26. A method as claimed in any of claims 16 to 25, comprising performing an encoding or latent prediction operation on the tiles to form the data units.

27. A method as claimed in claim 26, wherein the latent prediction operation comprises one of a hyper-encoding operation and a multistage context modelling operation.

28. A method as claimed in any of claims 16 to 27, comprising determining whether to form the data units as claimed in claim 17 or as claimed in claim 21 in dependence on the extent of the image in the first dimension.

29. An image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

30. An image processing device as claimed in claim 29, wherein the processor(s) is / are configured to form the data units such that in the first dimension of the image one tile represents an edge region of the image space and padding data.

31. An image processing device as claimed in claim 30, wherein the padding data is of a constant value.

32. An image processing device as claimed in claim 31, wherein the padding data repeats data of the edge region extending in a second dimension orthogonal to the first dimension.

33. An image processing device as claimed in any of claims 30 to 32, wherein the processors) is / are configured to form the encoded image data to determine an extent of the padding data in the first dimension either (i) in accordance with a value signalled in the encoded image data or (ii) in dependence on an extent of the image in the first dimension signalled in the encoded image data.

34. An image processing device as claimed in claim 29, wherein the processor(s) is / are configured to form the data units such that in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other.

35. An image processing device as claimed in claim 34, wherein one tile of the first pair has an edge extending in a second dimension orthogonal with the first dimension that is coincident with an edge of the image.

36. An image processing device as claimed in claim 34 or 35, wherein the processors) is / are configured to determine the width of the overlap between the tiles of the first pair in accordance with a value signalled in the encoded image data.

37. An image processing device as claimed in any of claims 29 to 36, wherein the processors) is / are configured to select as the said one of the tiles of the set a tile whose identity is signalled as such in the encoded image data.

38. An image processing device as claimed in any of claims 29 to 37, wherein the processor s) is / are configured to form the data units such that in a second dimension of the image orthogonal to the first dimension: (i) each tile overlaps at least one neighbouring tile and (ii) the tiles are of the same extent.

39. An image processing device as claimed in any of claims 29 to 38, wherein the processor s) is / are configured to perform a decoding or latent prediction operation on the tiles to form the data units.

40. An image processing device as claimed in claim 39, wherein the latent prediction operation comprises one of a hyperdecoding operation and a multistage context modelling operation.

41. An image processing device as claimed in any of claims 29 to 40, wherein the processor(s) is / are configured to determine whether to form the data units as claimed in claim 30 or as claimed in claim 34 in dependence on one of (i) the extent of the image in the first dimension and (ii) a value signalled in the encoded image data.

42. An image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension of the image one tile representing an edge region of the image space and padding data such that all the tiles of the set are of the same extent in the first dimension.

43. An image processing device for decoding encoded image data to recover an image, the image processing device comprising one or more processors configured to process data units in a latent space, each data unit representing a respective tile in the image and such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, and in the first dimension the overlap between a first pair of two tiles that are adjacent to each other is larger than the overlap between a second pair of two tiles that are adjacent to each other such that all the tiles of the set are of the same extent in the first dimension.

44. A method for decoding encoded image data to recover an image by processing data units in a latent space, each data unit representing a respective tile in the image, the method being such that for a set of tiles extending fully across the image in a first dimension, each of the tiles of the set overlapping at least one neighbouring tile of the set, the extent of the image in the first dimension uniquely represented by one of the tiles of the set differs from that represented by any of the other tiles of the set such that all the tiles of the set are of the same extent in the first dimension.

45. A data carrier storing data defining instructions for one or more processors to perform the steps of any of claims 15 to 28

Citation Information

Patent Citations

  • Decoding and encoding of neural-network-based bitstreams

    WO2022128105A1

  • Efficiently performing inference computations of a fully convolutional network for inputs with different sizes

    WO2023075742A1

  • Parallel processing of image regions with neural networks – decoding, post filtering, and rdoq

    WO2024002496A1