Neural network with variable number of channels and method of operation thereof
By designing the number of channels conversion formula in the neural network, ensuring that the number of channels per layer is an integer, the accuracy and data loss caused by channel number arithmetic errors and rounding processing are solved, and higher decoding accuracy and calculation speed are achieved.
Patent Information
- Application Number
- CN202380074079.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-09-11
- Publication Date
- 2025-06-17
AI Technical Summary
The arithmetic errors in the number of channels during encoding and decoding of existing neural networks lead to reduced decoding accuracy, and traditional methods require rounding processing to lead to data loss.
A neural network is designed where the number of input channels for each neural network layer can be converted into integer output channels through a specific formula, ensuring that the number of channels is always integer and avoiding rounding.
Improves the overall accuracy of video decoding, avoids data loss, and improves calculation speed.
Smart Images

Figure CN120167064A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This patent application claims the priority of International Patent Application No. PCT / EP2022 / 079255, filed on October 20, 2022. The disclosure of the above patent application is incorporated herein by reference in its entirety. Technical field
[0003] Embodiments of the present invention generally relate to the field of encoding and decoding data based on a neural network architecture. Specifically, some embodiments relate to methods and apparatuses for encoding and decoding images and / or videos in a bitstream using multiple processing layers. Background art
[0004] For decades, hybrid image and video codecs have been used to compress image and video data. In these codecs, the signal is typically encoded block by block by predicting the blocks and by further decoding only the difference between the original block and its prediction. Specifically, such decoding may include transformation, quantization, and generating a bitstream, typically including some entropy decoding. Generally, the three components (transformation, quantization, and entropy decoding) of the hybrid decoding method are optimized separately. Modern video compression standards, such as High-Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC), also use a transform representation to decode the predicted residual signal.
[0005] Recently, neural network architectures have been applied to image and / or video decoding. Generally, these neural network (NN)-based methods can be applied to image and video decoding in various different ways. For example, some end-to-end optimized image or video decoding frameworks have been discussed. In addition, deep learning has been used to determine or optimize some parts of the end-to-end decoding framework, such as the selection of prediction parameters or compression, etc. In addition, some neural network-based methods for hybrid image and video decoding frameworks have been discussed, for example, implemented as a trained deep learning model for intra-frame prediction or inter-frame prediction in image or video decoding.
[0006] A common feature of the end-to-end optimized image or video decoding applications discussed above is that some feature map data is generated, which will be transmitted between the encoder and the decoder.
[0007] A neural network is a machine learning model that uses one or more layers of non-linear units, based on which these models can predict an output for the received input. Some neural networks include one or more hidden layers in addition to the output layer. A corresponding feature map can be provided as the output of each hidden layer. Such a corresponding feature map of each hidden layer can be used as the input to subsequent layers in the network (i.e., subsequent hidden layers or the output layer). Each layer of the network generates an output from the received input according to the current values of the corresponding set of parameters. In a neural network partitioned between devices (e.g., between an encoder and a decoder, between a device and the cloud, or between different devices), the feature map at the output of the partition location (e.g., the first device) is compressed and transmitted to the remaining layers of the neural network (e.g., transmitted to the second device).
[0008] It may be necessary to further improve the encoding and decoding using a trained network architecture. Summary of the Invention
[0009] The present invention is set forth in the appended set of claims. The above and other objects are achieved by the subject matter claimed in the independent claims. Other implementations are apparent from the dependent claims, the description, and the drawings.
[0010] Specific embodiments are outlined in the appended independent claims, and other embodiments are outlined in the dependent claims.
[0011] According to a first aspect, there is provided a neural network including a first neural network layer and a second neural network. The first neural network is configured to obtain (receive) a first number of channels C in as input and output a second number of channels C out , where the first number of channels is different from the second number of channels, and C out = p * C in / q, where C in is a multiple of q, and C in、 C out , p, and q are (positive) integers. The second neural network layer is configured to obtain (receive) the second number of channels C out as input. For example, the C in input channels can be or include color components, such as RGB or YGB color components. For example, the C out output channels are in a latent space (see the detailed description below). For example, the integer q is equal to 2. For example, p / q = 5 / 2 or p / q = 2.
[0012] Contrary to the prior art, the neural network according to the first aspect ensures that the number of channels in the neural network is an integer on the output side of the first neural network layer, wherein the number of channels on the input side of the first neural network layer in the neural network changes to the number of channels on the output side of the first neural network layer. For example, for each neural network layer in the neural network layers of the neural network, for its respective input channel number C in and its respective output channel number C out , the condition C out = p * C in / q can be satisfied, where C in is a multiple of q, and C in , C out , p, and q are (positive) integers. Alternatively, for the neural network layers of the neural network, the condition is satisfied, and the number of channels of one neural network layer in the neural network layer changes to the number of channels of another neural network layer. Contrary to the prior art, all the channel numbers in the neural network can be restricted to integer values, and the number of channels of one neural network layer in the neural network changes to the number of channels of another neural network layer.
[0013] In many applications, especially in the context of video decoding, even very small arithmetic errors cannot be tolerated. For example, the quality of the synthesized video frames obtained according to the autoencoder and / or the very large scale decoder (see the description below) may be severely affected by arithmetic / numerical errors caused by floating-point operations, especially the floating-point operations performed at the decoding end on a computer hardware architecture that is different from the computer hardware architecture used at the encoding end. Specifically, the number of channels in the neural network is restricted to an integer, and the number of channels of one neural network layer in the neural network changes to the number of channels of another neural network layer, which improves the overall decoding accuracy and avoids the need for rounding processing of fractional channels, which usually results in data / information loss.
[0014] The features described with reference to the following implementation manners can be appropriately combined.
[0015] According to one implementation manner, the neural network further includes a third neural network layer, the second neural network layer is further configured to output a third channel number C', and the third neural network layer is configured to obtain (receive) the third channel number C', where C' = p' * C out / q', C out is a multiple of q', and C', p', and q' are (positive) integers. For example, p' / q' = 2 or p' / q' = 5 / 2. Similar conditions can be allowed for the output channels of the third neural network layer. The third channel number C' can be related to the first channel number C in and / or the second channel number Cout is different. Condition C' = p' * C out / q' advantageously ensures that for the first neural network layer and the second neural network layer, both the number of input channels and the number of output channels are integers.
[0016] When the neural network includes first, second, and third neural network layers, different additional conditions can also be satisfied. According to one implementation, the number of third channels is less than the number of second channels, and the number of second channels is less than the number of first channels. Therefore, a continuous reduction in the tensor size can be achieved while keeping the number of channels as integer values. According to another implementation, the number of third channels is less than the number of second channels, and the number of second channels is greater than the number of first channels. Therefore, an increase in the tensor size (e.g., for time - assumed storage) is achieved, followed by a decrease in the tensor size (e.g., after the input data is used for hypothesis selection), while keeping the number of channels as integer values. According to another implementation, the number of third channels is greater than the number of second channels, and the number of second channels is greater than the number of first channels. Therefore, a continuous increase in the tensor size can be achieved while keeping the number of channels as integer values.
[0017] According to another implementation, the neural network does not allow the number of third channels to be greater than the number of second channels, and does not allow the number of second channels to be less than the number of first channels. In other words, in the neural network provided by the first aspect and its implementations, the situation where the number of third channels is greater than the number of second channels and the number of second channels is less than the number of first channels is not applicable. Therefore, data loss and (inaccurate) recreation of data can be prevented.
[0018] According to another implementation, the number of first channels C in and the number of second channels C out are multiples of the block group size. The same can be allowed for the number of third channels C'. In addition, all the channel numbers of the neural network can be multiples of the block group size. For example, the block group size is 16 or 32. Generally, the first channel number C in and / or the second channel number C outThe third channel number C' and / or all channel numbers can be multiples of 16 or 32. The block group size defines the number of output channels that can be processed in parallel by the relevant computing module. The entire output channel number can be split into block groups to process each block group with the block group size in parallel. Making the channel number a multiple of the block group size can further increase the speed of the involved computations (e.g., the computations involved in the video decoding process). Generally, providing a channel number that is a multiple of 16 or 32 can increase the speed of the involved computations (e.g., the computations involved in the video decoding process).
[0019] As already mentioned, the neural network according to the first aspect and its implementations can be advantageously used for video decoding purposes. According to one implementation, the first neural network layer or the subnet of the neural network is a very large scale decoder subnet. According to another implementation, the first neural network layer or the subnet of the neural network is a prediction fusion subnet. However, this application is not limited to the first neural network layer or the subnet of the neural network being a very large scale decoder subnet or a prediction fusion subnet.
[0020] It should be noted that the following hierarchical structure of terms is generally used herein. A neural network includes (neural network) subnets, also referred to as nets, while a (neural network) subnet includes one or more neural network layers. A subnet can include one or more other subnets. Sometimes, the term neural network layer can be used for a neural network subnet.
[0021] According to another implementation, the first neural network layer includes a data path for at least one of the primary component and the secondary component (corresponding to the first channel number C in ). The primary component can be one of the RGB or YUV color components, and the secondary component can be the other of the RGB or YUV color components. The decoding of the secondary component can be based on the primary component, and the tensor size involved in the processing of the secondary component can be smaller than the tensor size involved in the processing of the primary component.
[0022] According to another implementation of the neural network according to the first aspect, the neural network includes at least one neural network subnet composed of consecutive neural network layers; for the at least one neural network subnet, for consecutive neural network layers, at least one of the following conditions is satisfied, where the channel number of one neural network layer in the consecutive neural network layers becomes the channel number of another neural network layer:
[0023] (a) Each neural network layer in the consecutive neural network layers is used to output only channels whose number is smaller or larger than the number of channels received from the previous neural network layer in the consecutive neural network layers in the processing order.
[0024] (b) The continuous neural network layer is composed of a first subset of the continuous neural network layer and a second subset of the continuous neural network layer in the processing order;
[0025] (i) Each neural network layer in the first subset of the continuous neural network layer is used to output only channels with a quantity greater than the quantity of channels received from the previous neural network layer in the first subset of the continuous neural network layer in the processing order;
[0026] (ii) Each neural network layer in the second subset of the continuous neural network layer is used to output only channels with a quantity smaller than the quantity of channels received from the previous neural network layer in the second subset of the continuous neural network layer in the processing order;
[0027] (c) The continuous neural network layer is composed of a first subset of the continuous neural network layer and a second subset of the continuous neural network layer in the processing order;
[0028] (i) Each neural network layer in the first subset of the continuous neural network layer is used to output only channels with a quantity smaller than the quantity of channels received from the previous neural network layer in the first subset of the continuous neural network layer in the processing order;
[0029] (ii) None of the neural network layers in the second subset of the continuous neural network layer is used to output channels with a quantity greater than the quantity of channels received from the previous neural network layer in the second subset of the continuous neural network layer in the processing order;
[0030] (iii) The first neural network layer in the second subset of the continuous neural network layer in the processing order is used to output only channels with a quantity smaller than the quantity of channels received from the last neural network layer in the first subset of the continuous neural network layer in the processing order.
[0031] Therefore, the additional conditions of the above implementation are further developed.
[0032] According to another implementation, each neural network layer in the continuous neural network layer is used to output a multiple of 16 or 32 channels.
[0033] According to another implementation, the at least one neural network subnet is one of a very large-scale decoder subnet and a prediction fusion subnet.
[0034] According to a second aspect, a method for operating a neural network having a variable number of channels in neural network layers is provided. The method for operating a neural network according to the second aspect and its implementations described below can be implemented in a neural network according to the first aspect and its implementations. Additionally, a neural network according to the first aspect and its implementations can be used by the method according to the second aspect. The method according to the second aspect and its implementations provides advantages similar to those described above.
[0035] According to the second aspect, the method for operating a neural network having a variable number of channels in neural network layers includes the following steps: a first neural network layer obtains (receives) a first number of channels C in as input; the first neural network layer outputs a second number of channels C out , where the first number of channels is different from the second number of channels, C out =p*C in / q, where C in is a multiple of q, C in , C out , p, and q are (positive) integers; a second neural network layer obtains (receives) the second number of channels C out as input. The integer q can be equal to 2. For example, p / q = 5 / 2 or p / q = 2.
[0036] According to one implementation, the method according to the second aspect further includes: the second neural network layer outputs a third number of channels C', and a third neural network layer obtains (receives) the third number of channels C', where C' = p'*C out / q', C out is a multiple of q', and C', p', and q' are (positive) integers. For example, p' / q' = 2 or p' / q' = 5 / 2.
[0037] According to another implementation, one of the following conditions is satisfied: the third number of channels is less than the second number of channels, and the second number of channels is less than the first number of channels; or the third number of channels is less than the second number of channels, and the second number of channels is greater than the first number of channels; or the third number of channels is greater than the second number of channels, and the second number of channels is greater than the first number of channels.
[0038] According to another implementation, it is not applicable that the third number of channels is greater than the second number of channels and the second number of channels is less than the first number of channels.
[0039] According to another implementation, the first number of channels C in and the second number of channels C outIt is a multiple of the block group size. The block group size can be 16 or 32.
[0040] According to another implementation, the first neural network layer or the subnet of the neural network is a very large-scale decoder network.
[0041] According to another implementation, the first neural network layer or the subnet of the neural network is a prediction fusion network.
[0042] According to another implementation, the first neural network layer includes data paths for main components or / and secondary components.
[0043] According to another implementation, the neural network includes at least one neural network subnet composed of consecutive neural network layers; for the at least one neural network subnet, for consecutive neural network layers, at least one of the following conditions is satisfied, and the number of channels of one neural network layer in the consecutive neural network layers becomes the number of channels of another neural network layer:
[0044] (a) Each neural network layer in the consecutive neural network layers only outputs channels whose number is smaller or larger than the number of channels received from the previous neural network layer in the consecutive neural network layers in the processing order.
[0045] (b) The consecutive neural network layers are composed of a first subset of consecutive neural network layers and a second subset of consecutive neural network layers in the processing order;
[0046] (i) Each neural network layer in the consecutive neural network layers in the first subset only outputs channels whose number is larger than the number of channels received from the previous neural network layer in the consecutive neural network layers in the first subset in the processing order.
[0047] (ii) Each neural network layer in the consecutive neural network layers in the second subset only outputs channels whose number is smaller than the number of channels received from the previous neural network layer in the consecutive neural network layers in the second subset in the processing order.
[0048] (c) The consecutive neural network layers are composed of a first subset of consecutive neural network layers and a second subset of consecutive neural network layers in the processing order;
[0049] (i) Each neural network layer in the consecutive neural network layers in the first subset only outputs channels whose number is smaller than the number of channels received from the previous neural network layer in the consecutive neural network layers in the first subset in the processing order.
[0050] (ii) None of the consecutive neural network layers in the second subset outputs channels with a quantity greater than the quantity of channels received by the consecutive neural network layer in the second subset at the previous neural network layer in the processing order.
[0051] (iv) The first neural network layer in the consecutive neural network layers in the second subset in the processing order only outputs channels with a quantity smaller than the quantity of channels received by the last neural network layer in the consecutive neural network layers in the first subset in the processing order.
[0052] According to another implementation, the method includes: Each neural network layer in the consecutive neural network layers outputs a multiple of 16 or 32 channels.
[0053] According to another implementation, the at least one neural network subnet is one of a very large-scale decoder subnet and a prediction fusion subnet.
[0054] In addition, a method for encoding data is provided, including the steps of operating a neural network according to the second aspect or any of its implementations.
[0055] In addition, a method for decoding encoded data is provided, including the steps of operating a neural network according to the second aspect or any of its implementations.
[0056] In addition, a computer program product including program code stored in a non-transitory medium is provided, wherein when the program is executed on one or more processors, it executes the method of operating a neural network according to the second aspect or any of its implementations.
[0057] In addition, a device for encoding data is provided, the device including a processing circuit for performing the steps of operating a neural network according to the second aspect or any of its implementations.
[0058] In addition, a device for decoding data is provided, the device including a processing circuit for performing the steps of operating a neural network according to the second aspect or any of its implementations.
[0059] In addition, a device for decoding at least a part of an encoded image is provided, including a processing circuit for providing an entropy model, the entropy model including performing the steps of the method according to the second aspect or any of its implementations; processing a bitstream through a neural network according to the provided entropy model to obtain a latent tensor representing a component of the image; processing the latent tensor to obtain a tensor representing the component of the image.
[0060] In addition, a device for encoding at least a part of an image is provided, including a neural network according to the first aspect or any implementation thereof.
[0061] In addition, a device for decoding at least a part of an encoded image is provided, including a neural network according to the first aspect or any implementation thereof.
[0062] Details of one or more embodiments are set forth in the drawings and the description. Other features, objects, and advantages will be apparent from the description, the drawings, and the claims. Description of the Drawings
[0063] The embodiments of the present invention will be described in detail below with reference to the drawings, where:
[0064] Figure 1 A schematic diagram of channels processed by layers of a neural network;
[0065] Figure 2 A schematic diagram of an autoencoder type of neural network;
[0066] Figure 3A A schematic diagram of an exemplary network architecture including an encoding end and a decoding end with a hyperprior model;
[0067] Figure 3B A schematic diagram of a general network architecture including an encoding end with a hyperprior model;
[0068] Figure 3C A schematic diagram of a general network architecture including a decoding end with a hyperprior model;
[0069] Figure 4 A schematic diagram of an exemplary network architecture including an encoding end and a decoding end with a hyperprior model;
[0070] Figure 5 A block diagram of the structure of a cloud technology solution for machine-based tasks such as machine vision tasks;
[0071] Figure 6A A block diagram of an end-to-end video compression framework based on a neural network;
[0072] Figure 6B A block diagram of some exemplary details of a neural network application for motion field compression;
[0073] Figure 6C A block diagram of some exemplary details of a neural network application for motion compensation;
[0074] Figure 7A An example showing an AI-learnable image codec architecture;
[0075] Figure 7B illustrates several examples of convolution and transposed convolution (C in >C out and C in <C out );
[0076] Figure 8 is a block diagram of an exemplary encoding terminal subnet;
[0077] Figure 9a is a block diagram of an exemplary decoding terminal subnet, where the decoding terminal subnet includes a very large scale decoder subnet and a prediction fusion subnet;
[0078] Figure 9b is a block diagram of an alternative architecture of the very large scale decoder subnet compared to that shown in Figure 9a ;
[0079] Figure 9c is a block diagram of an alternative architecture of the prediction fusion subnet compared to that shown in Figure 9a ;
[0080] Figure 10 illustrates some exemplary common NN elements;
[0081] Figure 11 illustrates an example of the implementation manner of an embodiment compared to the traditional method;
[0082] Figure 12 illustrates another example of the implementation manner of an embodiment compared to the traditional method;
[0083] Figure 13 is a block diagram of an example of an encoding device or a decoding device;
[0084] Figure 14 is a block diagram of another example of an encoding device or a decoding device;
[0085] Figure 15 is a block diagram of an exemplary video decoding system for implementing the embodiment of the present invention;
[0086] Figure 16 is a block diagram of another exemplary video decoding system for implementing the embodiment of the present invention;
[0087] Figure 17 is a block diagram of an example of an encoding device or a decoding device;
[0088] Figure 18 is a block diagram of another example of an encoding device or a decoding device;
[0089] Figure 19 is a block diagram of another example of an encoding device or a decoding device;
[0090] Figure 20 Block diagram of a neural network including a first neural network layer and a second neural network layer provided for an embodiment;
[0091] Figure 21 Flowchart of a method for operating a neural network having a variable number of neural network layer channels provided for an embodiment.
[0092] Similar reference numerals and names in different figures may represent similar elements. Detailed Description
[0093] In the following description, reference is made to the accompanying drawings which form a part hereof, and which show, by way of illustration, specific aspects of embodiments of the present invention or specific aspects in which embodiments of the present invention may be used. It is to be understood that embodiments of the present invention may be used in other aspects and include structural or logical changes not depicted in the drawings. Accordingly, the following detailed description should not be taken in a limiting sense, and the scope of the present invention is defined by the appended claims.
[0094] For example, it should be understood that the disclosure related to the described method is equally applicable to a device or system corresponding to the method for performing the method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, e.g., functional units, for performing the one or more described method steps (e.g., one unit for performing the one or more steps, or multiple units each performing one or more of the multiple steps), even if such one or more units are not explicitly depicted or described in the figures. On the other hand, for example, if a specific device is described in terms of one or more units (e.g., functional units), the corresponding method may include a step for performing the functions of the one or more units (e.g., one step for performing the functions of the one or more units, or multiple steps each performing the functions of one or more of the multiple units), even if such one or more steps are not explicitly depicted or described in the figures. Additionally, it should be understood that, unless otherwise specifically stated, the features of the various exemplary embodiments and / or aspects described herein may be combined with each other.
[0095] In the following, an overview of some technical terms and frameworks that may be used in embodiments of the present invention is provided.
[0096] Artificial Neural Network
[0097] An artificial neural network (ANN), or connectionist system, is a computational system vaguely inspired by the biological neural networks that constitute animal brains. These systems "learn" to perform tasks by example, generally without being programmed with task-specific rules. For example, in image recognition, these systems might learn to identify images containing cats by analyzing exemplary images manually labeled as "cat" or "no cat" and using the results to identify cats in other images. These systems do not know beforehand that cats have fur, tails, whiskers, and cat faces, etc. Instead, these systems automatically generate recognition features from the examples they process.
[0098] ANNs are based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like a synapse in a biological brain, can send a signal to other neurons. An artificial neuron receives signals, processes the signal, and can send a signal to neurons connected to it.
[0099] In an ANN implementation, the "signal" at a connection is a real number, and the output of each neuron is computed by some non-linear function of its input sum. These connections are called edges. Neurons and edges typically have weights that are adjusted as learning proceeds. The weights increase or decrease the signal strength at a connection. A neuron may have a threshold such that it only sends a signal if the aggregated signal exceeds that threshold. Typically, neurons are grouped into multiple layers. Different layers can perform different transformations on their inputs. Signals may travel from the first layer (input layer) to the last layer (output layer) after traversing the layers multiple times.
[0100] The original goal of the ANN approach was to solve problems in the same way as the human brain. Over time, the focus shifted to performing specific tasks, leading to a departure from biology. ANNs have been used for a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, board and video games, medical diagnosis, and even in activities traditionally considered reserved for humans, such as painting.
[0101] The name "convolutional neural network" (CNN) indicates that the network employs a mathematical operation called convolution. Convolution is a specialized linear operation. A convolutional network is a neural network that uses convolution in at least one of its layers instead of general matrix multiplication.
[0102] Figure 1 The general concept of the processing of a neural network (e.g., a CNN) is schematically illustrated. A convolutional neural network consists of an input layer, an output layer, and multiple hidden layers. The input layer provides the input (e.g.,Figure 1 a layer that processes a part 11) of the input image shown. The hidden layers in a CNN typically include a series of convolutional layers that perform convolution through multiplication or other dot products. The result of a layer is one or more feature maps (represented by empty solid rectangles), sometimes also called channels. Resampling (e.g., subsampling) may be involved in some or all of the layers. Thus, the feature maps may become smaller, as shown Figure 1 shown. It should be noted that convolution with a stride can also reduce the size of the input feature map (resampling). The activation functions in a CNN are rectified linear unit (ReLU) layers or Leaky ReLU, followed by additional convolutions, such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation functions and the final convolution. Although these layers are popularly called convolutions, this is just by convention. Mathematically, it is technically a sliding dot product or cross-correlation. This has important implications for the indexing in the matrix because it affects the way the weights are determined at a particular index point.
[0103] When programming a CNN to process images, as shown Figure 1 shown, the input is a tensor of shape (number of images) × (image width) × (image height) × (image depth). It should be known that the image depth can be composed of the channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map with the shape (number of images) × (feature map width) × (feature map height) × (feature map channels). The convolutional layer in a neural network should have the following properties. A convolutional kernel defined by width and height (hyperparameters). The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) can be equal to the number of channels (depth) of the input feature map.
[0104] In the past, traditional multilayer perceptron (MLP) models were used for image recognition. However, due to the full connection between nodes, they were affected by high dimensions and could not be fully extended in higher-resolution images. An image of 1000×1000 pixels with RGB color channels has 3 million weights, which is too high to be processed efficiently on a large scale in a fully connected case. In addition, this network architecture does not consider the spatial structure of the data, resulting in the same processing of input pixels that are far apart and those that are close together. This ignores the reference locality in the image data, both computationally and semantically. Therefore, the full connection of neurons is wasteful for purposes such as image recognition dominated by spatially local input patterns.
[0105] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons, specifically designed to mimic the behavior of the visual cortex. These models alleviate the challenges posed by the MLP architecture by exploiting the strong spatial local correlations present in natural images. The convolutional layer is the core building block of a CNN. The parameters of this layer consist of a set of learnable filters (the aforementioned kernels), which have a small receptive field but extend across the entire depth of the input volume. During the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the filter entries and the input, and generating a two-dimensional activation map for that filter. Thus, the network learns filters that activate when they detect certain specific types of features at a particular spatial location in the input.
[0106] Stacking the activation maps of all filters along the depth dimension forms the complete output volume of the convolutional layer. Thus, each entry in the output volume can also be interpreted as the output of a neuron that looks at a small region in the input and shares parameters with the neurons in the same activation map. A feature map or activation map is the output activation of a specified filter. Feature maps and activations have the same meaning. In some papers, feature maps are called activation maps because it is a mapping corresponding to the activation of different parts of an image, and it is also a feature map because it is also a mapping where certain features are found in an image. High activation means that a certain feature has been found.
[0107] Another important concept in CNNs is pooling, which is a form of non-linear downsampling. There are several non-linear functions to implement pooling, and max pooling is the most common. The input image is divided into a set of non-overlapping rectangles, and the maximum value is output for each such sub-region.
[0108] Intuitively, the exact location of a feature is less important than its approximate location relative to other features. This is the idea behind using pooling in convolutional neural networks. Pooling layers are used to gradually reduce the spatial size of the representation, reducing the number of parameters, memory footprint, and computational load in the network, and thus are also used to control overfitting. In a CNN architecture, it is common to insert pooling layers periodically between consecutive convolutional layers. The pooling operation provides another form of translation invariance.
[0109] Pooling layers operate independently on each depth strip of the input and resize it spatially. The most common form is a pooling layer that applies a filter of size 2×2 with a stride of 2 at each depth strip of the input, both along the width and height, discarding 75% of the activations. In this case, each max operation is over 4 numbers. The depth dimension remains unchanged. In addition to max pooling, pooling units can also use other functions, such as average pooling or Pooling. Average pooling was often used in the past, but has been used less frequently recently compared to max pooling, and in fact the latter often performs better. Due to the significant reduction in representation size, there has recently been a trend to use smaller filters or to discard the pooling layer altogether. "Region of Interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important part of object detection in convolutional neural networks based on the Fast R-CNN architecture.
[0110] The above ReLU is an abbreviation for the rectified linear unit, which applies a non-saturating activation function. It effectively removes negative values from the activation map by setting them to 0. It increases the non-linear nature of the decision function and the overall network without affecting the receptive field of the convolutional layer. Other functions are also used to increase non-linearity, such as the saturated hyperbolic tangent and sigmoid functions. ReLU is generally more popular than other functions because it trains neural networks several times faster without significantly affecting generalization accuracy.
[0111] Leaky rectified linear unit or Leaky ReLU is an activation function based on ReLU, but it has a small slope for negative values instead of a flat slope. The slope coefficient is determined before training, i.e., it is not learned during training. This type of activation function is common in tasks where it has sparse gradients, such as training generative adversarial networks. LeakyReLU applies an element-wise function:
[0112] LeakyReLU(x) = max(0,x) + negative_slope * min(0,x) or
[0113]
[0114] where the parameters are:
[0115] negative_slope: Controls the angle of the negative slope. Default value: 1e–2
[0116] inplace: Can choose to perform the operation in-place. Default value: False.
[0117] After several convolutional layers and max pooling layers, high-level reasoning in the neural network is done through fully connected layers. Neurons in the fully connected layer are connected to all activations in the previous layer, as shown in a regular (non-convolutional) artificial neural network. Thus, these activations can be computed as an affine transformation, a matrix multiplication followed by a bias offset (a vector addition of a learned or fixed bias term).
[0118] The "loss layer" (including the calculation of the loss function) specifies how training penalizes the deviation between the predicted (output) label and the true label, typically the last layer of a neural network. Various loss functions suitable for different tasks can be used. Softmax loss is used to predict a single class out of K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values in [0,1]. Euclidean loss is used for regression to real-valued labels.
[0119] In summary, Figure 1 The data flow in a typical convolutional neural network is shown. First, the input image passes through a convolutional layer and is abstracted into a feature map including several channels, corresponding to multiple filters in a set of learnable filters of that layer. Then, the feature map is subsampled using a pooling layer, etc., which reduces the dimension of each channel in the feature map. The data then reaches another convolutional layer, which can have a different number of output channels. As mentioned above, the number of input channels and output channels are hyperparameters of the layer. To establish the connection of the network, these parameters need to be synchronized between two connected layers, such that the number of input channels of the current layer should be equal to the number of output channels of the previous layer. For the first layer processing input data such as images, the number of input channels is usually equal to the number of channels of the data representation. For example, 3 channels are used for the RGB or YUV representation of an image or video, or 1 channel is used for the grayscale image or video representation. The channels obtained by one or more convolutional layers (and possibly resampling layers) can be passed to the output layer. In some implementations, such an output layer can be convolutional or resampling. In one exemplary and non-limiting implementation, the output layer is a fully connected layer.
[0120] Autoencoders and Unsupervised Learning
[0121] An autoencoder is a type of artificial neural network for learning efficient data decoding in an unsupervised manner. A schematic diagram is as Figure 2 shown. The autoencoder includes an encoding end 210 and a decoding end 250. The encoding end 210 has an input x input into the input layer of the encoder subnet 220, and the decoding end 250 has an output x' output from the decoder subnet 260. The purpose of the autoencoder is to learn a representation (encoding) 230 of the dataset x by training the networks 220, 260 to ignore the signal "noise", typically for dimensionality reduction. Together with the reduction (encoder) side subnet 220, the reconstruction (decoder) side subnet 260 is learned, where the autoencoder attempts to generate a representation x' as close as possible to its original input x from the reduced encoding 230, hence the name. In the simplest case, given a hidden layer, the encoder stage of the autoencoder takes the input x and maps it to h
[0122] h = σ(Wx + b).
[0123] This image h is often referred to as the code 230, latent variable, or latent representation. Here, σ is an element-wise activation function, e.g., the sigmoid function or the rectified linear unit. W is a weight matrix and b is a bias vector. The weights and biases are typically initialized randomly and then iteratively updated during training via backpropagation. Subsequently, the decoder stage of the autoencoder maps h to a reconstruction x′ of the same shape as x:
[0124] x′ = σ′(W′h′ + b′)
[0125] where the σ′, W′, and b′ of the decoder can be independent of the corresponding σ, W, and b of the encoder.
[0126] Variational autoencoder models make strong assumptions about the distribution of the latent variables. These models use variational methods for latent representation learning, resulting in an additional loss component and a specific estimator for the training algorithm, called the Stochastic Gradient Variational Bayes (SGVB) estimator. Assume that the data is generated by a directed graphical model p θ (x|h), and the encoder is learning an approximation q θ (h|x) of the posterior distribution p φ (h|x), where φ and θ denote the parameters of the encoder (recognition model) and decoder (generative model), respectively. The probability distribution of the latent vectors of VAEs is generally closer to the probability distribution of the training data than that of standard autoencoders. The objective of VAE has the following form:
[0127]
[0128] Here, D KL denotes the Kullback–Leibler divergence. The prior of the latent variable is typically set to a centered isotropic multivariate Gaussian Typically, the shapes of the variational and likelihood distributions are chosen such that they are factorized Gaussians:
[0129]
[0130] where ρ(x) and ω 2 (x) are the encoder outputs, while μ(h) and σ 2 (h) are the decoder outputs.
[0131] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, have interested researchers in applying neural network-based techniques to image and video compression tasks. For example, end-to-end optimized image compression using a variational autoencoder-based network has been proposed.
[0132] Therefore, data compression is considered a fundamental and well-studied problem in engineering, typically to design a code with minimum entropy for a given discrete data set. This technical solution relies heavily on knowledge of the probability structure of the data, so the problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as a vector of image pixel intensities) must be quantized into a finite set of discrete values, which introduces error.
[0133] In this case, the lossy compression problem, two conflicting costs must be traded off: the entropy of the discretized representation (rate) and the error introduced by quantization (distortion). Different compression applications, such as data storage or transmission over a finite-capacity channel, require different balances of rate and distortion.
[0134] Joint optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional space is intractable. Therefore, most existing image compression methods operate by linearly transforming the data vector into a suitable continuous-valued representation, independently quantizing its elements, and then encoding the resulting discrete representation using a lossless entropy code. Due to the central role of the transform, this scheme is called transform coding.
[0135] For example, JPEG uses the discrete cosine transform on pixel blocks, and JPEG 2000 uses multi-scale orthogonal wavelet decomposition. Typically, the three components of a transform coding method (transform, quantization, and entropy coding) are optimized separately (usually by manually adjusting parameters). Modern video compression standards, such as HEVC, VVC, and EVC, also use transform representations to decode the predicted residual signals. Several transforms are used for this purpose, such as the discrete cosine transform (DCT) and the discrete sine transform (DST), as well as the low-frequency non-separable manually optimized transform (LFNST).
[0136] Variational Image Compression
[0137] The Variational Auto-Encoder (VAE) framework can be regarded as a non-linear transformation coding model. The transformation process can mainly be divided into four parts. This is illustrated in the Figure 3A which is given as an example.
[0138] The transformation process can mainly be divided into four parts: Figure 3A illustrates the VAE framework. In Figure 3A the encoder 101 maps the input image x to a latent representation (represented by y) through the function y = f(x). Hereinafter, this latent representation can also be referred to as a part of or a point in the "latent space". The function f() is a transformation function that converts the input signal x into a representation y that can be further compressed. The quantizer 102 transforms the latent representation y into a quantized latent representation where the (discrete) value is Q represents the quantizer function. The entropy model or hyper encoder / decoder (also called the hyper prior) 103 estimates the distribution of the quantized latent representation to obtain the minimum rate achievable by lossless entropy source decoding.
[0139] The latent space can be understood as the representation of compressed data, where similar data points are closer in the latent space. The latent space is very useful for learning data features and finding a simpler representation of the data for analysis. The quantized latent representation T, and the side information of the hyper prior 3 are included in the bitstream 2 (after being binarized) using arithmetic coding (AE). In addition, a decoder 104 is also provided, which transforms the quantized latent representation into a reconstructed image signal which is an estimate of the input image x. It is desired that x is as close as possible to that is, the reconstruction quality is as high as possible. However, the higher the similarity between Figure 3A and x, the greater the amount of side information required for transmission. The side information includes the bitstream 1 and the bitstream 2 shown in Figure 3A which are generated by the encoder and transmitted to the decoder. Generally, the greater the amount of side information, the higher the reconstruction quality. However, a large amount of side information means a low compression ratio. Therefore,
[0140] In Figure 3A the component AE 105 is an arithmetic coding module that converts the samples of the quantized latent representation and the side information into a binarized representation bitstream 1. For example, and Examples can include integers or floating - point numbers. One purpose of the arithmetic coding module is to convert the sample values into a binary - digit string (through a binarization process), and then the binary - digit string is included in the bitstream, which can include other parts corresponding to the encoded image or other side information).
[0141] Arithmetic decoding (AD) 106 is a process of restoring the binarization process, in which the binary digits are converted back to sample values. Arithmetic decoding is provided by the arithmetic decoding module 106.
[0142] It should be noted that the present invention is not limited to this specific framework. In addition, the present invention is not limited to image or video compression and can also be applied to object detection, image generation, and recognition systems.
[0143] In Figure 3A there are two sub - networks cascaded with each other. In this context, a sub - network is a logical division between parts of the entire network. For example, in Figure 3A modules 101, 102, 104, 105, and 106 are referred to as the "encoder / decoder" sub - network. The "encoder / decoder" sub - network is responsible for encoding (generating) and decoding (parsing) the first bitstream "bitstream 1". Figure 3A The second network in includes modules 103, 108, 109, 110, and 107 and is referred to as the "super encoder / decoder" sub - network. The second sub - network is responsible for generating the second bitstream "bitstream 2". The purposes of these two sub - networks are different.
[0144] The first sub - network is responsible for:
[0145] · Transforming the input image x by 101 into its latent representation y (which is easier to compress x),
[0146] · Quantizing the latent representation y by 102 into a quantized latent representation
[0147] ● Using the arithmetic coding module 105 to compress the quantized latent representation with AE to obtain the bitstream "bitstream 1".
[0148] ● Using the arithmetic decoding module 106 to parse the bitstream 1 through AD,
[0149] ● Reconstructing the reconstructed image by 104 using the parsed data
[0150] The purpose of the second subnet is to obtain the statistical properties of the "bitstream 1" samples (e.g., the mean, variance, and correlation between the bitstream 1 samples), so that the first subnet can compress the bitstream 1 more efficiently. The second subnet generates a second bitstream "bitstream 2", which includes the information (e.g., the mean, variance, and correlation between the bitstream 1 samples).
[0151] The second network includes an encoding part, and the encoding part includes transforming the quantized latent representation by transformation 103 into side information z, quantizing the side information z into quantized side information and encoding the quantized side information 109 (e.g., binarization) into the bitstream 2. In this example, the binarization is performed by arithmetic encoding (AE). The decoding part of the second network includes arithmetic decoding (AD) 110, which transforms the input bitstream 2 into decoded quantized side information which may be the same because the arithmetic encoding and arithmetic decoding operations are lossless compression methods. Then, the decoded quantized side information is transformed 107 into decoded side information representing the statistical properties (e.g., the mean of the samples of, or the variance of the sample values, etc.). Then, the decoded latent representation is provided to the above-mentioned arithmetic encoder 105 and arithmetic decoder 106 to control the probability model of.
[0152] Figure 3A Describes an example of a variational auto encoder (VAE), and its details may vary in different implementations. For example, in a specific implementation, there may be other components to more efficiently obtain the statistical properties of the bitstream 1 samples. In this implementation, there may be a context modeler whose purpose is to extract relevant information from the bitstream 1. The statistical information provided by the second subnet can be used by the arithmetic encoder (AE) 105 and arithmetic decoder (AD) 106 components.
[0153] Figure 3A The encoder and decoder are depicted in a single figure. Those skilled in the art understand that the encoder and decoder can and often are embedded in different devices from each other.
[0154] Figure 3B Describes the encoder, Figure 3CDescribes the decoder component of the VAE framework. According to some embodiments, the encoder receives an image as input. The input image may include one or more channels, for example, color channels or other types of channels, such as depth channels or motion information channels, etc. The output of the encoder (as Figure 3B shown) is bitstream 1 and bitstream 2. Bitstream 1 is the output of the first subnet of the encoder, and bitstream 2 is the output of the second subnet of the encoder.
[0155] Similarly, in Figure 3C , the two bitstreams (bitstream 1 and bitstream 2) are received as input, and a reconstructed (decoded) image is generated at the output As mentioned above, the VAE can be divided into different logical units that perform different operations. This is illustrated in Figure 3B and Figure 3C , Figure 3B depicts components involved in encoding signals such as video and provides encoding information. Then, this encoding information is received by components such as the decoder in Figure 3C for decoding. It should be noted that the functions of the components of the encoder and decoder represented by the numbers 12x and 14x can correspond to the components represented by the number 10x mentioned above in Figure 3A .
[0156] Specifically, as shown in Figure 3B , the encoder includes encoder 121, which transforms the input x into signal y and then provides signal y to quantizer 322. Quantizer 122 provides information to arithmetic coding module 125 and super encoder 123. Super encoder 123 provides the above-mentioned bitstream 2 to super decoder 147, and super decoder 147 in turn provides information to arithmetic coding module 105 (125).
[0157] The output of the arithmetic coding module is bitstream 1. Bitstream 1 and bitstream 2 are the outputs of signal encoding, and then this output is provided (transmitted) to the decoding process. Although unit 101 (121) is called an "encoder", the complete subnet described in Figure 3B can also be called an "encoder". The encoding process generally refers to the unit (module) that converts the input into an encoded (e.g., compressed) output. As can be seen from Figure 3B , unit 121 can actually be regarded as the core of the entire subnet because it performs the conversion of input x to y, which is a compressed version of x. For example, the compression in encoder 121 can be achieved by applying a neural network or any processing network that generally has one or more layers. In such a network, the compression can be performed through cascaded processing including downsampling, which reduces the size and / or the number of channels of the input. Therefore, for example, the encoder can be called a neural network (NN)-based encoder, etc.
[0158] The remaining parts in the figure (quantization unit, super encoder, super decoder, arithmetic encoder / decoder) are all parts that improve the efficiency of the encoding process or are responsible for converting the compressed output y into a series of bits (bitstream). Quantization can be provided to further compress the output of the NN encoder 121 through lossy compression. The AE 125 can perform binarization together with the super encoder 123 and the super decoder 127 used to configure the AE 125, and the quantized signal can be further compressed through lossless compression. Therefore, the entire subnet in Figure 3B can also be referred to as an "encoder".
[0159] Most deep learning (DL)-based image / video compression systems reduce the dimension of the signal before converting it into binarized digits (bits). For example, in the VAE framework, the encoder that performs the non-linear transformation maps the input image x to y, where the width and height of y are smaller than those of x. Since y has a smaller width and height, its size is smaller, and the (size) dimension of the signal is reduced, making it easier to compress the signal y. It should be noted that generally, the encoder does not necessarily need to reduce the size in two (or usually all) dimensions. Instead, some exemplary implementations can provide an encoder that reduces the size only in one (or usually a subset) dimension.
[0160] In the arXiv e-print version published by J. Balle, L. Valero Laparra, and E. P. Simoncelli (2015), "Density Modeling of Images Using a Generalized Normalization Transformation", at the 4th International Conference on Learning Representations (hereinafter referred to as "Balle") in 2016, the authors proposed an end-to-end optimization framework for an image compression model based on non-linear transformation. The authors optimized the mean squared error (MSE), but used a more flexible transformation constructed by linear convolution and non-linear cascading. Specifically, the authors used the generalized divisive normalization (GDN) joint non-linearity, which was inspired by the neuron model in the biological visual system and has been proven to be effective in Gaussianizing image density. This cascaded transformation is followed by uniform scalar quantization (i.e., each element is rounded to the nearest integer), which helps to achieve a parametric form of vector quantization on the original image space. The compressed image is reconstructed from these quantized values using an approximate parametric non-linear inverse transformation.
[0161] An example of the VAE framework is as follows Figure 4 as shown, which utilizes 6 downsampling layers, labeled 401 to 406. The network architecture includes a hyperprior model. On the left (g a , g s ) shows the image autoencoder architecture, and on the right (h a , h s ) corresponds to the autoencoder that implements the hyperprior. The factored prior model uses the same architecture for the analysis and synthesis transforms g a and g s . Q represents quantization, and AE and AD represent the arithmetic encoder and arithmetic decoder respectively. The encoder inputs the input image x into g a , generating a response y (latent representation) with spatially varying standard deviation. The encoding g a includes multiple convolutional layers, with subsampling and generalized divisive normalization (GDN) as the activation function.
[0162] The response is fed into h a to summarize the standard deviation distribution in z. Then z is quantized, compressed, and transmitted as side information. Then, the encoder uses the quantization vector to estimate i.e., the spatial distribution of the standard deviation, to obtain the probability value (or frequency value) for arithmetic decoding (AE) and use it to compress and transmit the quantized image representation (or latent representation). The decoder first recovers from the compressed signal. Then, it uses h s to obtain which provides it with the correct probability estimate to successfully recover . Then, is fed into g s to obtain the reconstructed image.
[0163] The layers that include downsampling are indicated by a downward arrow in the layer description. The layer description "Conv N,k1,2↓" means that the layer is a convolutional layer with N channels and a convolution kernel size of k1×k1. For example, k1 can be equal to 5 and k2 can be equal to 3. As described above, 2↓ means that downsampling by a factor of 2 is performed in this layer. Downsampling by a factor of 2 causes one dimension of the input signal to be reduced by half at the output. In Figure 4In it, 2↓ means that both the width and height of the input image are reduced by half. Since there are 6 downsampling layers, if the width and height of the input image 414 (also denoted as x) are given by w and h, the width and height of the output signal z^413 are equal to w / 64 and h / 64 respectively. The modules represented by AE and AD are the arithmetic encoder and the arithmetic decoder, which will be explained with reference to Figures 3A to 3C Arithmetic encoders and decoders are specific implementations of entropy decoding. AE and AD can be replaced by other entropy decoding methods. In information theory, entropy coding is a lossless data compression scheme used to convert the values of symbols into a binary representation, which is a reversible process. In addition, "Q" in the figure corresponds to the quantization operation mentioned above regarding Figure 4 and is further explained in the "Quantization" section above. In addition, the quantization operation and the corresponding quantization unit as part of component 413 or 415 do not necessarily exist and / or can be replaced by another unit.
[0164] In Figure 4 a decoder including upsampling layers 407 to 412 is also shown. Another layer 420 is provided between upsampling layers 411 and 410 in the processing order of the input, which is implemented as a convolutional layer but does not provide upsampling for the received input. The corresponding convolutional layer 430 for the decoder is also shown. Such a layer can be provided in the NN to perform an operation on the input that does not change the input size but changes specific features. However, it is not necessary to provide such a layer.
[0165] When seen from the processing order of the code stream 2 passing through the decoder, the upsampling layers run in the reverse order, that is, from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide upsampling with an upsampling ratio of 2, denoted by ↑. Of course, not all upsampling layers necessarily have the same upsampling ratio, and other upsampling ratios such as 3, 4, 8, etc. can also be used. Layers 407 to 412 are implemented as convolutional layers (conv). Specifically, since they can provide an operation opposite to that of the encoder on the input, the upsampling layers can apply a transposed convolution operation to the received input so that its size increases by a factor corresponding to the upsampling ratio. However, the present invention is generally not limited to transposed convolution, and upsampling can be performed in any other way, for example, by bilinear interpolation between two adjacent samples, or by nearest neighbor sample replication, etc.
[0166] In the first subnet, after some convolutional layers (401 to 403) at the encoding end, generalized divisive normalization (GDN) is applied, and inverse GDN (IGDN) is applied at the decoding end. In the second subnet, the activation function applied is ReLu. It should be noted that the present invention is not limited to this implementation manner, and generally, other activation functions can be used to replace GDN or ReLu.
[0167] Cloud technology solutions applicable to machine tasks
[0168] Machine Video Coding (VCM) is another popular direction in computer science today. The main idea behind this approach is to transmit an encoded representation of image or video information for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. Compared with traditional image and video decoding for human perception, the quality metric is the performance of computer vision tasks, e.g., object detection accuracy, rather than reconstruction quality. This is shown in Figure 5 as follows.
[0169] Machine video coding, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks in mobile cloud infrastructure. By partitioning the network between the mobile side 510 and the cloud side 590 (e.g., cloud server), the computational workload can be distributed, thus minimizing the overall energy and / or latency of the system. Generally speaking, collaborative intelligence is a paradigm in which the processing of a neural network is distributed among two or more different computing nodes (e.g., devices, but generally any functionally defined node). Here, the term "node" does not refer to the above-mentioned neural network nodes. Instead, the (computing) nodes here refer to (physically or at least logically) separate devices / modules that implement parts of the neural network. Such devices can be different servers, different end-user devices, mixtures of servers and / or user devices and / or clouds and / or processors, etc. In other words, the computing nodes can be regarded as nodes belonging to the same neural network and communicate with each other to transmit decoded data within / for the neural network. For example, in order to be able to perform complex calculations, one or more layers can be executed on a first device (e.g., a device on the mobile side 510), and one or more layers can be executed on another device (e.g., a cloud server on the cloud side 590). However, the distribution can also be more refined, and a single layer can be executed on multiple devices. In the present invention, the term "multiple" means two or more than two. In some prior art solutions, a part of the neural network function is executed in a device (user device or edge device, etc.) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a collection of processing or computing systems located outside the device, and the device is operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training. In this case, the data flow is two-way: from the cloud to the mobile during backpropagation in training, from the mobile to the cloud during forward pass in training (as shown in Figure 5 ), and inference.
[0170] Some works have proposed semantic image compression by encoding deep features and then reconstructing the input image from these features. Compression based on uniform quantization followed by context-based adaptive arithmetic coding (CABAC) of H.264 is shown. In some scenarios, it may be more efficient to transmit the output of the hidden layer (deep feature map) 550 from the mobile part 510 to the cloud 590 rather than sending compressed natural image data to the cloud and performing object detection using the reconstructed image. Therefore, it may be beneficial to compress the data (features) generated by the mobile side 510, and the mobile side 510 may include a quantization layer 520 for this purpose. Accordingly, the cloud side 590 may include an inverse quantization layer 560. Efficient compression of the feature map is beneficial for image and video compression and reconstruction, whether for human perception or machine vision. Entropy coding methods (such as arithmetic coding) are popular methods for compressing deep features (i.e., feature maps).
[0171] Today, video content contributes more than 80% to Internet traffic, and this proportion is expected to rise further. Therefore, it is crucial to establish an efficient video compression system and generate higher-quality frames within a given bandwidth budget. In addition, most video-related computer vision tasks, such as video object detection or video object tracking, are sensitive to the quality of compressed videos, and efficient video compression may bring benefits to other computer vision tasks. At the same time, video compression technology also helps with action recognition and model compression. However, in the past few decades, video compression algorithms have relied on handcrafted modules, such as block-based motion estimation and Discrete Cosine Transform (DCT), to reduce redundancy in video sequences, as described above. Although each module is well designed, the entire compression system is not optimized end-to-end. It is hoped to further improve video compression performance by jointly optimizing the entire compression system.
[0172] End-to-end image or video compression
[0173] DNN-based image compression methods can utilize large-scale end-to-end training and highly non-linear transformations, which are not used in traditional methods. However, it is not common to directly apply these techniques to build an end-to-end learning system for video compression. First, learning how to generate and compress motion information tailored for video compression remains an open question. Video compression methods rely heavily on motion information to reduce temporal redundancy in video sequences.
[0174] A simple technical solution is to use learning-based optical flow to represent motion information. However, the current learning-based optical flow methods aim to generate the flow field as accurately as possible. Precise optical flow is usually not the best choice for specific video tasks. In addition, compared with the motion information in traditional compression systems, the data volume of optical flow has increased significantly. Directly applying existing compression methods to compress optical flow values will significantly increase the number of bits required to store motion information. Secondly, it is not clear how to construct a DNN-based video compression system through rate-distortion objectives that minimize residuals and motion information. The purpose of rate-distortion optimization (RDO) is to achieve higher-quality (i.e., less distorted) reconstructed frames when given the number of compressed bits (or bitrate). RDO is very important for video compression performance. To leverage the power of end-to-end training of learning-based compression systems, an RDO strategy is needed to optimize the entire system.
[0175] In the proceedings of the 2019 IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), "DVC: An End-to-end Deep Video Compression Framework" on pages 11006 to 11015 by Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao, the authors proposed an end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual decoding.
[0176] This encoder is as Figure 6A shown. Specifically, Figure 6A shows the overall structure of the end-to-end trainable video compression framework. To compress motion information, a CNN is specified to transform the optical flow v t into a corresponding representation m t that is suitable for better compression. Specifically, an autoencoder-style network is used to compress the optical flow. The motion vector (MV) compression network is as Figure 6B shown. The network architecture is somewhat similar to the g Figure 4 in a / g s Specifically, the optical flow v tIt is fed into a series of convolutional operations and non - linear transformations including GDN and IGDN. The number of output channels c of the convolution (deconvolution) is exemplified as 128, but in this example, the last deconvolutional layer is equal to 2. The kernel size is k, for example, k = 3. Given the optical flow size of M×N×2, the MV encoder will generate a motion representation m of size M / 16×N / 16×128 t . Then, the motion representation is quantized (Q), entropy - decoded and sent as the bitstream. The MV decoder receives the quantized representation and uses the MV encoder to reconstruct the motion information Generally, the values of k and c can be different from the above - mentioned embodiments known in the art.
[0177] Figure 6C The structure of the motion compensation part is shown. Here, using the previously reconstructed frame x t-1 and the reconstructed motion information, the warping unit generates a warped frame (usually, by means of an interpolation filter, for example, a bilinear interpolation filter). Then, a separate CNN with three inputs generates a predicted image. The architecture of the motion compensation CNN is also as Figure 6C shown.
[0178] The residual information between the original frame and the predicted frame is encoded by the residual encoder network. A highly non - linear neural network is used to transform the residual into the corresponding latent representation. Compared with the discrete cosine transform in traditional video compression systems, this method can better utilize the power of non - linear transformation and achieve higher compression efficiency.
[0179] As can be seen from the above overview, considering different parts of the video frame, including motion estimation, motion compensation, and residual coding, the CNN - based architecture can be applied to image and video compression. Entropy decoding is a popular method for data compression, widely adopted in the industry, and is also applicable to the compression of feature maps for human perception or computer vision tasks.
[0180] Machine Video Coding
[0181] Machine Video Coding (VCM) is another popular direction in computer science today. The main idea behind this method is to transmit the encoded representation of image or video information for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. Compared with traditional image and video decoding for human perception, the quality metric is the performance of computer vision tasks, such as object detection accuracy, rather than the reconstruction quality.
[0182] A recent study proposed a new deployment paradigm called collaborative intelligence, which partitions deep models between the mobile and the cloud. Extensive experiments under various hardware configurations and wireless connection modes have shown that the optimal operating points in terms of energy consumption and / or computational latency involve partitioning the model, typically at some point deep in the network. The common scenarios today, where the model is either fully in the cloud or fully in the mobile, are rarely (if ever) optimal. The concept of collaborative intelligence has also been extended to model training. In this case, the data flow is two-way: from the cloud to the mobile during backpropagation in training, from the mobile to the cloud during the forward pass in training, and for inference.
[0183] In the context of recent deep models for object detection, lossy compression of deep feature data was studied based on HEVC intra coding. As the compression level increases, the detection performance degrades, and compression-enhanced training was proposed to minimize this loss by generating a model that is more stable to quantization noise in the feature values. However, this is still a sub-optimal technical solution because the codec used is very complex and optimized for natural scene compression rather than deep feature compression.
[0184] By a method that utilizes the popular YOLOv2 network for object detection tasks, the trade-off between compression efficiency and recognition accuracy was studied, addressing the deep feature compression problem of collaborative intelligence. Here, the term "deep feature" has the same meaning as "feature map". The word "deep" comes from the idea of collaborative intelligence in the case of capturing the output feature maps of certain hidden (deep) layers and transmitting them to the cloud for inference. This seems to be more efficient than sending compressed natural image data to the cloud and performing object detection using the reconstructed images.
[0185] Efficient compression of feature maps is beneficial for image and video compression and reconstruction, whether for human perception or machine vision. The drawbacks of state-of-the-art autoencoder-based compression methods also apply to machine vision tasks.
[0186] The quality metrics in the JPEG AI Common Training and Test Conditions (CTTC) are selected based on their correlation with human perception. Instead of a single metric, JPEG AI uses 7 different metrics that are sensitive to different types of artifacts. It should be noted that most quality metrics are calculated in the YUV color space, and some use only the luminance component. The JPEG AI learnable image codec is built on the assumption that one color component contains most of the information and has a greater impact on human perception than other color components. To extract this strongest color component, the RGB input first enters a color transformation. The default color transformation is RGB→YUV BT.709 (full range). Adaptive color transformations with signal color transformation matrices are supported.
[0187] Figure 7A An example of the AI learnable image codec architecture is shown. The primary and secondary components are encoded separately using networks with the same architecture but different numbers of channels. In Figure 7A , the neural subnets / data streams used only at the encoding end are marked by dashed boxes / lines. The solid boxes represent the subnets used on the decoder (or at the decoding and encoding ends). All boxes with the same name are subnets with the same architecture, except for the input / output tensor sizes and the number of channels. In the present invention, the subnets are also referred to as Nets.
[0188] Inside the neural network, the number of channels of the latent tensor can vary layer by layer. Figure 7B Several typical cases of convolution and transposed convolution with different numbers of channels in the input and output tensors are shown. Generally, the neural network algorithm includes the relationship equation between C in (the number of channels of the input tensor) and C out (the number of channels of the output tensor), which is usually a ratio C out = p * C in / q.
[0189] In learnable image coding, the input signal to be encoded is denoted as x, the latent space tensor in the variational autoencoder bottleneck is y, and the hyperparameter (the tensor in the hyperencoder bottleneck) is z. The prediction fusion network produces a prediction μ of the latent space tensor y. The residual signal r = y - μ is quantized and encoded, assuming a Gaussian distribution with mean and variance σ equal to zero, which is the output of the hyperscale decoder network.
[0190] The input image has width W and height H. The primary component and the related tensors are marked as "Y" in Figure 7A . The primary component is encoded at full resolution (the size of x Y is H × W). The output of the analysis transformation network (for the primary component is Cp ×h Y ×w Y )。 The number of channels allocated to the main component decoding is C p = 128, and the output of the super encoder network of the main component is C p ×h hpY ×w hpY 。
[0191] The main components are jointly decoded. The secondary components and the associated tensors are labeled "UV" in Figure 7A . The secondary components are decoded at a lower resolution. Therefore, both the secondary tensor color components and the main tensor color components before the analysis transformation network pass through the downsampling module. The input to the analysis transformation network of the secondary components is a concatenated tensor of three color components (at a reduced resolution), and the size of this tensor is W / 2 × W / 2 × 3. The output of the analysis transformation network of the secondary components is C s ×h YUV ×w UV 。 The number of channels C allocated to the secondary component decoding s = 64. The output of the super encoder network of the main component is C s ×h hpUV ×w hpUV 。
[0192] The prediction subnet (e.g., the prediction fusion network in Figure 7A ) runs independently and has the same architecture for the main and secondary components. The prediction subnet outputs a tensor of the same size w × h × C as the latent space tensor y. The prediction subnet receives tensors of the same spatial size but with four times the number of channels w × h × 4C. Half of the channels are the output of the context model network and the other half are the output of the super decoder network.
[0193] Convolution with a downsampling stride is part of the JPEG AI of the learnable codec architecture. To keep the tensor size always a multiple of the downsampling stride, a padding layer is inserted before the convolution with a downsampling stride.
[0194] Figure 8 is a block diagram of an exemplary encoding terminal subnet. For example, the analysis transformation network can include a sequence of four convolutions with a downsampling stride of 2. The sizes of the tensors in different parts of the analysis transformation network are as Figure 8 shown. It can be seen that the peak memory usage occurs after the first convolutional layer.
[0195] As Figure 8 shown, each of these convolutions is preceded by a padding layer. At the padding layer, the tensor size changes as follows: h t = ceil(h t-1 , s), w t = ceil(wt-1 , s), where s is the downsampling convolution stride (s = 2).
[0196] Analyze that the non-linear activation layer in the transformation network is a residual non-linear unit with an attention mechanism (ResAU) (for example, Figure 10 the ResAU block in). ResAU consists of an element-wise non-linear operation (ReLU), a convolutional layer, an element-wise multiplication operation (tahn), and a residual connection. The detailed structure of ResAU used in the analysis transformation network (also known as the transformation module) is as Figure 10 shown.
[0197] A residual non-local attention block (RNAB) is used in the analysis transformation network. RNAB is located between the first downsampling convolution and the second downsampling convolution and acts as an attention mechanism. Since RNAB (for example, Figure 10 the RNAB in) includes a downsampling convolution and an upsampling convolution, if there is a residual block (for example, Figure 10 the RB depicted in), a padding layer is inserted before RNAB. RNAB also includes a series of residual blocks (for example, Figure 10 the RB depicted in).
[0198] The encoding end may also include a super encoder network. For example, as Figure 8 shown, the super encoder network may include a sequence that has two convolutions with a downsampling stride of 2, three convolutions without a change in tensor size, and ReLU as the activation (for example, Figure 10 the ReLU in). The sizes of the tensors in different parts of the analysis transformation are as Figure 8 shown. Similar to the analysis transformation network, there is a padding layer before each downsampling convolution.
[0199] Use a fixed probability density model to encode / decode bitstream #1 and bitstream #2 (on Figure 7A ). The discretized cumulative distribution function is stored in a predetermined fixed table for parsing the quantized hyperprior tensor Then, the quantized hyperprior tensor is processed by a very large-scale decoder network, which is an NN-based subnet for generating the Gaussian variance σ. Then, the quantized residual latent samples are obtained by applying arithmetic decoding to the second bitstream (bitstream #3 and bitstream #4) Assumed to be a zero-mean Gaussian distribution It should be noted that the entire entropy decoding process can be carried out before the start of the latent sample prediction process.
[0200] Figure 9aIt is a block diagram of the decoding terminal network. The very large-scale decoder network 910 includes a sequence that has two transposed convolutions with an upsampling stride of 2, two convolutions without tensor size changes, and Leaky ReLU as the activation that sizes the tensors of different parts of the analysis transform (e.g., Figure 10 the LeakyReLU in
[0201] ), as shown in the block diagram of Figure 9. Symmetric to the hyper-encoder, a cropping operation follows each upsampling convolutional layer. At the start of the latent sample prediction process, the very large-scale decoder network 910 performs an inverse transform operation on the hyper prior latent . The output of this process is concatenated with the output of the context model sub-network, and then this output is processed by the prediction fusion network 920 to generate the predicted value μ. Then the predicted value is added to the quantized residual sample
[0202] to obtain the quantized latent sample
[0203] The very large-scale decoder network (Figure 9) includes a sequence that has two transposed convolutions with an upsampling stride of 2, three convolutions without tensor size changes, and Leaky ReLU as the activation. Symmetric to the hyper-encoder, a cropping operation follows each upsampling convolutional layer.
[0204] As shown in Figure 9, the synthesis transform network includes a sequence of 4 convolutions with an upsampling stride of 2. The sizes of the tensors in different parts of the synthesis transform network are as shown in Figure 9. It can be noted that the peak memory usage (which is also the maximum number of multiplication operations) occurs before the last upsampling convolutional layer. The number of channels for the principal component synthesis transform is as follows: = C in = 128; C out = 1. The number of channels for the secondary component synthesis transform is as follows: C = 64; C in = 128 + 64; C out = 2.
[0205] It should be noted that the latent sample prediction process is an autoregressive process. However, the quantized latent samples of different rows can be processed in parallel.
[0206] Respectively, Figure 9b and Figure 9c show alternative architectures of the very large-scale decoder sub-network 930 and the prediction fusion sub-network 940 that can implement the embodiments of the present invention.
[0207] Figure 8 and Figure 9a 、Figure 9b and Figure 9c common NN elements in Figure 10 as shown in
[0208] In JPEG AIVMuC - 2.0, for example, in the "ultra - large - scale decoder network" and the "prediction fusion (aggregation) network", the tensor size in the channel dimension is a decimal. In the actual implementation, the decimal is rounded down and data is lost.
[0209] This problem is specific to the CCS architecture adopted by JPEG AIVMuC - 2.0 because the number of channels in the primary component pipeline (Cp) and the secondary component pipeline (Cs) is not a multiple of 3. In almost all papers, a channel number of C = 192 (a multiple of 3 because the three colors RGB are decoded together) is used.
[0210] The number of channels in a neural network (NN) sub - network usually varies layer by layer. A typical example is the aggregation sub - network, which works by fusing multiple hypotheses generated by previous sub - networks. In the present invention, the terms "aggregation" and "fusion" have the same meaning.
[0211] According to an embodiment, the NN design principles include at least one of the following conditions:
[0212] 1. If, within a sub - network, the number of channels varies layer by layer, then it (the number of channels) must be an integer at each layer. Only when C in is a multiple of q, the channel number C of the sub - network design out = p * C in / q. This can prevent data loss. For example, q equals 2.
[0213] 2. If, within a sub - network, the number of channels changes, then within a sub - network, the number of channels should not first decrease and then increase:
[0214] · If the tensor sizes (in the processing order) are C0, C1 …… Ck–1, Ck …… CN, then
[0215] · C0 > C1 > …… > Ck–1 > Ck > …… > CN indicates a decrease in the tensor size after the input data is utilized → is allowed.
[0216] · C0 < C1 < …… < Ck–1 > Ck > …… > CN indicates an increase in the tensor size for storing temporary hypotheses, then a decrease in the tensor size after the input data is utilized → is allowed.
[0217] · C0 > C1 > …… < Ck–1 <ck>……>CN indicates a decrease in the tensor size (data loss), then an increase in the tensor size
[0218] (data creation) → should not be allowed.
[0219] This can prevent data loss and further reconstruction of data.
[0220] 3. If, within a subnet, the number of channels changes, then at each layer, the number of channels must be a multiple of the block group size. For each specific implementation of the dedicated hardware for neural network algorithms, the block group size is different. The commonly used block group size is 16. In the near future, with the development of dedicated hardware incrementers, it is expected that the block group size = 32. The block group size is a power of 2.
[0221] This can provide efficient utilization of NPU / GPU resources.
[0222] To implement the above NN design principle, an embodiment of the present invention provides a neural network, as Figures 1 to 10 shown. The neural network includes a first neural network layer and a second neural network layer. The neural network layer can obtain the first number of channels (C in ) as input and output the second number of channels (C out ), where the first number of channels is different from the second number of channels, C out = p * C in / q, where C in is a multiple of q, and C in、 C out 、p, and q are integers. The second neural network layer is used to obtain the second number of channels (C out ) as input.
[0223] For example, the first neural network layer is an ultra-large decoder network, as shown in FIGS. 7 and 9. Alternatively, the first neural network layer can be a prediction fusion (aggregation) network, for example, as shown in FIGS. 7 and 9.
[0224] The first neural network layer includes data paths for the main component or / and the secondary component.
[0225] The second neural network layer can have a structure similar to that of the first neural network layer.
[0226] The neural network further includes a third neural network layer. The second neural network layer is also used to output a third number of channels (C'); the third neural network layer is used to obtain the third number of channels (C'), where C' = p' * C out / q', where C' is a multiple of q', and C', p', and q' are integers.
[0227] The third neural network layer may have a structure similar to that of the first neural network layer.
[0228] As described above, the number of third channels is less than the number of second channels, and the number of second channels is less than the number of first channels; or the number of third channels is less than the number of second channels, and the number of second channels is greater than the number of first channels; or
[0229] the number of third channels is greater than the number of second channels, and the number of second channels is greater than the number of first channels.
[0230] However, it is not applicable that the number of third channels is greater than the number of second channels and the number of second channels is less than the number of first channels.
[0231] In a possible implementation, the number of first channels (C in ) and the number of second channels (C out ) are multiples of the block group size. 8. The block group size can be 16 or 32.
[0232] To implement the embodiments of the present invention, Figure 11 and Figure 12 two examples are provided in Figure 11 and Figure 12 As shown in Figure 11 , in the conventional method, when Cp = 128, the output of Conv 5 / 3C×3×3 in the very large-scale decoder network is 213.3333, and when Cs = 64, it is 106.6667. Therefore, the output of the number of channels is a non-integer. In the Figure 11 -provided embodiment, when Cp = 128, the output of Conv 5 / 2C×3×3 in the very large-scale decoder network is 320, and when Cs = 64, it is 160. Therefore, the output of the number of channels is an integer.
[0233] Similarly, in the prediction fusion (aggregation) network, in the conventional method, when Cp = 128, the output of Conv 10 / 3C×1×1 in the very large-scale decoder network is 426.6667, and when Cs = 64, it is 213.3333. Therefore, the output of the number of channels is a non-integer. In the Figure 11 -provided embodiment, when Cp = 128, the output of Conv 7 / 2C×1×1 is 448, and when Cs = 64, it is 224. Therefore, the output of the number of channels is an integer.
[0234] Figure 11 is an alternative design for the very large-scale decoder network and the prediction fusion (aggregation) network. Figure 12 Another option for the very large-scale decoder network and the prediction fusion (aggregation) network is provided. Figure 11 The very large-scale decoder network in Figure 12 works together with the very large scale decoder network and / or the prediction fusion (aggregation) network in Figure 11 . Similarly, the prediction fusion (aggregation) network in Figure 12 can also work together with the very large scale decoder network and / or the prediction fusion (aggregation) network in
[0235] Figure 11 and Figure 12 is implemented from right to left.
[0236] In Figure 11 the example of the very large scale decoder network, p / q = 2, and p' / q' = 5 / 2. In Figure 12 the example of the very large scale decoder network, p / q = 5 / 2 and p' / q' = 2.
[0237] As Figure 11 shown in the example, the tensor sizes of the number of channels (C) in the very large scale decoder network can sequentially include 3 / 2C, 2C, 5 / 2C, and 3 / 2C. As Figure 12 shown in the example, the tensor sizes of the number of channels (C) in the very large scale decoder network can sequentially include 3 / 2C, 5 / 2C, 2C, and 3 / 2C.
[0238] In Figure 11 the example of the prediction fusion (aggregation) network, p / q = 7 / 2 and p' / q' = 3. In Figure 12 the example of the prediction fusion (aggregation) network, p / q = 9 / 2 and p' / q' = 7 / 2.
[0239] As Figure 11 shown in the example, the tensor sizes of the number of channels (C) in the prediction fusion (aggregation) network can sequentially include 7 / 2C, 3C, and 5 / 2C. As Figure 12 shown in the example, the tensor sizes of the number of channels (C) in the prediction fusion (aggregation) network can sequentially include 9 / 2C, 7 / 2C, and 5 / 2C.
[0240] In the present invention, if the number of channels inside the sub-network is variable, data loss can be prevented by keeping the number of channels as an integer in each layer. If the number of channels inside the sub-network is variable, then by keeping the number of channels as a multiple of the block group size in each layer, it provides efficient NPU / GPU resource utilization. In addition, it is not allowed that the third number of channels is greater than the second number of channels and the second number of channels is less than the first number of channels. This can prevent data loss and further reconstruct data.
[0241] Functional module
[0242] Variable Bitrate Module
[0243] The encoder can output bitstreams with different bitrates. Therefore, in some methods, the output of the encoding network is scaled (e.g., each channel is multiplied by a corresponding scale factor, which is also called the target gain value), and the input of the decoding network is inversely scaled (e.g., each channel is multiplied by the reciprocal of the corresponding scale factor, which is also called the target inverse gain value), as Figure 13 shown. Among them, the scale factor can be preset. Different quality levels or quantization parameters correspond to different target gain values. If the output of the encoding network is scaled to a smaller value, the bitstream size can be reduced. Otherwise, the bitstream size may increase.
[0244] Color Format Conversion
[0245] RGB and YUV are common color spaces. The conversion between RGB and YUV can be carried out according to the formulas specified in standards such as CCIR 601 and BT.709.
[0246] Independent Structure of Luminance and Chrominance
[0247] Some VAE-based codecs use the YUV color space as the input of the encoder and the output of the decoder, as Figure 14 shown. The Y component represents luminance, and the UV components represent chrominance. The resolution of the UV components can be equal to or less than that of the Y component. Typical formats include YUV4:4:4, YUV4:2:2, and YUV4:2:0. The Y component is converted into a feature map F_Y through the network, and the entropy encoding module generates the bitstream of the Y component according to the feature map F_Y. The UV components are converted into a feature map F_UV through another network, and the entropy encoding module generates the bitstream of the UV components according to the feature map F_UV. In this structure, the feature maps of the Y component and the UV components can be independently quantized, so as to flexibly allocate bits for luminance and chrominance. For example, for color-sensitive images, the quantization of the feature mapping of the UV components can be reduced, and the number of bits of the bitstream of the UV components can be increased to improve the reconstruction quality of the UV components and obtain better visual effects.
[0248] In some other methods, the encoder concatenates the Y component and the UV components and sends them to the UV component processing module (for converting image information into a feature map). In addition, the decoder concatenates the reconstructed feature map of the Y component and the reconstructed feature map of the UV components and sends them to the UV component processing module 2 (for converting the feature map into image information). In this method, the correlation between the Y component and the UV components can be utilized to reduce the bitstream of the UV components.
[0249] Similarly, although the figures depict operations in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageously employed. Additionally, the separation or integration of various system modules and components in the embodiments described above should not be understood to be required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0250] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. For example, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In some implementations, multitasking and parallel processing can be advantageously employed.
[0251] Figure 15 A corresponding system in which the above encoder-decoder processing chain can be deployed is shown. Figure 15 FIG. is a schematic block diagram of an exemplary decoding system, such as a video, image, audio, and / or other decoding system (or simply referred to as a decoding system) that can utilize the technology of the present application. The video encoder 20 (or simply referred to as encoder 20) and the video decoder 30 (or simply referred to as decoder 30) of the video decoding system 10 represent examples of devices that can be used to perform various techniques according to the various examples described in the present application. For example, video encoding and decoding can employ a neural network, which can be distributed and can perform the above-described bitstream parsing and / or bitstream generation to transmit feature maps between distributed computing nodes (two or more).
[0252] As Figure 15 shown, the decoding system 10 includes a source device 12, for example, the source device 12 is configured to provide encoded image data 21 to a destination device 14 for decoding the encoded image data 13.
[0253] The source device 12 includes an encoder 20, and additionally, optionally, may include a pre-processor (or pre-processing unit) 18 such as an image source 16, an image pre-processor 18, a communication interface, or a communication unit 22.
[0254] The image source 16 may include or be any type of image capture device, such as a camera for capturing real-world images, and / or any type of image generation device, such as a computer graphics processor for generating computer animated images, or any other type of device for acquiring and / or providing real-world images, computer-generated images (e.g., screen content, virtual reality (VR) images), and / or any combination thereof (e.g., augmented reality (AR) images). The image source may be any type of memory / storage that stores any of the above images.
[0255] Distinct from the processing performed by the pre-processor 18 and the pre-processing unit 18, the image or image data 17 may also be referred to as the raw image or raw image data 17.
[0256] The pre-processor 18 is used to receive the (raw) image data 17 and pre-process the image data 17 to obtain pre-processed image 19 or pre-processed image data 19. The pre-processing performed by the pre-processor 18 may include trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising, etc. It should be understood that the pre-processing unit 18 may be an optional component. It should be noted that the pre-processing may also use a neural network indicated by a presence indicator (e.g., in any one of FIGS. Figure 1 to 7).
[0257] The video encoder 20 is used to receive the pre-processed image data 19 and provide encoded image data 21.
[0258] The communication interface 22 in the source device 12 can be used to: receive the encoded image data 21 and send the encoded image data 21 (or any other processed version) to another device such as the destination device 14 or any other device through the communication channel 13 for storage or direct reconstruction.
[0259] The destination device 14 includes a decoder 30 (e.g., a video decoder 30), and additionally, optionally, may include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0260] The communication interface 28 in the destination device 14 is used to directly receive the encoded image data 21 (or any other processed version) from the source device 12 or from any other source device such as a storage device. For example, the storage device is an encoded image data storage device, and provide the encoded image data 21 to the decoder 30.
[0261] The communication interfaces 22 and 28 can be used to send or receive the encoded image data 21 or the encoded data 13 via a direct communication link (e.g., a direct wired or wireless connection) between the source device 12 and the destination device 14, or via any type of network (e.g., a wired or wireless network or any combination thereof, or any type of private and public network), or any combination thereof.
[0262] For example, the communication interface 22 can be used to encapsulate the encoded image data 21 into a suitable format such as a message, and / or process the encoded image data using any type of transport encoding or processing for transmission over the communication link or communication network.
[0263] For example, the communication interface 28 corresponding to the communication interface 22 can be used to receive the transmitted data and process the transmitted data using any type of corresponding transport decoding or processing and / or de-encapsulation to obtain the encoded image data 21.
[0264] Both the communication interface 22 and the communication interface 28 can be configured as Figure 15 a unidirectional communication interface as indicated by the arrow of the communication channel 13 pointing from the source device 12 to the destination device 14 in, or configured as a bidirectional communication interface, and can be used to send and receive messages, etc., to establish a connection, confirm, and exchange any other information related to the communication link and / or data transmission (e.g., encoded image data transmission), etc. The decoder 30 is used to receive the encoded image data 21 and provide the decoded image data 31 or the decoded image 31.
[0265] The post-processor 32 of the destination device 14 is used to post-process the decoded image data 31 (also referred to as the reconstructed image data) (e.g., the decoded image 31) to obtain the post-processed image data 33 (e.g., the post-processed image 33). The post-processing performed by the post-processing unit 32 can include color format conversion (e.g., from YCbCr to RGB), color grading, trimming, or resampling, or any other processing to enable the decoded image data 31 to be displayed by a display device 34, etc.
[0266] The display device 34 in the destination device 14 is configured to receive the post-processed image data 33 and display an image to a user, viewer, or the like. The display device 34 may be or may include any type of display for presenting the reconstructed image, such as an integrated or external display or monitor. For example, the display may include a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro-LED display, a liquid crystal on silicon (LCoS) display, a digital light processor (DLP), or any other type of display.
[0267] Although Figure 15 the source device 12 and the destination device 14 are described as separate devices, embodiments of the device may also include the source device 12 and the destination device 14 or both the corresponding functions of the source device 12 and the corresponding functions of the destination device 14. In these embodiments, the source device 12 or the corresponding function and the destination device 14 or the corresponding function may be implemented using the same hardware and / or software or by separate hardware and / or software or any combination thereof.
[0268] Based on the foregoing description, those skilled in the art will clearly appreciate that Figure 15 the presence and (exact) partitioning of the different units or functions in the source device 12 and / or the destination device 14, as shown, may vary depending on the actual device and application.
[0269] The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30) or both the encoder 20 and the decoder 30 can be implemented by processing circuitry, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, video decoding dedicated processors, or any combination thereof. The encoder 20 can be implemented by processing circuitry 46 to embody various modules including a neural network or portions thereof. The decoder 30 can be implemented by processing circuitry 46 to embody any decoding system or subsystem described herein. The processing circuitry can be used to perform various operations that will be discussed later. When the techniques are implemented partially in software, the device can store the instructions of the software in a suitable non-transitory computer-readable storage medium and can execute the instructions in hardware using one or more processors to perform the techniques of the present invention. The video encoder 20 or the video decoder 30 can be integrated as part of a combined encoder / decoder (codec) in a single device, as Figure 16 shown.
[0270] The source device 12 and the destination device 14 can include any one of a variety of devices, including any type of handheld or fixed device, such as a laptop or notebook computer, a mobile phone, a smartphone, a tablet (tablet / tablet computer), a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device (such as a content service server or a content distribution server), a broadcast receiver device, a broadcast transmitter device, etc., and may or may not use any type of operating system. In some cases, the source device 12 and the destination device 14 can be equipped with components for wireless communication. Thus, the source device 12 and the destination device 14 can be wireless communication devices.
[0271] In some cases, Figure 15 The illustrated video decoding system 10 is merely exemplary, and the techniques provided by this application are applicable to video decoding setups (e.g., video encoding or video decoding) that do not necessarily include any data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory, sent over a network, and so on. A video encoding device may encode data and store the data in memory, and / or a video decoding device may retrieve data from memory and decode the data. In some examples, encoding and decoding are performed by devices that do not communicate with each other but merely encode data into memory and / or retrieve data from memory and decode the data.
[0272] Figure 17 FIG. is a schematic diagram of a video decoding device 8000 provided for an embodiment of the present invention. The video decoding device 8000 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video decoding device 8000 may be a decoder (such as Figure 15 video decoder 30) or an encoder (such as Figure 15 video encoder 20).
[0273] The video decoding device 8000 includes an in-port 8010 (or input port 8010) for receiving data and a receiving unit (Rx) 8020, a processor, logic unit, or central processing unit (CPU) 8030 for processing data, a transmitting unit (Tx) 8040 and an out-port 8050 (or output port 8050) for transmitting data, and a memory 8060 for storing data. The video decoding device 8000 may also include optical-to-electrical (OE) components and electrical-to-optical (EO) components coupled to the in-port 8010, the receiving unit 8020, the transmitting unit 8040, and the out-port 8050 for an exit or entry of optical or electrical signals.
[0274] The processor 8030 is implemented by hardware and software. The processor 8030 can be implemented as one or more CPU chips, cores (e.g., multi-core processors), FPGAs, ASICs, and DSPs. The processor 8030 communicates with the input port 8010, the receiving unit 8020, the transmitting unit 8040, the output port 8050, and the memory 8060. The processor 8030 includes a neural-network based codec 8070. The neural-network based codec 8070 implements the embodiments disclosed above. For example, the neural-network based codec 8070 performs, processes, prepares, or provides various decoding operations. Thus, the neural-network based codec 8070 provides a substantial improvement to the functionality of the video decoding device 8000 and enables the transition of the video decoding device 8000 to different states. Alternatively, the neural-network based codec 8070 is implemented as instructions stored in the memory 8060 and executed by the processor 8030.
[0275] The memory 8060 may include one or more magnetic disks, tape drives, and solid state drives, and can be used as an overflow data storage device for storing such programs when selected for execution, and for storing instructions and data read during program execution. For example, the memory 8060 can be volatile and / or non-volatile, and can be read-only memory (ROM), random access memory (RAM), ternary content-addressable memory (TCAM), and / or static random-access memory (SRAM).
[0276] Figure 18 A simplified block diagram of an apparatus provided for an exemplary embodiment, the apparatus being usable as Figure 15 either or both of the source device 12 and the destination device 14 in
[0277] The processor 9002 in the apparatus 9000 can be a central processing unit. Alternatively, the processor 9002 can be any other type of device or devices, existing or to be developed in the future, that can manipulate or process information. Although the disclosed implementations can be implemented using a single processor such as the processor 9002 shown in the figure, using more than one processor can increase speed and efficiency.
[0278] In one implementation, the memory 9004 in the apparatus 9000 may be a read only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as the memory 9004. The memory 9004 may include code and data 9006 that can be accessed by the processor 9002 via the bus 9012. The memory 9004 may further include an operating system 9008 and application programs 9010, and the application programs 9010 include at least one program that causes the processor 9002 to execute the methods described herein. For example, the application programs 9010 may include Applications 1 to N, and further include a video decoding application that executes the methods described herein.
[0279] The apparatus 9000 may further include one or more output devices, such as a display 9018. In one example, the display 9018 may be a touch-sensitive display that combines a display with a touch-sensitive element that can be used to sense touch inputs. The display 9018 may be coupled to the processor 9002 via the bus 9012.
[0280] Although the bus 9012 in the apparatus 9000 is described herein as a single bus, the bus 9012 may include multiple buses. In addition, the auxiliary memory may be directly coupled to other components of the apparatus 9000 or may be accessed via a network, and may include a single integrated unit (such as a memory card) or multiple units (such as multiple memory cards). Therefore, the apparatus 9000 may have various configurations.
[0281] Figure 19 Block diagram of the video decoding system 10000 provided for the embodiments of the present invention.
[0282] Figure 20 The neural network 2000 provided for the embodiments of the present invention is shown. The neural network 2000 includes a first neural network layer 2010 and a second neural network 2020. The first neural network 2010 is configured to obtain (receive) a first number of channels C in as an input and output a second number of channels C out , where the first number of channels is different from the second number of channels, and C out = p*C in / q, where C in is a multiple of q, and C in , C out , p, and q are (positive) integers. The second neural network layer 2020 is configured to obtain (receive) the second number of channels C out as an input. For example, the C in input channels may be or include color components, such as RGB or YGB color components. For example, C out The number of output channels is in the latent space. For the C out number of output channels received by the first neural network layer 2010 as C in The number of output channels of the second neural network layer 2020 with the number of input channels can also meet the condition C out = p * C in / q. According to one embodiment, for the corresponding C of each neural network layer of the neural network 2000 in The number of input channels and C out The number of output channels meets this condition. Generally, the number of channels or the number of output channels can be a multiple of 16 or 32.
[0283] Compared with the prior art, restricting the variable number of channels to an integer can significantly improve the accuracy and the calculation speed.
[0284] The neural network 2000 may include Figures 1 to 10 The neural network architecture in one example shown in Figures 15 to 19 Or a device shown in
[0285] Figure 21 A method 2100 for operating a neural network with a variable number of channels in the neural network layer is shown in. For example, the method 2100 is used to operate the neural network 2000 including Figure 20 The first neural network layer 2010 and the second neural network layer 2020 shown in. The method 2100 includes the following steps: The first neural network layer obtains (receives) 2110 the first number of channels C in As input; The first neural network layer outputs 2120 the second number of channels C out , where the first number of channels is different from the second number of channels, C out = p * C in / q, where C in Is a multiple of q, C in , C out , p and q are (positive) integers; The second neural network layer obtains (receives) 2130 the second number of channels C out As input. According to one embodiment, the method 2100 further includes outputting C' = p' * C out / q', where C out Is a multiple of q', and C', p' and q' are (positive) integers, which are used as channels by the second neural network layer. The video encoding method or the video decoding method can advantageously include Figure 21 The method 2100 shown in.< / ck>
Claims
1. A neural network (2000), characterized in that, Including: The first neural network layer (2010) is used to obtain the first number of channels C in as input and output the second number of channels C out , where the first number of channels is different from the second number of channels, and C out = p * C in / q, where C in is a multiple of q, and C in , C out , p, and q are integers; The second neural network layer (2020) is used to obtain the second number of channels C out as input.
2. The neural network (2000) according to claim 1, characterized in that, q is equal to 2.
3. The neural network (2000) according to claim 1 or 2, characterized in that, The neural network (2000) further includes a third neural network layer; The second neural network layer (2020) is further configured to output a third channel number C'; The third neural network layer is used to obtain the third number of channels C', where C' = p' * C out / q', and C out is a multiple of q', and C', p', and q' are integers.
4. The neural network (2000) according to claim 3, characterized in that, The third channel number is less than the second channel number, and the second channel number is less than the first channel number; or The third channel number is less than the second channel number, and the second channel number is greater than the first channel number; or The third channel number is greater than the second channel number, and the second channel number is greater than the first channel number.
5. The neural network (2000) according to claim 3 or 4, characterized in that, The neural network (2000) does not allow the third channel number to be greater than the second channel number and does not allow the second channel number to be less than the first channel number.
6. The neural network (2000) according to any one of claims 3 to 5, characterized in that, p / q = 5 / 2 and p' / q' = 2; or p / q = 2 and p' / q' = 5 / 2.
7. The neural network (2000) according to any one of claims 1 to 6, characterized in that, The number C of the first channels in and the number C of the second channels out are multiples of the block group size.
8. The neural network (2000) according to claim 7, characterized in that, The block group size is 16 or 32.
9. The neural network (2000) according to any one of claims 1 to 8, characterized in that, The first neural network layer (2010) or a subnet of the neural network (2000) is a very large scale decoder subnet (910, 930).
10. The neural network (2000) according to any one of claims 1 to 8, characterized in that, The first neural network layer (2010) or a subnet of the neural network (2000) is a prediction fusion subnet (920, 940).
11. The neural network (2000) according to any one of claims 1 to 10, characterized in that, The first neural network layer (2010) includes a data path for at least one of a main component and a secondary component.
12. The neural network (2000) according to claim 1 or 2, characterized in that, The neural network (2000) includes at least one neural network subnet composed of consecutive neural network layers; For the at least one neural network subnet, for consecutive neural network layers, at least one of the following conditions is satisfied, and the channel number of one neural network layer in the consecutive neural network layers becomes the channel number of another neural network layer: (d) Each neural network layer in the consecutive neural network layers is configured to output only channels whose number is smaller or larger than the number of channels received from the previous neural network layer in the consecutive neural network layers in the processing order; (e) The consecutive neural network layers are composed of a first subset of consecutive neural network layers and a second subset of the consecutive neural network layers in the processing order; (iii) Each neural network layer in the consecutive neural network layers in the first subset is configured to output only channels whose number is larger than the number of channels received from the previous neural network layer in the consecutive neural network layers in the first subset in the processing order; (iv) Each neural network layer in the consecutive neural network layers in the second subset is configured to output only channels whose number is smaller than the number of channels received from the previous neural network layer in the consecutive neural network layers in the second subset in the processing order; (f) The consecutive neural network layers are composed of a first subset of consecutive neural network layers and a second subset of the consecutive neural network layers in the processing order; (v) Each neural network layer in the consecutive neural network layers in the first subset is configured to output only channels whose number is smaller than the number of channels received from the previous neural network layer in the consecutive neural network layers in the first subset in the processing order; (vi) None of the consecutive neural network layers in the second subset are used to output a number of channels greater than the number of channels received by the previous neural network layer in the consecutive neural network layers of the second subset in the processing order; (vii) The first neural network layer in the consecutive neural network layers of the second subset in the processing order is used to output only a number of channels smaller than the number of channels received by the last neural network layer in the consecutive neural network layers of the first subset in the processing order.
13. The neural network (2000) according to claim 12, characterized in that, Each neural network layer in the consecutive neural network layers is used to output a multiple of 16 or 32 channels.
14. The neural network (2000) according to claim 12 or 13, characterized in that, The at least one neural network subnet is one of the very large scale decoder subnets (910, 930) and the prediction fusion subnets (920, 940).
15. A method (2100) of operating a neural network (2000) having a variable number of neural network layer channels, characterized in that, Comprising: The first neural network layer (2010) obtains the first number of channels C in as input The first neural network layer (2010) outputs a second number of channels C out , where the first number of channels is different from the second number of channels, C out = p * C in / q, C in is a multiple of q, C in , C out , p, and q are integers; The second neural network layer (2020) obtains the second channel number C out as input.
16. The method (2100) according to claim 15, characterized in that, q equals 2.
17. The method (2100) according to claim 15 or 16, characterized in that, The method (2100) further comprises: The second neural network layer (2020) outputs a third channel number C'; The third neural network layer obtains the number C' of the third channels, where C' = p' * C out / q', and C out is a multiple of q', and C', p', and q' are integers.
18. The method (2100) according to claim 17, characterized in that, The third channel number is less than the second channel number, and the second channel number is less than the first channel number; or The third channel number is less than the second channel number, and the second channel number is greater than the first channel number; or The third channel number is greater than the second channel number, and the second channel number is greater than the first channel number.
19. The method (2100) according to claim 17 or 18, characterized in that, The case where the third channel number is greater than the second channel number and the second channel number is less than the first channel number is not applicable.
20. The method (2100) according to any one of claims 17 to 19, characterized in that, p / q = 5 / 2 and p' / q' = 2; or p / q = 2 and p' / q' = 5 / 2.
21. The method (2100) according to any one of claims 15 to 20, characterized in that, The number C of the first channels in and the number C of the second channels out are multiples of the block group size.
22. The method (2100) according to claim 21, characterized in that, The block group size is 16 or 32.
23. The method (2100) according to any one of claims 15 to 22, characterized in that, The first neural network layer (2010) or the subnet of the neural network is the very large scale decoder subnet (910, 930).
24. The method (2100) according to any one of claims 15 to 22, characterized in that, The first neural network layer (2010) or the subnet of the neural network is the prediction fusion subnet (920, 940).
25. The method (2100) according to any one of claims 15 to 24, characterized in that, The first neural network layer (2010) includes a data path for the main component and / or the secondary component.
26. The method (2100) according to claim 15 or 16, characterized in that, The neural network includes at least one neural network subnet composed of consecutive neural network layers; For the at least one neural network subnet, for consecutive neural network layers, at least one of the following conditions is satisfied, and the number of channels of one neural network layer in the consecutive neural network layers becomes the number of channels of another neural network layer: (c) Each neural network layer in the consecutive neural network layers only outputs a number of channels smaller or larger than the number of channels received by the previous neural network layer in the consecutive neural network layers in the processing order; (d) The consecutive neural network layers are composed of a first subset of consecutive neural network layers and a second subset of consecutive neural network layers in the processing order; (i) Each neural network layer in the consecutive neural network layers of the first subset only outputs a number of channels greater than the number of channels received by the previous neural network layer in the consecutive neural network layers of the first subset in the processing order; (ii) Each neural network layer in the second subset of the consecutive neural network layers outputs only a number of channels that is smaller than the number of channels received at the previous neural network layer in the second subset of the consecutive neural network layers in the processing order. (c) The consecutive neural network layers consist of a first subset of consecutive neural network layers and a second subset of the consecutive neural network layers in the processing order. (j) Each neural network layer in the first subset of the consecutive neural network layers outputs only a number of channels that is smaller than the number of channels received at the previous neural network layer in the first subset of the consecutive neural network layers in the processing order. (ii) None of the neural network layers in the second subset of the consecutive neural network layers outputs a number of channels that is larger than the number of channels received at the previous neural network layer in the second subset of the consecutive neural network layers in the processing order. (viii) The first neural network layer in the second subset of the consecutive neural network layers in the processing order outputs only a number of channels that is smaller than the number of channels received at the last neural network layer in the first subset of the consecutive neural network layers in the processing order.
27. The method (2100) according to claim 26, characterized in that, Comprising: Each neural network layer in the consecutive neural network layers outputs a multiple of 16 or 32 channels.
28. The method (2100) according to claim 26 or 27, characterized in that, The at least one neural network subnet is one of a very large scale decoder subnet (910, 930) and a prediction fusion subnet (920, 940).
29. A method (2100) for encoding data, characterized in that, Comprising steps of a method (2100) for operating a neural network (2000) according to any one of claims 15 to 28.
30. A method (2100) for decoding encoded data, characterized in that, Comprising steps of a method (2100) for operating a neural network (2000) according to any one of claims 15 to 28.
31. A computer program product comprising program code stored in a non-transitory medium, characterized in that, When the program code is executed on one or more processors, it executes the method (2100) according to any one of claims 15 to 28.
32. An apparatus (20) for encoding data, characterized in that, The apparatus comprises a processing circuit for performing the steps of the method (2100) according to any one of claims 15 to 28.
33. An apparatus (30) for decoding data, characterized in that, The apparatus comprises a processing circuit for performing the steps of the method (2100) according to any one of claims 15 to 28.
34. An apparatus (30) for decoding at least a part of an encoded image, characterized in that, Comprising a processing circuit for providing an entropy model, the entropy model comprising performing the steps of the method (2100) according to any one of claims 15 to 28; processing a bitstream through a neural network (2000) according to the provided entropy model to obtain a latent tensor representing a component of the image; and processing the latent tensor to obtain a tensor representing the component of the image.
35. An apparatus (20) for encoding at least a part of an image, characterized in that,Comprising a neural network (2000) according to any one of claims 1 to 14.
36. An apparatus (30) for decoding at least a part of an encoded image, characterized in that, Comprising a neural network (2000) according to any one of claims 1 to 14.