Neural Network for Multirate Computer Vision Tasks in Compressed Regions

The NIC framework optimizes image/video compression using ANNs for improved rate-distortion performance, enabling efficient multi-rate compression and direct computer vision tasks on compressed images, addressing the challenges of expertise and complexity in existing tools.

JP7697637B2Active Publication Date: 2025-06-24TENCENT AMERICA LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023564403
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-03-22
Filing Date
2023-03-23
Publication Date
2025-06-24
Estimated Expiration
2043-03-23

AI Technical Summary

Technical Problem

Existing image/video compression tools require significant expertise and time for improvement, and optimizing video codecs as a whole can be challenging, leading to suboptimal performance.

Method used

A neural image compression (NIC) framework using artificial neural networks (ANNs) is trained end-to-end to optimize various modules jointly for improved rate-distortion performance, allowing for multi-rate compression and computer vision tasks directly on compressed images without reconstruction.

Benefits of technology

The NIC framework enables efficient image/video compression with reduced computational complexity and faster processing times for computer vision tasks, achieving optimized compression rates and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697637000018
    Figure 0007697637000018
  • Figure 0007697637000019
    Figure 0007697637000019
  • Figure 0007697637000020
    Figure 0007697637000020
Patent Text Reader

Abstract

In some examples, the processing circuitry decodes from a coded bitstream carrying a compressed image an index pointing to a value in a set of values ​​for the parameter. Changing the value of the parameter adjusts a compression rate of the compressed image. The compressed image is generated by a neural network-based encoder based on the parameter. The processing circuitry inputs the value of the parameter to a multi-rate compressed domain computer vision task decoder. The multi-rate compressed domain computer vision task decoder includes one or more neural networks for performing computer vision tasks from the compressed image according to corresponding values ​​of the parameters used to generate the compressed image. The multi-rate compressed domain computer vision task decoder generates a computer vision task result according to the compressed image in the coded bitstream and the value of the parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Incorporation by Reference This application claims the benefit of priority to U.S. Patent Application No. 18 / 124,828, filed on March 22, 2023, entitled "MULTI-RATE OF COMPUTER VISON TASK NEURAL NETWORKS IN COMPRESSION DOMAIN", which claims the benefit of priority to Provisional Application No. 63 / 331,168, filed on April 14, 2022, entitled "Multi-rate of Computer Vison Task Neural Networks in Compression Domain". The disclosures of these prior applications are hereby incorporated by reference in their entirety.

[0002] This disclosure generally describes embodiments related to image / video processing.

Background Art

[0003] The description of the background art provided herein is for the purpose of generally presenting the content of the present disclosure. The work of the inventors whose names are currently described may not be recognized as prior art at the time of filing, either explicitly or implicitly, in the same manner as the aspects described in this background art paragraph that are not within the scope of the work described in this background art paragraph, and are not admitted as prior art to the present disclosure.

[0004] Image / video compression can help transmit image / video files between different devices, storage, and networks while minimizing quality degradation. Improving image / video compression tools can require a lot of expertise, effort, and time. Applying machine learning techniques to image / video compression can simplify and speed up the improvement of compression tools.

Summary of the Invention

[0005] Aspects of the present disclosure provide methods and apparatus for image / video processing (e.g., encoding and decoding). In some examples, an apparatus for image / video processing includes a processing circuit. The processing circuit decodes an index that refers to a value within a set of parameter values from a coded bitstream that conveys a compressed image. By changing the parameter values, the compression rate of the compressed image is adjusted. The compressed image is generated by a neural network-based encoder based on the parameters. The processing circuit inputs the parameter values to a multi-rate compressed region computer vision task decoder within a compressed region computer vision task framework (CDCVTF). The multi-rate compressed region computer vision task decoder includes one or more neural networks for performing computer vision tasks from the compressed image according to corresponding values of parameters used to generate the compressed image at a plurality of different compression rates. The multi-rate compressed region computer vision task decoder generates computer vision task results according to the compressed image within the coded bitstream compressed at a corresponding compression rate from a plurality of different compression rates based on the parameter values.

[0006] In some examples, a first neural network within the multi-rate compressed region computer vision task decoder converts the parameter values into a tensor. The tensor is input to one or more layers of a second neural network within the multi-rate compressed region computer vision task decoder. The second neural network generates computer vision task results according to the compressed image and the tensor.

[0007] In some examples, the first neural network includes one or more convolutional layers.

[0008] In some examples, the first neural network includes convolutional layers having activation functions.

[0009] In some examples, the second neural network is configured to generate computer vision task results without generating a reconstructed image from the compressed image.

[0010] In some examples, the second neural network is configured to generate a reconstructed image from the compressed image and generate computer vision task results from the reconstructed image.

[0011] In some examples, the neural network-based encoder is based on an encoder model in a neural image compression (NIC) framework, the multirate compression region computer vision task decoder is based on a decoder model in the NIC framework, and the NIC framework is trained end-to-end.

[0012] In some examples, the decoder model of the multirate compression region computer vision task decoder is trained separately from the encoder model of the neural network-based encoder.

[0013] In some examples, the parameter is a hyperparameter for weighting distortion in the calculation of rate-distortion loss.

[0014] In some examples, the computer vision task includes at least one of image classification, image noise removal, object detection, and super-resolution.

[0015] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for encoding and / or decoding an image / video.

Brief Description of the Drawings

[0016] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Mode for Carrying Out the Invention

[0017] According to one aspect of the present disclosure, some video codecs may be difficult to optimize as a whole. For example, even if a single module (e.g., an encoder) of a video codec is improved, the coding gain of the overall performance may not be improved. In contrast, in an artificial neural network (ANN)-based video / image coding framework, a machine learning process is executed, and then various modules of the ANN-based video / image coding framework are jointly optimized from input to output to improve the final goal (e.g., rate-distortion performance such as the rate-distortion loss L described in the present disclosure). For example, the learning process or training process (e.g., a machine learning process) is executed on the ANN-based video / image coding framework, and the modules of the ANN-based video / image coding framework are jointly optimized to achieve an overall optimized rate-distortion performance. Thus, the result of the optimization can be end-to-end (E2E) optimized neural image compression (NIC).

[0018] In the following description, the ANN-based video / image coding framework is represented by a neural image compression (NIC) framework. Although image compression (e.g., encoding and decoding) will be described in the following description, it should be noted that the technology of image compression can also be appropriately applied to video compression.

[0019] According to some aspects of the present disclosure, the NIC framework can be trained in an offline training process and / or an online training process. In the offline training process, a set of previously collected training images can be used to train the NIC framework and optimize the NIC framework. In some examples, the parameters of the NIC framework determined by the offline training process can be referred to as pre-trained parameters, and the NIC framework including the pre-trained parameters can be referred to as a pre-trained NIC framework. The pre-trained NIC framework can be used for image compression operations.

[0020] In some examples, when one or more images (also referred to as one or more target images) are available for the image compression operation, the pre-trained NIC framework can be further trained based on the one or more target images in the online training process to adjust the parameters of the NIC framework. The parameters of the NIC framework adjusted by the online training process can be referred to as online-trained parameters, and the NIC framework including the online-trained parameters can be referred to as an online-trained NIC framework. The online-trained NIC framework can then perform an image compression operation on one or more target images. Some aspects of the present disclosure provide techniques for online-training-based encoder adjustment in neural image compression.

[0021] A neural network refers to a computational architecture that models the biological brain. A neural network is a model implemented in software or hardware that emulates the computational capabilities of biological systems using a large number of artificial neurons connected via connection lines. Artificial neurons, called nodes, are connected to each other and operate collectively to process input data. A neural network (NN) is also known as an artificial neural network (ANN).

[0022] The nodes within an ANN can be arranged in any suitable architecture. In some embodiments, the nodes within an ANN are arranged in layers that include an input layer that receives input signals into the ANN and an output layer that outputs output signals from the ANN. In one embodiment, the ANN further includes a layer called a hidden layer between the input layer and the output layer. Different layers can perform different types of transformations on their respective inputs. Signals can be transmitted from the input layer to the output layer.

[0023] An ANN that includes multiple layers between the input layer and the output layer can be referred to as a deep neural network (DNN). A DNN can have any suitable structure. In some examples, the DNN is configured with a feed-forward neural network structure in which data flows from the input layer to the output layer without loop-back. In some examples, the DNN is configured with a fully connected network structure in which each node in one layer is connected to all the nodes in the next layer. In some examples, the DNN is configured with a recurrent neural network (RNN) structure in which data can flow in any direction.

[0024] An ANN that includes at least a convolutional layer that performs a convolutional operation can be called a convolutional neural network (CNN). A CNN can include an input layer, an output layer, and hidden layers between the input layer and the output layer. The hidden layer can include a convolutional layer (such as used in an encoder) that performs a convolution such as a two-dimensional (2D) convolution. In one embodiment, the 2D convolution performed in the convolutional layer is performed between a convolutional kernel (also called a filter or channel such as a 5×5 matrix) and an input signal (such as a 2D matrix such as a 256×256 matrix) for the convolutional layer. The dimension of the convolutional kernel (such as 5×5) is smaller than the dimension of the input signal (such as 256×256). During the convolutional operation, a dot product operation is performed on the convolutional kernel and a patch (such as a 5×5 region) within an input signal (such as a 256×256 matrix) of the same size as the convolutional kernel, and an output signal for input to the next layer is generated. A patch (such as a 5×5 region) within an input signal (such as a 256×256 matrix) that is the size of the convolutional kernel can be called the receptive field of each node in the next layer.

[0025] During convolution, a dot product of the convolutional kernel and the corresponding receptive field within the input signal is calculated. The convolutional kernel includes weights as elements, and each element of the convolutional kernel is a weight applied to the corresponding sample within the receptive field. For example, a convolutional kernel represented by a 5×5 matrix has 25 weights. In some examples, a bias is applied to the output signal of the convolutional layer, and the output signal is based on the sum of the dot product and the bias.

[0026] In some examples, the convolutional kernel can be shifted along the input signal (e.g., a 2D matrix) by a size called the stride, and thus the convolution operation generates a feature map or activation map (e.g., another 2D matrix), which then contributes to the input of the next layer of the CNN. For example, the input signal is a 2D block with 256×256 samples, and the stride is 2 samples (e.g., stride 2). When the stride is 2, the convolutional kernel shifts 2 samples at a time along the X direction (e.g., the horizontal direction) and / or the Y direction (e.g., the vertical direction).

[0027] In some examples, multiple convolutional kernels can be applied to the input signal in the same convolutional layer to generate multiple feature maps respectively, and each feature map can represent a specific feature of the input signal. In some examples, the convolutional kernel can correspond to the feature map. A convolutional layer that includes N convolutional kernels (or N channels), (each convolutional kernel has M×M samples,) and has a stride S can be specified as Conv:M×M cN sS. For example, a convolutional layer that includes 192 convolutional kernels (or 192 channels), (each convolutional kernel has 5×5 samples,) and has a stride of 2 is specified as Conv:5×5 c192 s2. The hidden layer can include a deconvolutional layer (e.g., used in a decoder) that performs deconvolution such as 2D deconvolution. Deconvolution is the inverse of convolution. A deconvolutional layer that includes 192 deconvolutional kernels (or 192 channels), (each deconvolutional kernel has 5×5 samples,) and has a stride of 2 is specified as DeConv:5×5 c192 s2.

[0028] In a CNN, a relatively large number of nodes can share the same filter (e.g., the same weights) and the same bias (if a bias is used), and thus a single bias and a single weight vector can be used across all receptive fields that share the same filter, reducing the memory footprint. For example, in the case of an input signal having 100×100 samples, a convolutional layer with a convolutional kernel having 5×5 samples has 25 learnable parameters (e.g., weights). If a bias is used, then next, one channel uses 26 learnable parameters (e.g., 25 weights and 1 bias). If there are N convolutional kernels in the convolutional layer, the total number of learnable parameters is 26×N. The number of learnable parameters is relatively small compared to a fully connected feedforward neural network layer. For example, in the case of a fully connected feedforward layer, 100×100 (i.e., 10,000) weights are used to generate a resulting signal for inputting to each node in the next layer. If there are L nodes in the next layer, then the total number of learnable parameters is 10,000×L.

[0029] A CNN can further include one or more other layers such as a pooling layer, a fully connected layer(s) that can connect all nodes of one layer to all nodes of another layer, and / or a normalization layer. The layers of a CNN can be arranged in any suitable order and suitable architecture (feedforward architecture, recurrent architecture, etc.). In one example, after a convolutional layer, other layers such as a pooling layer, a fully connected layer, and / or a normalization layer follow.

[0030] The dimensionality of data can be reduced by using a pooling layer to combine the outputs from multiple nodes in one layer into a single node in the next layer. The pooling operation for a pooling layer with a feature map as input is described below. This description can be appropriately applied to other input signals as well. The feature map can be divided into sub-regions (e.g., rectangular sub-regions), and the features within each sub-region can be independently downsampled (or pooled) to a single value by taking, for example, the average value in the case of average pooling or the maximum value in the case of max pooling.

[0031] The pooling layer can perform pooling such as local pooling, global pooling, max pooling, and / or average pooling. Pooling is a form of non-linear downsampling. Local pooling combines a small number of nodes (e.g., a local cluster of nodes such as 2×2 nodes) within the feature map. In global pooling, for example, all nodes of the feature map can be combined.

[0032] The pooling layer can reduce the size of the representation, thus reducing the number of parameters, the memory footprint, and the computational complexity in a CNN. In one example, a pooling layer is inserted between consecutive convolutional layers of a CNN. In one example, an activation function such as a rectified linear unit (ReLU) layer follows the pooling layer. In one example, the pooling layer is omitted between consecutive convolutional layers of a CNN.

[0033] The normalization layer can be, for example, ReLU, leaky ReLU, generalized divisive normalization (GDN), or inverse GDN (IGDN). ReLU can remove negative values from an input signal such as a feature map by applying a non-saturating activation function to set negative values to zero. Leaky ReLU can have a small slope (e.g., 0.01) instead of a flat slope (e.g., 0) for negative values. Thus, when the value x is greater than 0, the output from leaky ReLU is then x. Otherwise, the output from leaky ReLU is the value x multiplied by a small slope (e.g., 0.01). In one example, since the slope is determined before training, it is not learned during training.

[0034] The NIC framework can be compatible with a compression model for image compression. The NIC framework receives an input image x and outputs a reconstructed image x- corresponding to the input image x (where x- represents

Number

Number

[0035] In some examples, the NIC framework can use a variational autoencoder (VAE) structure. In the VAE structure, the entire input image x can be input into a neural network encoder. The entire input image x can pass through a set of neural network layers that function as a black box (of the neural network encoder) to calculate a compressed representation x^. The compressed representation x^ is the output of the neural network encoder. The neural network decoder can receive the entire compressed representation x^ as input. The compressed representation x^ can pass through another set of neural network layers that function as another black box (of the neural network decoder) to calculate a reconstructed image x-. The rate distortion (R-D) loss

Number

Number

Number

[0036] A neural network (e.g., an ANN) can learn how to perform a task from examples without performing task-specific programming. An ANN can be composed of connected nodes or artificial neurons. Signals can be transmitted from a first node to a second node (e.g., a receiving node) through connections between the nodes, and the signals can be modified by weights that can be represented by connection weight coefficients. The receiving node can process the signals from the nodes (i.e., the input signals of the receiving node), transmit this signal to the receiving node, and then generate an output signal by applying a function to the input signal. The function can be a linear function. In one example, the output signal is the weighted sum of the input signals. In one example, the output signal is further modified by a bias represented by a bias term, and thus, the output signal is the sum of the bias and the weighted sum of the input signals. This function can include, for example, a non-linear operation on the weighted sum, or the sum of the bias and the weighted sum of the input signals. The output signal can be transmitted to a node (a downstream node connected to the receiving node). An ANN can be represented or configured by parameters (e.g., connection weights and / or biases). The weights and / or biases can be obtained by training (e.g., offline training, online training, etc.) the ANN using an example where the weights and / or biases can be adjusted iteratively. A trained ANN configured using the determined weights and / or determined biases can be used to perform a task.

[0037] FIG. 1 shows a NIC framework (100) (e.g., a NIC system) in some examples. The NIC framework (100) can be based on a neural network such as a DNN and / or a CNN. Using the NIC framework (100), an image can be compressed (e.g., encoded), and the compressed image (e.g., the encoded image) can be decompressed (e.g., decoded or reconstructed).

[0038] Specifically, in the example of FIG. 1, the compression model within the NIC framework (100) includes two levels, namely the main level of the compression model and the hyper level of the compression model. The main level of the compression model and the hyper level of the compression model can be implemented using neural networks. In FIG. 1, the neural network at the main level of the compression model is shown as the first sub-NN (151), and the hyper level of the compression model is shown as the second sub-NN (152).

[0039] The first sub-NN (151) can be similar to an autoencoder and can be trained to generate a compressed image x^ of the input image x and then decompress the compressed image (i.e., the encoded image) x^ to obtain a reconstructed image x-. The first sub-NN (151) can include a plurality of components (or modules) such as a main encoder neural network (or main encoder network) (111), a quantizer (112), an entropy encoder (113), an entropy decoder (114), and a main decoder neural network (or main encoder network) (115).

[0040] Referring to FIG. 1, the main encoder network (111) can generate a latent or latent representation y from the input image x (e.g., the image to be compressed or encoded). In one example, the main encoder network (111) is implemented using a CNN. The relationship between the latent representation y and the input image x can be described using Equation 2.

Equation

[0041] The latent representation y can be quantized using a quantizer (112) to generate a quantized latent y^ (where y^ is

Number

[0042] The encoded image (131) can be decompressed (e.g., entropy decoded) by an entropy decoder (114) to generate an output. The entropy decoder (114) can use entropy coding techniques such as Huffman coding or arithmetic coding that correspond to the entropy coding technique used by the entropy encoder (113). In one example, the entropy decoder (114) uses arithmetic decoding and is an arithmetic decoder. In one example, reversible compression is used by the entropy encoder (113) and reversible decompression is used by the entropy decoder (114), and noise, etc. due to the transmission of the encoded image (131) can be omitted, and the output from the entropy decoder (114) is the quantized latent y^.

[0043] The main decoder network (115) can decode the quantized latent y^ to generate a reconstructed image x-. In one example, the main decoder network (115) is implemented using a CNN. The relationship between the reconstructed image x- (i.e., the output of the main decoder network (115)) and the quantized latent y^ (i.e., the input of the main decoder network (115)) can be described using Equation 3.

Number

[0044] In some examples, the second sub-NN (152) can learn an entropy model (e.g., a previous probability model) for the quantized latent y^ used for entropy encoding. Thus, the entropy model can be a conditional entropy model that depends on the input image x, such as a Gaussian mixture model (GMM) or a Gaussian scale model (GSM).

[0045] In some examples, the second sub-NN (152) can include a context model NN (116), an entropy parameter NN (117), a hyper encoder network (121), a quantizer (122), an entropy encoder (123), an entropy decoder (124), and a hyper decoder network (125). The entropy model used in the context model NN (116) can be a potential autoregressive model (e.g., quantized potential ŷ). In one example, the hyper encoder network (121), the quantizer (122), the entropy encoder (123), the entropy decoder (124), and the hyper decoder network (125) form a hyperprior model that can be implemented using a hyper-level neural network (e.g., a hyper prior NN). The hyperprior model can represent information useful for correcting context-based predictions. Data from the context model NN (116) and the hyperprior model can be combined by the entropy parameter NN (117). The entropy parameter NN (117) can generate parameters such as the mean parameter and the scale parameter of an entropy model such as a conditional Gaussian entropy model (e.g., GMM).

[0046] Referring to FIG. 1, on the encoder side, the quantized potential ŷ from the quantizer (112) is supplied to the context model NN (116). On the decoder side, the quantized potential ŷ from the entropy decoder (114) is supplied to the context model NN (116). The context model NN (116) can be implemented using a neural network such as a CNN. The context model NN (116) outputs an output o <i based on the context ŷ cm,i which is the quantized potential ŷ available to the context model NN (116). <ican include the latent that was previously quantized on the encoder side, or the quantized latent that was previously entropy decoded on the decoder side. The output o of the context model NN(116) cm,i and the input (e.g., ŷ <i ) can be described using Equation 4.

Equation

[0047] The output o from the context model NN(116) cm,i and the output o from the hyper-decoder network (125) hc are supplied to the entropy parameter NN(117) to generate the output o ep . The entropy parameter NN(117) can be implemented using a neural network such as a CNN. The relationship between the output o ep of the entropy parameter NN(117) and the input (e.g., o cm,i and o hc ) can be described using Equation 5.

Equation

[0048] The second sub-NN (152) can be described as follows. The latent y can be supplied to the hyper-encoder network (121) to generate the hyper-latent z. In one example, the hyper-encoder network (121) is implemented using a neural network such as a CNN. The relationship between the hyper-latent z and the latent y can be described using Equation 6.

Equation

[0049] The hyper-latent z is quantized by the quantizer (122) to generate the quantized latent z^ (where z^ represents

Equation

[0050] Side information such as the symbolized bits (132) can be decompressed (e.g., entropy decoded) by the entropy decoder (124) to generate an output. The entropy decoder (124) can use entropy coding techniques such as Huffman coding and arithmetic coding. In one example, the entropy decoder (124) uses arithmetic decoding and is an arithmetic decoder. In one example, reversible compression is used in the entropy encoder (123), reversible decompression is used in the entropy decoder (124), noise due to transmission of side information etc. is omitted, and the output from the entropy decoder (124) can be the quantized latent z^. The hyper decoder network (125) can decode the quantized latent z^ to generate an output o hc can be generated. The output o hc The relationship between and the quantized latent z^ can be described using Equation 7.

Equation

[0051] As described above, the compressed bits or symbolized bits (132) can be added to the bitstream coded as side information, thereby enabling the entropy decoder (114) to use a conditional entropy model. In this way, the entropy model depends on the image and can be spatially adapted, and thus can be more accurate than a fixed entropy model.

[0052] The NIC framework (100) can be appropriately adapted, for example, to omit one or more components shown in FIG. 1, modify one or more components shown in FIG. 1, and / or include one or more components not shown in FIG. 1. In one example, the NIC framework that uses a fixed entropy model includes the first sub-NN (151) and does not include the second sub-NN (152). In one example, the NIC framework includes components within the NIC framework (100) excluding the entropy encoder (123) and the entropy decoder (124).

[0053] In one embodiment, one or more components within the NIC framework (100) shown in FIG. 1 are implemented using a neural network such as a CNN. Each NN-based component (within the NIC framework (e.g., the NIC framework (100)), such as the main encoder network (111), the main decoder network (115), the context model NN (116), the entropy parameter NN (117), the hyper encoder network (121), or the hyper decoder network (125)) can include any suitable architecture (e.g., having any suitable combination of layers), can include any suitable type of parameters (e.g., weights, biases, and / or a combination of weights and biases, etc.), and can include any suitable number of parameters.

[0054] In one embodiment, the main encoder network (111), the main decoder network (115), the context model NN (116), the entropy parameter NN (117), the hyper encoder network (121), and the hyper decoder network (125) are each implemented using a respective CNN.

[0055] Figure 2 shows an exemplary CNN of the main encoder network (111) according to an embodiment of the present disclosure. For example, the main encoder network (111) includes four sets of layers, and each set of layers includes a convolutional layer 5×5 c192 s2 followed by a subsequent GDN layer. One or more of the layers shown in Figure 2 may be modified and / or omitted. Additional layers can be added to the main encoder network (111).

[0056] Figure 3 shows an exemplary CNN of the main decoder network (115) according to an embodiment of the present disclosure. For example, the main decoder network (115) includes three sets of layers, and each set of layers includes a transposed convolutional layer 5×5 c192 s2 followed by a subsequent IGDN layer. Further, after the three sets of layers, a transposed convolutional layer 5×5 c3 s2 is followed by an IGDN layer. One or more of the layers shown in Figure 3 may be modified and / or omitted. Additional layers can be added to the main decoder network (115).

[0057] Figure 4 shows an exemplary CNN of the hyper encoder network (121) according to an embodiment of the present disclosure. For example, the hyper encoder network (121) includes a convolutional layer 3×3 c192 s1 followed by a subsequent leaky ReLU, a convolutional layer 5×5 c192 s2 followed by a subsequent leaky ReLU, and a convolutional layer 5×5 c192 s2. One or more of the layers shown in Figure 4 may be modified and / or omitted. Additional layers can be added to the hyper encoder network (121).

[0058] FIG. 5 shows an exemplary CNN of a hyper decoder network (125) according to an embodiment of the present disclosure. For example, the hyper decoder network (125) includes a transposed convolutional layer 5×5 c192 s2 followed by a leaky ReLU, a transposed convolutional layer 5×5 c288 s2 followed by a leaky ReLU, and a transposed convolutional layer 3×3 c384 s1. One or more of the layers 5 shown in FIG. 5 may be modified and / or omitted. Additional layers can be added to the hyper decoder network (125).

[0059] FIG. 6 shows an exemplary CNN of a context model NN (116) according to an embodiment of the present disclosure. For example, the context model NN (116) includes a masked convolution 5×5 c384 s1 for context prediction, and thus, the context y^ of Equation 4 <i includes a restricted context (e.g., a 5×5 convolutional kernel). The convolutional layer in FIG. 6 is modifiable. Additional layers can be added to the context model NN (1016).

[0060] FIG. 7 shows an exemplary CNN of an entropy parameter NN (117) according to an embodiment of the present disclosure. For example, the entropy parameter NN (117) includes a convolutional layer 1×1 c640 s1 followed by a leaky ReLU, a convolutional layer 1×1 c512 s1 followed by a leaky ReLU, and a convolutional layer 1×1 c384 s1. One or more of the layers shown in FIG. 7 may be modified and / or omitted. Additional layers can be added to the entropy parameter NN (117).

[0061] As described with reference to FIGS. 2 to 7, the NIC framework (100) can be implemented using a CNN. The NIC framework (100) can be appropriately adapted such that one or more components within the NIC framework (100) (e.g., (111), (115), (116), (117), (121), and / or (125)) are implemented using any suitable type of neural network (e.g., a CNN or a non-CNN-based neural network). One or more other components of the NIC framework (100) can be implemented using a neural network.

[0062] The NIC framework (100) including a neural network (e.g., a CNN) can be trained to learn the parameters used in the neural network. For example, when a CNN is used, the weights and biases used in the convolutional kernels of the main encoder network (111) (if biases are used in the main encoder network (111)), the weights and biases used in the convolutional kernels of the main decoder network (115) (if biases are used in the main decoder network (115)), the weights and biases used in the convolutional kernels of the hyper encoder network (121) (if biases are used in the hyper encoder network (121)), the weights and biases used in the convolutional kernels of the hyper decoder network (125) (if biases are used in the hyper decoder network (125)), the weights and biases used in the convolutional kernels of the context model NN (116) (if biases are used in the context model NN (116)), and the weights and biases used in the convolutional kernels of the entropy parameter NN (117) (if biases are used in the entropy parameter NN (117)), etc., which are represented by θ1 to θ6, can be learned in the training process (e.g., an offline training process, an online training process, etc.).

[0063] As an example, referring to FIG. 2, the main encoder network (111) includes four convolutional layers, and each convolutional layer has a 5×5 convolutional kernel and 192 channels. Thus, the number of weights used in the convolutional kernel of the main encoder network (111) is 19,200 (i.e., 4×5×5×192). The parameters used in the main encoder network (111) include 19,200 weights and optional biases. When biases and / or additional NNs are used in the main encoder network (111), additional parameters can be included.

[0064] Referring to FIG. 1, the NIC framework (100) includes at least one component or module built on a neural network. The at least one component can include one or more of a main encoder network (111), a main decoder network (115), a hyper encoder network (121), a hyper decoder network (125), a context model NN (116), and an entropy parameter NN (117). The at least one component can be trained individually. In one example, the training process is used to learn the parameters of each component individually. The at least one component can be trained jointly as a group. In one example, the training process is used to jointly learn the parameters of a subset of the at least one component. In one example, the training process is used to learn all the parameters of the at least one component, and thus is called end-to-end optimization.

[0065] In the training process of one or more components in the NIC framework (100), the weights (or weight coefficients) of the one or more components can be initialized. In one example, the weights are initialized based on a corresponding pre-trained neural network model (e.g., DNN model, CNN model). In one example, the weights are initialized by setting them to random numbers.

[0066] For example, after initializing the weights, one or more components can be trained using a set of training images. The set of training images can include any suitable images having any suitable size. In some examples, the set of training images includes images from raw images, natural images, and / or computer-generated images within a spatial region. In some examples, the set of training images includes images from residual images, or residual images having residual data in a spatial region. The residual data can be calculated by a residual calculator. In some examples, the raw images and / or the residual images containing residual data can be directly used to train a neural network in an NIC framework such as the NIC framework (100). Thus, a neural network can be trained in the NIC framework using raw images, residual images, images from raw images, and / or images from residual images.

[0067] For the sake of brevity, the following training processes (e.g., offline training process, online training process, etc.) will be described using training images as an example. This description can be appropriately adapted to the training block. The training image t in the set of training images can generate a compressed representation (e.g., encoded information into a bitstream) through the encoding process of FIG. 1. The encoded information can be used to calculate and reconstruct the reconstructed image t- through the decoding process described in FIG. 1 (where the reconstructed image t- is

Number

[0068] In the NIC framework (100), a balance is taken between two competing targets, for example, the reconstructed quality and the bit consumption. A quality loss function (e.g., distortion or distortion loss)

Number

[0069] In the case of neural image compression, a differentiable approximation of quantization can be used in end-to-end optimization. In various examples, in the training process of neural network-based image compression, noise injection is used to simulate quantization, and thus, quantization is simulated by noise injection instead of being performed by a quantizer (e.g., quantizer (112)). Thus, training with noise injection can variably approximate the quantization error. Since an entropy coder can be simulated using a bits per pixel (BPP) estimator, entropy coding is simulated by the BPP estimator instead of being performed by an entropy encoder (e.g., (113)) and an entropy decoder (e.g., (114)). Therefore, the rate loss R in the loss function L shown in Equation 1 during the training process can be estimated based on, for example, noise injection and the BPP estimator. Generally, increasing the rate R can reduce the distortion D, and decreasing the rate R can increase the distortion D. Thus, the combined R-D loss L can be optimized using the trade-off hyperparameter λ of Equation 1, where L can be optimized as the sum of λD and R. The training process can be used to adjust the parameters of one or more components (e.g., (111)(115)) within the NIC framework (100) such that the combined R-D loss L is minimized or optimized. In some examples, the trade-off hyperparameter λ can be used to optimize the combined rate distortion (R-D) loss as follows.

Number

[0070] Using various models, the distortion loss D and the rate loss R can be determined, and thus the combined R-D loss L of Equation 1 can be determined. In one example, the distortion loss

Number

[0071] In one example, the target of the training process is to train an encoding neural network (e.g., an encoding DNN) such as a video encoder used on the encoder side, and a decoding neural network (e.g., a decoding DNN) such as a video decoder used on the decoder side. As an example, referring to FIG. 1, the encoding neural network can include a main encoder network (111), a hyper encoder network (121), a hyper decoder network (125), a context model NN (116), and an entropy parameter NN (117). The decoding neural network can include a main decoder network (115), a hyper decoder network (125), a context model NN (116), and an entropy parameter NN (117). The video encoder and / or the video decoder can include other components, whether NN-based and / or non-NN-based.

[0072] The NIC framework (e.g., NIC framework (100)) can be trained in an end-to-end manner. In one example, the encoding neural network and the decoding neural network are updated together in the training process based on the gradients backpropagated in an end-to-end manner, for example, using the gradient descent algorithm. The gradient descent algorithm can repeatedly optimize the parameters of the NIC framework to find the minimum value of the differentiable function of the NIC framework (e.g., the minimum value of the rate distortion loss). For example, the gradient descent algorithm can repeatedly take steps in the opposite direction of the gradient (or approximate gradient) of the differentiable function at the current point.

[0073] After training the parameters of the neural network in the NIC framework (100), one or more components in the NIC framework (100) can be used to encode and / or decode an image. In one embodiment, on the encoder side, the image encoder is configured to encode the input image x into an encoded image (131) transmitted as a bitstream. The image encoder can include a plurality of components within the NIC framework (100). In one embodiment, on the decoder side, the corresponding image decoder is configured to decode the encoded image (131) conveyed as a bitstream into a reconstructed image x-. The image decoder can include a plurality of components within the NIC framework (100).

[0074] Note that the image encoder and the image decoder by the NIC framework can have corresponding structures.

[0075] FIG. 8 shows an exemplary image encoder (800) according to an embodiment of the present disclosure. The image encoder (800) includes a main encoder network (811), a quantizer (812), an entropy encoder (813), and a second sub-NN (852). The main encoder network (811) is configured similarly to the main encoder network (111), the quantizer (812) is configured similarly to the quantizer (112), the entropy encoder (813) is configured similarly to the entropy encoder (113), and the second sub-NN (852) is configured similarly to the second sub-NN (152). This has been described above with reference to FIG. 1 and is omitted here for clarity.

[0076] FIG. 9 shows an exemplary image decoder (900) according to an embodiment of the present disclosure. The image decoder (900) can correspond to the image encoder (800). The image decoder (900) can include a main decoder network (915), an entropy decoder (914), a context model NN (916), an entropy parameter NN (917), an entropy decoder (924), and a hyper decoder network (925). The main decoder network (915) is configured similarly to the main decoder network (115), the entropy decoder (914) is configured similarly to the entropy decoder (114), the context model NN (916) is configured similarly to the context model NN (116), the entropy parameter NN (917) is configured similarly to the entropy parameter NN (117), the entropy decoder (924) is configured similarly to the entropy decoder (124), and the hyper decoder network (925) is configured similarly to the hyper decoder network (125). This has been described above with reference to FIG. 1 and is omitted here for clarity.

[0077] Referring to FIGS. 8 to 9, on the encoder side, the image encoder (800) can generate an encoded image (831) and encoded bits (832) transmitted in a bitstream. On the decoder side, the image decoder (900) can receive and decode the encoded image (931) and encoded bits (932). The encoded image (931) and encoded bits (932) can be parsed from the received bitstream.

[0078] FIGS. 10 to 11 respectively show an exemplary image encoder (1000) and a corresponding image decoder (1100) according to an embodiment of the present disclosure. Referring to FIG. 10, the image encoder (1000) includes a main encoder network (1011), a quantizer (1012), and an entropy encoder (1013). The main encoder network (1011) is configured in the same manner as the main encoder network (111), the quantizer (1012) is configured in the same manner as the quantizer (112), and the entropy encoder (1013) is configured in the same manner as the entropy encoder (113). This has been described above with reference to FIG. 1 and is omitted here for clarity.

[0079] Referring to FIG. 11, the image decoder (1100) includes a main decoder network (1115) and an entropy decoder (1114). The main decoder network (1115) is configured in the same manner as the main decoder network (115), and the entropy decoder (1114) is configured in the same manner as the entropy decoder (114). This has been described above with reference to FIG. 1 and is omitted here for clarity.

[0080] Referring to FIGS. 10 and 11, the image encoder (1000) can generate an encoded image (1031) included in a bitstream. The image decoder (1100) can receive the bitstream and decode the encoded image (1131) conveyed in the bitstream.

[0081] According to some aspects of the present disclosure, image compression can remove the redundancy of an image, and thus can significantly reduce the number of bits that can be used to represent the compressed image. Image compression provides advantages for image transmission and storage. The compressed image can be referred to as an image in a compressed region. Image processing on the compressed image can be referred to as image processing in the compressed region. In some examples, image compression can be performed at different compression rates, and in some examples, the compressed region can be referred to as a multi-rate compressed region.

[0082] Computer vision (CV) is a field of artificial intelligence (AI) that uses a computer equipped with a neural network to detect, understand, and process the content of an image. CV tasks can include, but are not limited to, image classification, object detection, super-resolution (generating a high-resolution image from one or more low-resolution images), and image noise removal. In some related examples, CV tasks are performed on the original uncompressed image and the decompressed image such as the reconstructed image from the compressed image. In some examples, the compressed image is decompressed to generate a reconstructed image, and the CV task is performed on the reconstructed image. Reconstruction may require a large amount of computation. Performing the CV task in the compressed region without reconstructing the image reduces the computational complexity and shortens the waiting time of the CV task.

[0083] Some aspects of the present disclosure provide techniques for multi-rate computer vision task neural networks in a compression domain. In some examples, this technique can be used in an end-to-end (E2E) optimization framework that includes a model of a multi-rate compression domain CV task framework (CDCVTF). The E2E optimization framework includes an encoder and a decoder. The encoder can generate a coded bitstream of the input image, and the decoder can decode the coded bitstream to generate a CV task-based result. Both the encoder and the decoder support multi-rates of image compression. The end-to-end (E2E) optimization framework can be a pre-trained artificial neural network (ANN)-based framework.

[0084] In some examples, the CV task for a compressed image can be performed by two parts of a neural network. In some examples, the two parts of the neural network form an E2E framework that can be trained end-to-end. The two parts of the neural network include a first part of an image coding neural network and a second part of a CV task neural network. The image coding neural network is also called an image compression encoder. The CV task neural network is also called a CV task decoder. The image compression encoder can encode an image into a coded bitstream. The CV task decoder can decode the coded bitstream to generate a result of the CV task in the compression domain. The CV task decoder performs the CV task based on the compressed image.

[0085] In some examples, an encoder of a NIC framework compresses an image to generate a coded bitstream that conveys a compressed image or a compressed feature map. Further, the compressed image or the compressed feature map can be provided to a CV task neural network to generate a result of the CV task.

[0086] FIG. 12 shows, in some examples, a system (1200) for performing a CV task in a compressed domain. The system (1200) includes an image compression encoder (1220), an image compression decoder (1250), and a CV task decoder (1270). The image compression decoder (1250) can correspond to the image compression encoder (1220). For example, the image compression encoder (1220) can be configured similar to the image encoder (800), and the image compression decoder (1250) can be configured similar to the image decoder (900). In another example, the image compression encoder (1220) can be configured similar to the image encoder (1000), and the image compression decoder (1250) can be configured similar to the image decoder (1100). In some examples, the image compression encoder (1220) is configured as an encoding part (having an encoder model) of a NIC framework, and the image compression decoder (1250) is configured as a decoding part (having a decoder model) of the NIC framework. The NIC framework can be trained end-to-end to determine pre-trained parameters of an encoder model and a decoder model of an E2E optimization framework. The image compression encoder (1220) and the image compression decoder (1250) are configured according to the encoder model and the decoder model having the pre-trained parameters. The image compression encoder (1220) can receive an input image, compress the input image, and generate a coded bitstream that conveys a compressed image corresponding to the input image. The image compression decoder (1250) can receive the coded bitstream, decompress the compressed image, and generate a reconstructed image.

[0087] The CV task decoder (1270) is configured to decode a coded bitstream that conveys a compressed image and generate a CV task result corresponding to the input image. The CV task decoder (1270) may be a single-task decoder or a multi-task decoder. The CV tasks include, but are not limited to, super-resolution, object detection, image noise removal, and image classification, etc.

[0088] The CV task decoder (1270) includes a neural network (e.g., a CV task decoder model) trained based on training data. For example, the training data can include training images, compressed training images (e.g., by the image compression encoder (1220)), and guideline CV task results of the training images. For example, the CV task decoder (1270) can receive the compressed training images as input and generate the results of the training CV tasks. Next, the CV task decoder (1270) (e.g., having an adjustable neural network structure and adjustable parameters) is trained to minimize the loss between the guideline CV task results and the training CV task results. Through training, the structure and pre-trained parameters of the CV task decoder (1270) are determined.

[0089] It should be noted that the CV task decoder (1270) can have any suitable neural network structure to generate the CV task results. In some embodiments, the CV task decoder (1270) decodes the coded bitstream and directly generates the CV task results without performing image reconstruction. In some embodiments, the CV task decoder (1270) first decodes the coded bitstream to generate a decompressed image (also called a reconstructed image), and then applies the CV task model to the decompressed image to generate the CV task results.

[0090] In some examples, the image compression encoder (1220), the image compression decoder (1250), and the CV task decoder (1270) are in different electronic devices. For example, the image compression encoder (1220) is in a first device (1210), the image compression decoder (1250) is in a second device (1240), and the CV task decoder (1270) is in a third device (1260). In some examples, the image compression decoder (1250) and the CV task decoder (1270) may be in the same device. In some examples, the image compression encoder (1220) and the image compression decoder (1250) may be in the same device. In some examples, the image compression encoder (1220) and the CV task decoder (1270) may be in the same device. In some examples, the image compression encoder (1220), the image compression decoder (1250), and the CV task decoder (1270) may be in the same device.

[0091] FIG. 13 shows, in some examples, a system (1300) for performing CV tasks in a compression region. The system (1300) includes an image compression encoder (1320) and a CV task decoder (1370). In some examples, the image compression encoder (1320) is configured as an encoding part of the NIC framework (having an encoder model), and the CV task decoder (1370) is configured as a decoding part of the NIC framework (having a decoder model). The NIC framework can be trained end-to-end based on training data to determine pre-trained parameters of the encoder model and the decoder model of the E2E optimization framework. For example, the training data may include training images (uncompressed), and the results of the CV tasks of the guidelines for the training images. For example, the NIC framework can receive a training image as input and generate the results of the training CV tasks. Next, the NIC framework is trained (using an adjustable neural network structure and adjustable parameters), and the loss between the results of the guideline CV tasks and the results of the training CV tasks is minimized. The image compression encoder (1320) and the CV task decoder (1370) are configured according to the pre-trained parameters.

[0092] Next, the image compression encoder (1320) can receive an input image and generate a coded bitstream that conveys the compressed image or the compressed feature map. The CV task decoder (1370) can receive the coded bitstream, decompress the compressed image or the compressed feature map, and generate the CV task results.

[0093] The CV task decoder (1370) may be a single-task decoder or a multi-task decoder. The CV tasks may include, but are not limited to, super-resolution, object detection, image noise removal, image classification, etc.

[0094] Note that the CV task decoder (1370) can use any suitable neural network structure. In some embodiments, the CV task decoder (1370) decodes the coded bitstream to directly generate the CV task result without performing image reconstruction. In some embodiments, the CV task decoder (1370) first decodes the coded bitstream to generate a decompressed image (also referred to as a reconstructed image), and then applies the CV task model to the decompressed image to generate the CV task result.

[0095] In some examples, the image compression encoder (1320) and the CV task decoder (1370) are in different electronic devices. For example, the image compression encoder (1320) is in the first device (1310), and the CV task decoder (1370) is in the second device (1360). In some examples, the image compression encoder (1320) and the CV task decoder (1370) may be in the same device.

[0096] According to some aspects of the present disclosure, the compression regions in the systems (1200) and (1300) can be appropriately adjusted for multirate compression. In some examples, the hyperparameter λ is used to adjust the compression rate. In some examples, the compression rate is defined as the number of bits per pixel and can be calculated according to Equation 9. Compression rate = R / (W × H) Equation 9 Here, R is the bit consumption of the compressed image, which is also used in Equation 1, W is the width of the input image, and H is the height of the input image. According to Equations 1 and 9, as the hyperparameter λ increases, the compression rate also increases.

[0097] Some aspects of the present disclosure provide a multi-rate neural network model in a compressed region CV task framework (CDCVTF). In some examples, since the hyperparameter λ is an input to the CDCVTF, the CDCVTF is trained to understand the compression rate and then the value of the hyperparameter λ can be used to adjust the compression rate of the CDCVTF model. The hyperparameter λ is used in the following description to explain the multi-rate technique in the CDCVTF, and it should be noted that the technique can be appropriately adjusted to use other suitable parameters that can adjust the compression rate.

[0098] FIG. 14 shows, in some examples, a system (1400) for performing a CV task in a multi-rate compression region. The system (1400) includes a multi-rate image compression encoder (1411), a multi-rate image compression decoder (1441), and a multi-rate CV task decoder (1461). In some examples, the multi-rate image compression decoder (1441) can correspond to the multi-rate image compression encoder (1411). In some examples, the multi-rate image compression encoder (1411) is configured as an encoding part (having an encoder model) of a multi-rate NIC framework, and the multi-rate image compression decoder (1441) is configured as a decoding part (having a decoder model) of the multi-rate NIC framework. The multi-rate NIC framework can be trained end-to-end (e.g., E2E multi-rate NIC training) to determine pre-trained parameters of the encoder model and the decoder model of an E2E optimization framework for multi-rate image compression.

[0099] In the example of FIG. 14, the multi-rate image compression encoder (1411) includes a conversion module (1430) and an image compression encoder (1420). The conversion module (1430) includes a neural network that converts the hyperparameter λ into a tensor (1431). The image compression encoder (1420) can compress the input image based on the tensor (1431) to generate a coded bitstream that conveys the compressed image corresponding to the input image. In some examples, the neural network in the conversion module (1430) includes an adjustable structure or adjustable parameters that can be adjusted during E2E multi-rate NIC training to determine pre-trained parameters.

[0100] In the example of FIG. 14, the multi-rate image compression decoder (1441) includes a conversion module (1445) and an image compression decoder (1450). The conversion module (1445) includes a neural network that converts the hyperparameter λ into a tensor (1446). The image compression decoder (1450) can decompress the compressed image based on the tensor (1446) to generate a reconstructed image. In some examples, the neural network in the conversion module (1445) includes an adjustable structure or adjustable parameters that can be adjusted during E2E multi-rate NIC training to determine pre-trained parameters.

[0101] In the example of FIG. 14, the multi-rate CV task decoder (1461) includes a conversion module (1480) and a CV task decoder (1470). The conversion module (1480) includes a neural network that converts the hyperparameter λ into a tensor (1481). The CV task decoder (1470) can decode the compressed image based on the tensor (1481) to generate a CV task result. In some examples, the neural network in the conversion module (1480) includes an adjustable structure or adjustable parameters that can be adjusted during training (referred to as CV task training) to determine pre-trained parameters.

[0102] The multi-rate CV task decoder (1461) may be a single-task decoder or a multi-task decoder. The CV tasks may include, but are not limited to, super-resolution, object detection, image noise removal, image classification, etc.

[0103] The multi-rate CV task decoder (1461) can be trained in CV task training based on training data. For example, the training data may include training images, compressed training images with corresponding rates (e.g., based on the multi-rate image compression decoder (1441)), and guideline CV task results of the training images. For example, the multi-rate CV task decoder (1461) can receive a compressed training image with a corresponding rate as input and generate the result of the training CV task. Next, the multi-rate CV task decoder (1461) is trained (using an adjustable neural network structure and adjustable parameters) to minimize the loss between the guideline CV task result and the result of the training CV task.

[0104] In some embodiments, the multi-rate CV task decoder (1461) decodes the coded bitstream and directly generates the CV task result without performing image reconstruction. In some embodiments, the multi-rate CV task decoder (1461) first decodes the coded bitstream, generates the decompressed image, and then applies the CV task model to the decompressed image to generate the CV task result.

[0105] In some examples, the multi-rate image compression encoder (1411), the multi-rate image compression decoder (1441), and the multi-rate CV task decoder (1461) are in different electronic devices. For example, the multi-rate image compression encoder (1411) is in a first device, the multi-rate image compression decoder (1441) is in a second device, and the multi-rate CV task decoder (1461) is in a third device. In some examples, the multi-rate image compression decoder (1441) and the multi-rate CV task decoder (1461) may be in the same device. In some examples, the multi-rate image compression encoder (1411) and the multi-rate image compression decoder (1441) may be in the same device. In some examples, the multi-rate image compression encoder (1411) and the multi-rate CV task decoder (1461) may be in the same device. In some examples, the multi-rate image compression encoder (1411), the multi-rate image compression decoder (1441), and the multi-rate CV task decoder (1461) may be in the same device.

[0106] In some examples, the conversion module (1430), the conversion module (1445), and the conversion module (1480) can have the same neural network structure and the same pre-trained parameters. In some examples, the conversion module (1430), the conversion module (1445), and the conversion module (1480) can have the same neural network structure but different pre-trained parameters. In some examples, the conversion module (1430), the conversion module (1445), and the conversion module (1480) can have different neural network structures.

[0107] Note that the tensor generated according to the value of the hyperparameter λ can be provided to any suitable layer within the neural network. For example, the tensor (1481) generated according to the value of the hyperparameter λ can be provided to one or more layers within the neural network of the CV task decoder (1470).

[0108] FIG. 15 shows, in some examples, a system (1500) for performing a CV task in a multirate compression domain. The system (1500) includes a multirate image compression encoder (1511) and a multirate CV task decoder (1561). In some examples, the multirate image compression encoder (1611) is configured as an encoding part of a multirate NIC framework (having an encoder model), and the multirate CV task decoder (1561) is configured as a decoding part of a multirate NIC framework (having a decoder model). The multirate NIC framework can be trained end-to-end (e.g., E2E multirate NIC training) to determine pre-trained parameters of an encoder model and a decoder model of an E2E optimization framework for a multirate CV task.

[0109] In the example of FIG. 15, the multirate image compression encoder (1511) includes a conversion module (1530) and an image compression encoder (1520). The conversion module (1530) includes a neural network that converts a hyperparameter λ into a tensor (1531). The image compression encoder (1520) can compress an input image based on the tensor (1531) to generate a coded bitstream that conveys a compressed image or a compressed feature map corresponding to the input image. In some examples, the neural network within the conversion module (1530) includes an adjustable structure or adjustable parameters that can be adjusted during E2E multirate NIC training to determine pre-trained parameters.

[0110] In the example of FIG. 15, the multi-rate CV task decoder (1561) includes a conversion module (1580) and a CV task decoder (1570). The conversion module (1580) includes a neural network that converts the hyperparameter λ into a tensor (1581). The CV task decoder (1570) can decode a compressed image or a compressed feature map based on the tensor (1581) to generate a CV task result. In some examples, the neural network in the conversion module (1580) includes an adjustable structure or adjustable parameters that can be adjusted during E2E multi-rate NIC training to determine pre-trained parameters.

[0111] The multi-rate CV task decoder (1561) may be a single-task decoder or a multi-task decoder. The CV tasks include, but are not limited to, super-resolution, object detection, image noise removal, image classification, etc.

[0112] The multi-rate CV task decoder (1561) can be trained using the multi-rate image compression encoder (1511) in E2E multi-rate NIC training based on training data. For example, the training data can include training images, the value of the hyperparameter λ, and the guided CV task results of the training images. For example, the multi-rate image compression encoder (1511) can receive a training image with the value of the hyperparameter λ as input and generate a compressed feature map. The compressed feature map and the value of the hyperparameter λ are input to the multi-rate CV task decoder (1561) to generate the results of the training CV task corresponding to the training image and the value of the hyperparameter λ. The multi-rate CV task decoder (1561) and the multi-rate image compression encoder (1511) are trained (using an adjustable neural network structure and adjustable parameters) to minimize the loss between the guided CV task results and the results of the training CV task.

[0113] In some embodiments, the multi-rate CV task decoder (1561) decodes the coded bitstream and directly generates the CV task result without performing image reconstruction. In some embodiments, the multi-rate CV task decoder (1561) first decodes the coded bitstream, generates the decompressed image, and then applies the CV task model to the decompressed image to generate the CV task result.

[0114] In some examples, the multi-rate image compression encoder (1511) and the multi-rate CV task decoder (1561) are in different electronic devices. In some examples, the multi-rate image compression encoder (1511) and the multi-rate CV task decoder (1561) are in the same electronic device.

[0115] In some examples, the conversion module (1530) and the conversion module (1580) can have the same neural network structure and the same pre-trained parameters. In some examples, the conversion module (1530) and the conversion module (1580) have the same neural network structure but may have different pre-trained parameters. In some examples, the conversion module (1530) and the conversion module (1580) can have different neural network structures.

[0116] Note that the tensor generated by the conversion module according to the value of the hyperparameter λ can be provided to any suitable layer within the neural network. For example, the tensor (1581) generated according to the value of the hyperparameter λ can be provided to one or more layers within the neural network of the CV task decoder (1570).

[0117] In some examples, the value of the parameter that controls the compression rate is signaled within the coded bitstream that conveys the compressed image or compressed feature map. For example, the value of the hyperparameter λ is signaled within the coded bitstream from the encoder side such as the multi-rate image compression encoder (1411) and the multi-rate image compression encoder (1511). Next, the CV task decoders such as the multi-rate CV task decoder (1461) and the multi-rate CV task decoder (1561) can receive the hyperparameter λ information.

[0118] Note that the network architectures such as the neural network structures of the image compression encoder, the CV task decoder, and the conversion module can have any suitable structure. In one embodiment, the conversion module includes a set of convolutional layers. In another embodiment, the conversion module includes convolutional layers having activation functions.

[0119] In some embodiments, the hyperparameter λ is selected from a set of predefined values known to both the encoder side and the decoder side. For the λ value selected on the encoder side, the index of the value within the set is signaled in the coded bitstream. According to the index, the decoder can determine the value of the hyperparameter λ used for encoding and use the same value of the hyperparameter λ as the input to the decoder network.

[0120] In one example, eight values of the hyperparameter λ are predefined and arranged within the set. Both the encoder side and the decoder side have information about the set such as the eight values and the positions of the eight values within the set. Next, instead of transmitting the selected value of the hyperparameter λ on the encoder side, an index indicating the selected value within the set can be transmitted from the encoder side to the decoder set. For example, the index can be from 0 to 7. According to the index, the decoder can determine the value of the hyperparameter λ selected on the encoder side.

[0121] According to one aspect of the present disclosure, compared with a normal CDCVTF, a multi-rate CDCVTF includes a multi-rate function, and parameters such as a hyperparameter λ are inputs to the multi-rate CDCVTF. By changing the value of a parameter such as the value of the hyperparameter λ, the image compression rate can be adjusted.

[0122] FIG. 16 shows a flowchart showing an outline of a process (1600) according to an embodiment of the present disclosure. The process (1600) is an encoding process. The process (1600) can be executed by an electronic device. In some embodiments, the process (1600) is implemented by software instructions, and thus when the processing circuit executes the software instructions, the processing circuit executes the process (1600). The process starts at (S1601) and proceeds to (S1610).

[0123] (In S1610), the value of a parameter for adjusting the compression rate is input to a multi-rate image compression encoder such as a multi-rate image compression encoder (1411) and a multi-rate image compression encoder (1511). The multi-rate image compression encoder includes one or more neural networks for encoding an image using each value of the parameter.

[0124] (In S1620), the multi-rate image compression encoder encodes the input image into a compressed image according to the value of the parameter and conveys it in a coded bitstream. The value of the parameter adjusts the compression rate for encoding the input image into a compressed image.

[0125] (In S1630), an index is encoded into the coded bitstream, and the index refers to a value within a set of values of the parameter value (for example, a predefined value).

[0126] Next, the process (1600) proceeds to (S1699) and ends.

[0127] The process (1600) can be appropriately adapted to various scenarios, and the steps of the process (1600) can be adjusted accordingly. One or more steps in the process (1600) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement the process (1600). Additional steps can be added.

[0128] FIG. 17 shows a flowchart illustrating an overview of a process (1700) according to an embodiment of the present disclosure. The process (1700) is a decoding process. The process (1700) can be executed by an electronic device. In some embodiments, the process (1700) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (1700). The process starts at (S1701) and proceeds to (S1710).

[0129] (S1710), an index indicating a value within a set of parameter values is decoded from the coded bitstream carrying the compressed image. By changing the value of the parameter, the compression rate of the compressed image is adjusted. The compressed image is generated by a neural network-based encoder such as a multi-rate image compression encoder (1411) and a multi-rate image compression encoder (1511) according to the value of the parameter.

[0130] (S1720), the value of the parameter is input to a multi-rate compression region computer vision task decoder such as a multi-rate CV task decoder (1461) and a multi-rate CV task decoder (1561). The multi-rate compression region computer vision task decoder includes one or more neural networks for executing a computer vision task from the compressed image according to the corresponding value of the parameter used for generating the compressed image.

[0131] In (S1730), the multi-rate compression region computer vision task decoder generates a computer vision task result according to the compressed image and the parameter values in the encoded bitstream.

[0132] In some examples, a first neural network (e.g., conversion model (1480), conversion model (1580)) in the multi-rate compression region computer vision task decoder converts the parameter values into a tensor. The tensor is input into one or more layers of a second neural network (e.g., CV task decoder (1470), CV task decoder (1570)) in the multi-rate compression region computer vision task decoder. The second neural network generates a computer vision task result according to the compressed image and the tensor.

[0133] In some examples, the first neural network includes one or more convolutional layers. In some examples, the first neural network includes a convolutional layer with an activation function.

[0134] In some examples, the second neural network is configured to generate a computer vision task result without generating a reconstructed image from the compressed image.

[0135] In some examples, the second neural network is configured to generate a reconstructed image from the compressed image and generate a computer vision task result from the reconstructed image.

[0136] In some examples, the neural network-based encoder is based on the encoder model of the neural image compression (NIC) framework, the multi-rate compression region computer vision task decoder is based on the decoder model of the NIC framework, and the NIC framework is trained end-to-end as described with reference to FIG. 15.

[0137] In some examples, the decoder model of the multi-rate compression region computer vision task decoder is trained separately from the encoder model of the neural network-based encoder, as described with reference to FIG. 14.

[0138] In some examples, the parameter is a hyperparameter for weighting distortion in the calculation of the rate distortion loss, such as the hyperparameter λ of Equation 1.

[0139] Note that the computer vision task can be any suitable computer vision task such as image classification, image noise removal, object detection, and super-resolution.

[0140] Next, the process (1700) proceeds to (S1799) and ends.

[0141] The process (1700) can be appropriately adapted to various scenarios, and the steps of the process (1700) can be adjusted accordingly. One or more steps in the process (1700) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to execute the process (1700). Additional steps can be added.

[0142] The above-described technology can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 18 shows a computer system (1800) suitable for implementing a particular embodiment of the disclosed subject matter.

[0143] Computer software can be coded using any suitable machine code or computer language and, upon undergoing assembly, compilation, linking, or similar mechanisms, can create code that includes instructions that can be directly executed by one or more computer central processing units (CPUs), and graphics processing units (GPUs), etc., or can be executed through interpretation, execution of microcode, etc.

[0144] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0145] The components related to the computer system (1800) shown in FIG. 18 are essentially illustrative and are not intended to imply any limitation regarding the use or function scope of the computer software implementing the embodiments of the present disclosure. Also, the component configuration should not be construed as having dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1800).

[0146] The computer system (1800) can include a specific human interface input device. Such a human interface input device can respond to input by one or more human users, for example, via tactile input (keystrokes, swipes, movement of a data glove, etc.), voice input (voice, applause, etc.), visual input (gestures, etc.), olfactory input (not shown), etc. The human interface device can also be used to capture specific media that is not necessarily directly related to conscious human input, such as voice (voice, music, ambient sound, etc.), images (scanned images, photographic images obtained from a still camera, etc.), videos (2D videos, 3D videos including stereoscopic videos, etc.).

[0147] The input human interface device can include one or more of a keyboard (1801), a mouse (1802), a trackpad (1803), a touch screen (1810), a data glove (not shown), a joystick (1805), a microphone (1806), a scanner (1807), a camera (1808) (only one of each is shown).

[0148] The computer system (1800) can also include certain human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices include tactile output devices (e.g., tactile feedback by a touch screen (1810), a data glove (not shown), or a joystick (1805), and there can also be a tactile feedback device that does not function as an input device), audio output devices (a speaker (1809), headphones (not shown), etc.), visual output devices (screens (1810) including a CRT screen, an LCD screen, a plasma screen, an OLED screen, etc., regardless of the presence or absence of each touch screen input function and regardless of the presence or absence of each tactile feedback function, some of which can output two-dimensional visual output or three-dimensional or higher-dimensional output through means such as stereographic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown) can be included.

[0149] The computer system (1800) can also include a human-accessible storage device, an optical medium including a CD / DVD, a ROM / RW (1820) with a CD / DVD, or a related medium such as a similar medium (1821), a thumb drive (1822), a removable hard drive or a solid state drive (1823), a conventional magnetic medium such as a tape and a floppy disk (not shown), a special ROM / ASIC / PLD-based device such as a security dongle (not shown), etc.

[0150] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0151] The computer system (1800) can also include an interface (1854) to one or more communication networks (1855). The network can be, for example, wireless, wired, optical, etc. Further, the network can be a local, wide area, metropolitan, vehicle and industrial, real-time, delay tolerant, etc. network. Examples of networks include local area networks such as Ethernet, wireless LAN, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial including CANBus, etc. A particular network generally requires an external network interface adapter connected to a particular general-purpose data port or peripheral bus (1849) (e.g., a USB port of the computer system (1800), etc.). Others are generally integrated into the core of the computer system (1800) by connecting to a system bus (e.g., an Ethernet interface to a PC computer system, or a cellular network interface to a smartphone computer system) as described below. Using any of these networks, the computer system (1800) can communicate with other entities. Such communication can be unidirectional, receive-only (such as TV broadcast), transmit-only (from a CANBus to a particular CANBus device, etc.), or bidirectional (such as other computer systems using a local or wide area digital network). Specific protocols and protocol stacks can be used for each of these networks and network interfaces as described above.

[0152] The foregoing human interface device, the memory device accessible by a human, and the network interface can be attached to the core (1840) of the computer system (1800).

[0153] The core (1840) can include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a special programmable processing device in the form of a field programmable gate array (FPGA) (1843), a hardware accelerator (1844) for a specific task, and a graphics adapter (1850), etc. These devices can be connected via a system bus (1848) together with a read-only memory (ROM) (1845), a random access memory (1846), an internal mass storage device such as an internal hard drive inaccessible to the user, an SSD, etc. (1847). In some computer systems, the system bus (1848) can be made accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be directly connected to the system bus (1848) of the core or can be connected via a peripheral bus (1849). In one example, a screen (1810) can be connected to a graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.

[0154] The CPU (1841), GPU (1842), FPGA (1843), and accelerator (1844) can execute specific instructions that can configure the aforementioned computer code in combination. The computer code can be stored in the ROM (1845) or the RAM (1846). Migration data can be stored in the RAM (1846), but persistent data can be stored, for example, in an internal mass storage device (1847). Fast storage and retrieval to / from any memory device is made possible by the use of a cache memory that can be closely related to one or more CPUs (1841), GPUs (1842), mass storage (1847), ROM (1845), and RAM (1846), etc.

[0155] A computer-readable medium can have computer code thereon for performing various operations implemented by a computer. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or may be of the kind well known and available to those having skill in the art of computer software technology.

[0156] By way of example and not limitation, a computer system having an architecture (1800), particularly a core (1840), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, and accelerators, etc.) executing software incorporated in one or more tangible computer-readable media. Such computer-readable media may be media associated with a mass storage device accessible by a user as introduced above, or may be a specific storage device of the core (1840) having a non-transitory nature such as an on-core mass storage device (1847) or a ROM (1845). The software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core (1840). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can cause the core (1840), particularly a processor (including a CPU, GPU, and FPGA, etc.) therein, to execute a specific process or a specific part of a specific process described herein that includes the definition of a data structure stored in the RAM (1846) and modifies such a data structure according to a process defined by the software. Additionally or alternatively, the computer system can provide functionality as a result of logic hardwired or otherwise incorporated in a circuit (e.g., an accelerator (1844)), which can operate instead of or in conjunction with software to execute a specific process or a specific part of a specific process described herein. References to software may include logic as necessary, and vice versa. References to computer-readable media can, as necessary, include a circuit (such as an integrated circuit (IC), etc.) storing software for execution, a circuit embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0157] Although several exemplary embodiments of the present disclosure have been described, there are changes, substitutions, and various alternative equivalents that are included within the scope of the present disclosure. Thus, those skilled in the art will understand that, although not explicitly illustrated or described herein, many systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure can be envisioned.

Claims

1. A method for image processing, the method comprising: decoding, from a coded bitstream carrying a compressed image, an index pointing to a value within a set of parameter values, wherein adjusting the compression rate of the compressed image by changing the values of the parameters, the compressed image being generated by a neural network-based encoder based on the parameters, the parameters being parameters for weighting distortion in the calculation of rate distortion loss; inputting the values of the parameters into a multi-rate compression region computer vision task decoder within a compression region computer vision task framework (CDC-VTF), the multi-rate compression region computer vision task decoder including one or more neural networks for performing a computer vision task from the compressed image according to corresponding values of the parameters used to generate the compressed image at a plurality of different compression rates; the multi-rate compression region computer vision task decoder generating a computer vision task result according to the compressed image in the coded bitstream compressed at a corresponding compression rate from the plurality of different compression rates based on the values of the parameters; the step of generating the computer vision task result includes: a first neural network within the multi-rate compression region computer vision task decoder converting the values of the parameters into a tensor; inputting the tensor into one or more layers of a second neural network within the multi-rate compression region computer vision task decoder; the second neural network generating the computer vision task result according to the compressed image and the tensor; A method.

2. The method according to claim 1, wherein the first neural network includes one or more convolutional layers.

3. The method according to claim 1, wherein the first neural network includes a convolutional layer having an activation function.

4. The method according to claim 1, wherein the second neural network is configured to generate the computer vision task result without generating a reconstructed image from the compressed image. **Claim 5** The method according to claim 1, wherein the second neural network is configured to generate a reconstructed image from the compressed image and generate the computer vision task result from the reconstructed image. **Claim 6** The method according to claim 1, wherein the neural network-based encoder is based on an encoder model in a neural image compression (NIC) framework, the multirate compression region computer vision task decoder is based on a decoder model in the NIC framework, and the NIC framework is trained end-to-end. **Claim 7** The method according to claim 1, wherein the decoder model of the multirate compression region computer vision task decoder is trained separately from the encoder model of the neural network-based encoder. **Claim 8** The method according to claim 1, wherein the computer vision task includes at least one of image classification, image noise removal, object detection, and super-resolution. **Claim 9** A device for image processing, the device including a processing circuit, the processing circuit being configured to perform the method according to any one of claims 1 to 8. Device. **Claim 10** A non-transitory computer-readable storage medium storing a program for image processing, wherein when the program is executed by a processing circuit, the processing circuit is caused to execute the method according to any one of claims 1 to 8. Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Visual fog

    US20200250003A1

  • Image decoding device and image encoding device

    WO2015098713A1

  • Method and apparatus for multi-rate neural image compression with stackable nested model structures

    WO2022035571A1