Methods and apparatuses for image encoding and decoding
The method allows neural network-based image encoding to adjust compression quality flexibly, addressing the need for retraining by using preset values to obtain target gain vectors, thus reducing costs and maintaining image quality.
Patent Information
- Application Number
- JP2024576850
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-06-30
- Publication Date
- 2025-07-03
- Estimated Expiration
- 2042-06-30
AI Technical Summary
Existing neural network-based image and video coding methods require retraining when different compression qualities are needed, leading to high computational costs and increased storage requirements.
A method for image encoding using a neural network that allows encoding at any desired compression quality or bitstream size by obtaining a target gain vector based on first encoding parameters outside the pre-training range, using preset minimum and maximum values to flexibly adjust compression quality without retraining the model.
Enables flexible encoding and decoding of images at any desired quality without retraining, reducing storage and transmission costs while maintaining image quality.
Smart Images

Figure 2025520847000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to the field of encoding and decoding that are database-ized on a neural network architecture. In particular, some embodiments relate to methods and apparatuses for such encoding and decoding of images and / or videos from bitstreams using multiple processing layers.
Background Art
[0002] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, signals are typically encoded in block units by predicting blocks and further encoding only the difference between the original block and its prediction. In particular, such encoding can include the transformation, quantization, and generation of a bitstream, usually including some form of entropy encoding. Typically, the three components (transformation, quantization, and entropy encoding) of a hybrid encoding method are optimized separately. Modern video compression standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use transform representations to encode the residual signal after prediction.
[0003] In recent years, neural network architectures have been applied to image and / or video coding. Generally, these neural network (NN)-based approaches can be applied to image and video coding in a variety of different ways. For example, several end-to-end optimized image or video coding frameworks have been considered. Further, deep learning has been used to determine or optimize some parts of an end-to-end coding framework such as the selection or compression of prediction parameters. Also, several neural network-based approaches have been considered for use in hybrid image and video coding frameworks, for example, for implementation as a trained deep learning model for intra or inter prediction in image or video coding.
[0004] The end-to-end optimized image or video coding applications described above have in common that they generate some feature map data to be transmitted between an encoder and a decoder.
[0005] A neural network is a machine learning model that uses one or more layers of non-linear units, based on which the machine learning model can predict the output of the received input. Some neural networks include one or more hidden layers in addition to the output layer. As the output of each hidden layer, a corresponding feature map can be provided. Such corresponding feature maps of each hidden layer can be used as input to subsequent layers in the network, i.e., subsequent hidden layers or the output layer. Each layer of the network generates an output from the received input according to the current values of its respective parameter set. In a neural network that is split between devices, for example, between an encoder and a decoder, between a device and the cloud, or between different devices, the feature map at the output of the split location (e.g., the first device) is compressed and transmitted to the remaining layers of the neural network (e.g., the second device).
[0006] Typically, a model is trained for a desired compression quality or bitstream size. If different compression qualities are required, a new model needs to be trained, which requires a lot of time and computational cost. Also, depending on the number of models, the amount of storage required increases.
[0007] Further improvements in encoding and decoding using a trained network architecture may be desirable. SUMMARY OF THE INVENTION
[0008] The present disclosure provides a method and apparatus for improving the flexibility of a pre-trained image or video encoding model. Thus, without sacrificing image quality, the storage and transmission costs of the encoded image or video are reduced, or the distortion of the reconstructed image is reduced without more bits.
[0009] The foregoing and other objects are achieved by the subject matter of the independent claims. Further embodiments will become apparent from the dependent claims, the description, and the figures.
[0010] Certain embodiments are outlined in the appended independent claims, and other embodiments are outlined in the dependent claims. MEANS FOR SOLVING THE PROBLEM
[0011] According to a first aspect, the present disclosure relates to a method for image encoding using a neural network. The method may be executed by an encoding device. The method includes obtaining an image, obtaining first encoding parameters of the image, where the value of the first encoding parameters is less than a preset minimum value or greater than a preset maximum value, obtaining a target gain vector based on the first encoding parameters, and encoding the image based on the target gain vector. The first encoding parameter (denoted as β) is used to select the compression quality. The larger β is, the larger the bitstream size becomes and the better the quality of the reconstructed data becomes.
[0012] Such a method can give more flexibility to a pre-trained encoding model because it enables encoding an image at any desired compression quality or bitstream size without training a new encoding model, especially when the compression quality indicated by the first encoding parameter is outside the pre-training range. Using this method, the pre-trained model can be used to encode an image at any desired quality, which can be flexibly deployed in various scenarios.
[0013] Each of the foregoing and other embodiments can optionally include one or more of the following features, either alone or in combination.
[0014] In a possible embodiment, a preset minimum value β s and a preset maximum value β t are stored in the encoding device. During the training stage of the encoding model (e.g., the encoder 101 in FIG. 3A), some values of β between [β s , β t are input into the model and trained to obtain corresponding gain vectors. Next, a pre-training set of possible values of β and corresponding gain vectors is obtained and stored in both the encoding device and the decoding device. The preset minimum value β s is the lower limit of the range of values of β, and the preset maximum value β t is the upper limit of the range of values of β.
[0015] In a possible embodiment, when the value of the first encoding parameter is smaller than the preset minimum value, obtaining a target gain vector based on the first encoding parameter includes obtaining a target gain vector based on the first encoding parameter, the preset minimum value, and a first gain vector corresponding to the preset minimum value.
[0016] Such a method enables encoding an image with a smaller bitstream size and lower compression quality, which can meet the encoding requirements in poor communication network situations.
[0017] In a possible embodiment, a preset minimum value and a first gain vector corresponding to the preset minimum value are stored in an encoding device or a remote device as a table.
[0018] In a possible embodiment, the relationship between the preset minimum value and the first gain vector is stored in the encoding device or the remote device. The relationship may be in the form of a function or a network model.
[0019] In a possible embodiment, obtaining a target gain vector based on a first encoding parameter, a preset minimum value, and a first gain vector includes obtaining a first ratio of the first encoding parameter to the preset minimum value, and obtaining a target gain vector based on the first ratio and the first gain vector.
[0020] In a possible embodiment, obtaining a target gain vector based on the first ratio and the first gain vector includes multiplying the first ratio by the first gain vector to obtain the target gain vector. The multiplication may be an element-wise multiplication operation.
[0021] In a possible embodiment, the target gain vector satisfies the following condition:
Equation
[0022] In a possible embodiment, when the value of the first encoding parameter is greater than a preset maximum value, obtaining a target gain vector based on the first encoding parameter includes obtaining the target gain vector based on the first encoding parameter, the preset maximum value, and a second gain vector corresponding to the preset maximum value.
[0023] Such a method enables encoding an image with higher compression quality, which can meet the encoding requirements for some high encoding quality situations such as encoding high-resolution movies or real-time transmission of sports events.
[0024] In a possible embodiment, the preset maximum value and the second gain vector corresponding to the preset maximum value are stored in the encoding device or the remote device as a table.
[0025] In a possible embodiment, the relationship between the preset maximum value and the second gain vector is stored in the encoding device or the remote device. The relationship may be in the form of a function or a network model.
[0026] In a possible embodiment, obtaining a target gain vector based on the first encoding parameter, the preset maximum value, and the second gain vector includes obtaining a second ratio of the first encoding parameter to the preset maximum value, and obtaining the target gain vector based on the second ratio and the second gain vector.
[0027] In a possible embodiment, obtaining a target gain vector based on the second ratio and the second gain vector includes multiplying the second ratio by the second gain vector to obtain the target gain vector. The multiplication may be an element-wise multiplication operation.
[0028] In a possible embodiment, the target gain vector satisfies the following conditions
Number
[0029] In a possible embodiment, when the value of the first encoding parameter is smaller than a preset minimum value, obtaining the target gain vector based on the first encoding parameter includes obtaining the target gain vector based on the first encoding parameter, the preset minimum value, the first gain vector corresponding to the preset minimum value, the third preset value closest to the preset minimum value, and the third gain vector corresponding to the third preset value. The third preset value is the pre-trained value closest to the preset minimum value among some pre-trained sets of values of β.
[0030] Such a method enables encoding an image with a smaller bitstream size and relatively good compression quality by using two closest preset values of the encoding parameter and their corresponding gain vectors.
[0031] In a possible embodiment, when the value of the first encoding parameter is greater than a preset maximum value, obtaining a target gain vector based on the first encoding parameter includes obtaining the target gain vector based on the first encoding parameter, the preset maximum value, a second gain vector corresponding to the preset maximum value, a fourth preset value closest to the preset maximum value, and a fourth gain vector corresponding to the fourth preset value. The fourth preset value is a pre-trained value closest to the preset maximum value among some pre-trained values of β.
[0032] Such a method enables encoding an image with a higher compression quality by using two closest preset values of the encoding parameter and their corresponding gain vectors.
[0033] In a possible embodiment, obtaining a target gain vector based on the first encoding parameter includes obtaining the target gain vector based on the first encoding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values include a preset minimum value and / or a preset maximum value.
[0034] Such a method can provide more encoding performance by using more pre-trained values of β and their corresponding gain vectors.
[0035] In a possible implementation according to any one of the first aspect and its possible implementations, encoding an image based on a target gain vector includes obtaining a first feature map of the image using a neural network, obtaining a second feature map based on the first feature map and the target gain vector, quantizing the second feature map to obtain a quantized second feature map, and encoding the quantized second feature map to obtain a bitstream. For example, the quantized second feature map may be encoded using entropy encoding.
[0036] In a possible implementation, obtaining a second feature map based on the first feature map and the target gain vector includes multiplying the target gain vector and the first feature map. The multiplication may be an element-wise multiplication operation.
[0037] In a possible implementation, the first feature map is a tensor having a shape of w×h×d, the target gain vector is a vector having a dimension of 1×d, w and h represent the width and height of the first feature map, and d represents the number of channels of the first feature map. The gain vector can be regarded as part of the model weights.
[0038] In a possible implementation, when the first feature map is a feature map of the luma samples of the image, d is equal to 128.
[0039] In a possible implementation, when the first feature map is a feature map of the chroma samples of the image, d is equal to 64.
[0040] In a possible embodiment, the method further includes obtaining a second encoding parameter of the image, where when the first encoding parameter is used to encode the luma samples of the image, the second encoding parameter is used to encode the chroma samples of the image, or when the first encoding parameter is used to encode the chroma samples of the image, the second encoding parameter is used to encode the luma samples of the image. At least one of the values of the first encoding parameter and the second encoding parameter is smaller than a preset minimum value or larger than a preset maximum value.
[0041] In a possible embodiment, the method further includes encoding the first encoding parameter into a bitstream. In a possible design, the first encoding parameter is directly encoded into the bitstream. In another possible design, the first encoding parameter is encoded into the bitstream as a first flag (e.g., the base-2 logarithm of the first encoding parameter) to save bits, and the first flag can be used to derive the first encoding parameter.
[0042] In a possible embodiment, the first encoding parameter is encoded in the picture parameter set (PPS) of the bitstream.
[0043] In a possible embodiment, the number of bits used to signal the first encoding parameter in the bitstream is 16 or less. When the first flag is used to derive the first encoding parameter, the bits used to signal the first encoding parameter in the bitstream can be reduced to 4 or less.
[0044] In a possible embodiment, the image includes both luma samples and chroma samples, and a second encoding parameter is also signaled in the bitstream. When the first encoding parameter is used to encode the luma samples of the image, the second encoding parameter is used to encode the chroma samples of the image, or when the first encoding parameter is used to encode the chroma samples of the image, the second encoding parameter is used to encode the luma samples of the image. At least one of the values of the first encoding parameter and the second encoding parameter is less than a preset minimum value or greater than a preset maximum value.
[0045] According to a second aspect, the present disclosure relates to a method for decoding a bitstream to obtain an image. The method is executed by a decoding device. The method includes obtaining a bitstream including encoded image data, analyzing the bitstream to obtain a first encoding parameter, the value of the first encoding parameter being less than a preset minimum value or greater than a preset maximum value, obtaining a target inverse gain vector based on the first encoding parameter, and obtaining an image based on the target inverse gain vector.
[0046] Such a method can give more flexibility to a pre-trained decoding model because it enables decoding an image at any desired compression quality or bitstream size without training a new decoding model, especially when the compression quality indicated by the first encoding parameter is outside the pre-trained range. Using this method, a pre-trained model can be used to encode or decode an image at any desired quality, which can be flexibly deployed in various scenarios.
[0047] The foregoing and other embodiments can each optionally include one or more of the following features, either alone or in combination.
[0048] In a possible implementation, a preset minimum value β s and a preset maximum value β t are stored in the decoding device. During the training stage of the decoding model (e.g., the decoder 104 in FIG. 3A), some values of β between s , β t are input into the model and trained to obtain the corresponding gain vectors. Next, a pre-training set of possible values of β and the corresponding gain vectors is obtained and stored in both the encoding device and the decoding device. The preset minimum value β s is the lower limit of the range of values of β, and the preset maximum value β t is the upper limit of the range of values of β.
[0049] In a possible implementation, when the value of the first encoding parameter is smaller than the preset minimum value, obtaining a target inverse gain vector based on the first encoding parameter includes obtaining the target inverse gain vector based on the first encoding parameter, the preset minimum value, and the first gain vector corresponding to the preset minimum value.
[0050] Such a method enables decoding an image with a smaller bitstream size and lower compression quality. This can meet the encoding requirements in poor communication network situations. Using this method, a pre-trained model can be used to encode or decode an image with any desired quality, which can be flexibly deployed in various scenarios.
[0051] In a possible implementation, the preset minimum value and the first gain vector corresponding to the preset minimum value are stored in the decoding device or a remote device as a table.
[0052] In a possible implementation, the relationship between the preset minimum value and the first gain vector is stored in the decoding device or a remote device. The relationship may be in the form of a function or a network model.
[0053] In a possible embodiment, obtaining a target inverse gain vector based on a first encoding parameter, a preset minimum value, and a first gain vector includes obtaining a first ratio of the first encoding parameter to the preset minimum value, and obtaining a target inverse gain vector based on the first ratio and the first gain vector.
[0054] In a possible embodiment, obtaining a target inverse gain vector based on the first ratio and the first gain vector includes multiplying the first ratio and the first gain vector to obtain a target gain vector, and obtaining a target inverse gain vector based on the target gain vector.
[0055] In a possible embodiment, the target gain vector satisfies the following conditions:
Number
[0056] In a possible embodiment, when the value of the first encoding parameter is greater than a preset maximum value, obtaining a target inverse gain vector based on the first encoding parameter includes obtaining a target inverse gain vector based on the first encoding parameter, the preset maximum value, and a second gain vector corresponding to the preset maximum value.
[0057] Such a method enables decoding an encoded image with a higher compression quality by using two closest preset values of encoding parameters and their corresponding gain vectors. Using this method, a pre-trained model can be used to encode or decode an image with any desired quality, which can be flexibly deployed in various scenarios.
[0058] In a possible implementation, obtaining a target inverse gain vector based on a first encoding parameter, a preset maximum value, and a second gain vector includes obtaining a second ratio of the first encoding parameter to the preset maximum value and obtaining a target inverse gain vector based on the second ratio and the second gain vector.
[0059] In a possible implementation, obtaining a target inverse gain vector based on the second ratio and the second gain vector includes multiplying the second ratio and the second gain vector to obtain a target gain vector and obtaining a target inverse gain vector based on the target gain vector.
[0060] In a possible implementation, the target gain vector satisfies the following condition,
Equation
[0061] In a possible implementation, the target inverse gain vector satisfies the following condition,
Equation
[0062] In a possible embodiment, when the value of the first encoding parameter is smaller than a preset minimum value, obtaining the target inverse gain vector based on the first encoding parameter includes the first encoding parameter, the preset minimum value, the first gain vector corresponding to the preset minimum value, the third preset value closest to the preset minimum value, and obtaining the target inverse gain vector based on the third gain vector corresponding to the third preset value.
[0063] In a possible embodiment, when the value of the first encoding parameter is larger than a preset maximum value, obtaining the target inverse gain vector based on the first encoding parameter includes the first encoding parameter, the preset maximum value, the second gain vector corresponding to the preset maximum value, the fourth preset value closest to the preset maximum value, and obtaining the target inverse gain vector based on the fourth gain vector corresponding to the fourth preset value.
[0064] In a possible embodiment, obtaining the target inverse gain vector based on the first encoding parameter includes obtaining the target inverse gain vector based on the first encoding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values include the preset minimum value and / or the preset maximum value.
[0065] In a possible embodiment, decrypting the bitstream to obtain an image based on the target inverse gain vector includes analyzing the bitstream to obtain a first latent representation of the image using entropy decoding, obtaining a second latent representation based on the first latent representation and the target inverse gain vector, and decrypting the second latent representation to obtain an image using a neural network.
[0066] In a possible embodiment, obtaining a second latent representation based on the first latent representation and the target inverse gain vector includes multiplying the target inverse gain vector by the first latent representation. The multiplication may be an element-wise multiplication operation.
[0067] In a possible embodiment, the first latent representation is a tensor having a shape of w×h×d, the target inverse gain vector is a vector having a dimension of 1×d, w and h represent the width and height of the first feature map, and d represents the number of channels of the first feature map.
[0068] In a possible embodiment, when the first latent representation is a feature map of the luma samples of the image, d is equal to 128.
[0069] In a possible embodiment, when the first latent representation is a feature map of the chroma samples of the image, d is equal to 64.
[0070] In a possible embodiment, the method further includes analyzing the bitstream to obtain a second encoding parameter, where when the first encoding parameter is used to decrypt the luma samples of the image, the second encoding parameter is used to decrypt the chroma samples of the image, or when the first encoding parameter is used to decrypt the chroma samples of the image, the second encoding parameter is used to decrypt the luma samples of the image.
[0071] The proposed method enables the decoding of an image using luma samples and chroma samples compressed with different qualities.
[0072] According to a third aspect, the present disclosure relates to an apparatus / device for decoding an image or video. Such an apparatus for decoding may refer to the same advantageous effects as the method for decoding according to the second aspect. Details are not described again here. The decoding apparatus provides technical means for performing the actions in the method defined according to the second aspect. Its function may be implemented by hardware or by hardware executing corresponding software. In a possible implementation, the decoding apparatus / device includes an entropy decoding module configured to analyze a bitstream to obtain a first latent representation of the image using entropy decoding, an inverse gain unit configured to obtain a second latent representation based on the first latent representation and a target inverse gain vector, and an image reconstruction module configured to decode the second latent representation to obtain the image using a neural network. These modules may be adapted to provide their respective functions corresponding to the method example according to the second aspect. For details, reference is made to the detailed description in the method example. Details are not described again here.
[0073] In a possible implementation, the decoding device further comprises an inverse gain vector acquisition module configured to acquire a target inverse gain vector based on the first encoding parameter.
[0074] According to a fourth aspect, the present disclosure relates to an apparatus / device for encoding an image or video. Such an apparatus for encoding may refer to the same advantageous effects as the method for encoding according to the first aspect. Details are not described again here. The encoding apparatus provides technical means for performing the actions in the method defined according to the first aspect. Its function may be implemented by hardware or by hardware that executes corresponding software. In a possible embodiment, the encoding apparatus includes a feature map acquisition module configured to acquire a first feature map from an input image, a gain unit configured to transform the first feature map based on a target gain vector to acquire a second feature map, a quantization module configured to quantize the second feature map to acquire a quantized second feature map, and an entropy encoding module configured to encode the quantized second feature map to acquire a bitstream (e.g., using entropy encoding). These modules may be adapted to provide their respective functions corresponding to the method example according to the first aspect. For details, reference is made to the detailed description in the method example. Details are not described again here.
[0075] In a possible embodiment, the encoding device may further include a gain vector acquisition module configured to acquire a target gain vector based on a first encoding parameter.
[0076] The method according to the first aspect of the present disclosure may be executed by the apparatus according to the fourth aspect of the present disclosure. Further features and embodiments of the method according to the first aspect of the present disclosure correspond to the respective features and embodiments of the apparatus according to the fourth aspect of the present disclosure. The advantages of the method according to the first aspect can be the same as the advantages of the corresponding embodiments of the apparatus according to the fourth aspect.
[0077] The method according to the second aspect of the present disclosure may be executed by the apparatus according to the third aspect of the present disclosure. Further features and embodiments of the method according to the second aspect of the present disclosure correspond to the respective features and embodiments of the apparatus according to the third aspect of the present disclosure. The advantages of the method according to the second aspect can be the same as the advantages of the corresponding embodiments of the apparatus according to the third aspect.
[0078] According to a fifth aspect, the present disclosure relates to a video stream or image decoding apparatus including a processor and a memory. The memory stores instructions that cause the processor to execute the method according to the second aspect.
[0079] According to a sixth aspect, the present disclosure relates to a video stream or image encoding apparatus including a processor and a memory. The memory stores instructions that cause the processor to execute the method according to the first aspect.
[0080] According to a seventh aspect, there is provided a computer-readable storage medium storing instructions that, when executed, cause one or more processors to encode video or image data. The instructions cause the one or more processors to execute the method according to the first aspect or the second aspect or any possible embodiment of the first aspect or the second aspect.
[0081] According to an eighth aspect, the present disclosure relates to a computer program product including program code for executing the method according to the first aspect or the second aspect or any possible embodiment of the first aspect or the second aspect when executed on a computer.
[0082] According to a ninth aspect, the present disclosure relates to an encoder including a processing circuit for executing the method according to the first aspect or the second aspect or any possible embodiment of the first aspect or the second aspect.
[0083] According to a tenth aspect, the present disclosure relates to a storage medium that stores a bitstream obtained using the method according to the first aspect or any possible embodiment of the first aspect.
[0084] According to an eleventh aspect, the present disclosure relates to a storage medium that stores a bitstream that can be decoded using the method according to the second aspect or any possible embodiment of the second aspect.
[0085] According to a twelfth aspect, the present disclosure relates to an encoded bitstream that includes encoded image data and a plurality of syntax elements, the plurality of syntax elements including a first flag (such as compression_quality_level) that indicates the compression quality of the encoded image data. The encoded bitstream may be obtained by performing the method according to the first aspect or any possible embodiment of the first aspect of the present disclosure. The encoded bitstream can be decoded by performing the method according to the second aspect or any possible embodiment of the second aspect of the present disclosure.
[0086] According to a thirteenth aspect, the present disclosure relates to an encoded bitstream that includes encoded image data and a plurality of syntax elements, the plurality of syntax elements including a first flag (such as compression_quality_level_luma) and a second flag (such as compression_quality_level_chroma), the first flag indicating the compression quality of the luma samples of the encoded image data and the second flag indicating the compression quality of the chroma samples of the encoded image data. The encoded bitstream may be obtained by performing the method according to the first aspect or any possible embodiment of the first aspect of the present disclosure. The encoded bitstream can be decoded by performing the method according to the second aspect or any possible embodiment of the second aspect of the present disclosure.
[0087] According to a 14th aspect, the present disclosure relates to an encoding apparatus comprising: a receiver unit configured to receive a picture to be encoded or a bitstream to be decoded; a transmitter unit coupled to the receiver unit, the transmitter unit being configured to transmit the bitstream to a decoder or to transmit the decoded picture to a display; a memory coupled to at least one of the receiver unit or the transmitter unit, the memory being configured to store instructions; and a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to perform a method according to the 1st aspect or the 2nd aspect or any possible implementation of the 1st aspect or the 2nd aspect.
[0088] According to a 15th aspect, the present disclosure relates to an encoding system comprising: an encoder; and a decoder in communication with the encoder, the encoder or the decoder including a decoding device according to the 3rd aspect or the 5th aspect of the present disclosure, an encoding device according to the 4th aspect or the 6th aspect of the present disclosure, or an encoding apparatus according to the 14th aspect of the present disclosure.
[0089] Details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will be apparent from the description, drawings, and claims.
[0090] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying figures and drawings.
Brief Description of the Drawings
[0091]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 6C
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
DETAILED DESCRIPTION OF THE INVENTION
[0092] Same reference numerals and signs in different drawings may indicate similar elements.
[0093] In the following description, reference is made to the accompanying drawings which form a part hereof and which illustrate specific aspects of embodiments of the present disclosure or specific aspects in which embodiments of the present disclosure may be used. It is understood that the embodiments of the present disclosure may be used in other aspects and may include structural or logical changes not shown in the figures. Accordingly, the following detailed description should not be construed in a limiting sense, and the scope of the present disclosure is defined by the appended claims.
[0094] For example, it is understood that the disclosure related to the described method may also apply to the corresponding device or system configured to perform this method, and vice versa. For example, if one or more specific method steps are described, the corresponding device may include one or more units, such as functional units, to perform the one or more method steps described, even if one or more of these units are not explicitly described or illustrated (e.g., one unit performs one or more steps, or multiple units each perform one or more of the multiple steps). On the other hand, for example, if a specific device is described based on one or more units, such as functional units, the corresponding method may include one step to perform the functions of the one or more units, even if one or more steps are not explicitly described or illustrated (e.g., one step performs the functions of one or more units, or multiple steps each perform one or more of the functions of one or more of the multiple units). Furthermore, it is understood that the features of the various exemplary embodiments and / or aspects described herein may be combined with each other, unless otherwise specifically stated.
[0095] In the specification, claims, and attached drawings of this application, terms such as "first" and "second" are intended to distinguish similar objects, but do not necessarily indicate a specific order or sequence. Terms used in this way are interchangeable in appropriate situations, and it should be understood that this is merely a method of distinction used when objects with the same attributes are described in the embodiments of this application. Furthermore, the terms "comprising", "having" and any other variations thereof mean that a process, method, system, product, or device comprising a series of units is not necessarily limited to those units, and may include other units not explicitly listed or inherent to such a process, method, product, or device, corresponding to non-exclusive inclusion.
[0096] In the description, claims, and appended drawings of this application, the term "and / or" is merely an associative relationship for explaining the associated objects. The term "and / or" indicates that three relationships can exist. For example, A and / or B can represent three cases, namely, only A exists, both A and B exist, and only B exists.
[0097] Hereinafter, some outlines of the technical terms and frameworks that can be used in the embodiments of the present disclosure are provided.
[0098] Artificial neural network An artificial neural network (ANN) or a connectionist system is a computing system vaguely inspired by the biological neural networks that make up the animal brain. Such systems generally "learn" to perform tasks by considering examples without being programmed with task-specific rules. For example, in image recognition, they can learn to identify images containing cats by analyzing exemplary images manually labeled as "cat" or "no cat" and using the results to identify cats in other images. They can do this without any prior knowledge of what a cat is, for example, that a cat has fur, a tail, whiskers, and a cat-like face. Instead, they automatically generate discriminative features from the examples they process.
[0099] An ANN is based on a collection of connected units or nodes called artificial neurons, which roughly model the neurons of the biological brain. Each connection, like the synapses of the biological brain, can transmit signals to other neurons. An artificial neuron that receives a signal can then process it and send a signal to the neurons connected to it.
[0100] In an ANN implementation, the "signal" in a connection is a real number, and the output of each neuron is calculated by some non-linear function of the sum of its inputs. This connection is called an edge. Neurons and edges typically have weights that are adjusted as learning progresses. The weights increase or decrease the strength of the signal in the connection. A neuron may have a threshold such that a signal is sent only if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals move from the first layer (input layer) to the last layer (output layer), sometimes crossing multiple layers in the process.
[0101] The original goal of the ANN approach was to solve problems in the same way that the human brain does. Over time, attention shifted to performing specific tasks, leading to a deviation from biology. ANNs are used in a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even in activities that have traditionally been left to humans, such as drawing.
[0102] The name "Convolutional Neural Network" (CNN) indicates that the network uses a mathematical operation called convolution. Convolution is a special type of linear operation. A convolutional network is a neural network that uses convolution in at least one of its layers instead of the general matrix multiplication.
[0103] Figure 1 schematically shows a general concept of processing by a neural network such as a CNN. A convolutional neural network consists of an input layer, an output layer, and a plurality of hidden layers. The input layer is a layer to which an input (such as a part of an image as shown in Figure 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that are convolved by multiplication or other dot products. The result of the layer is one or more feature maps (f.maps in Figure 1), which may also be called channels. Subsampling may exist in some or all of the layers. As a result, as shown in Figure 1, the feature maps can become smaller. The activation function in a CNN is usually a ReLU (Rectified Linear Unit) layer, followed by further convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are colloquially called convolutions, but this is just by convention. Mathematically, it is technically a sliding dot product or cross-correlation. This is important for matrix indices in terms of how weights are determined at specific index points.
[0104] When programming a CNN to process an image, as shown in Figure 1, the input is a tensor of the shape (number of images) × (image width) × (image height) × (image depth). It should be known that the image depth can be composed of channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map of the shape (number of images) × (feature map width) × (feature map height) × (feature map channels). The convolutional layer within the neural network should have the following attributes. A convolutional kernel (hyperparameter) defined by width and height. The number of input channels and output channels (hyperparameters). The depth of the convolutional filter (input channels) must be equal to the number of channels (depth) of the input feature map.
[0105] In the past, traditional multi-layer perceptron (MLP) models have been used for image recognition. However, due to the full connections between nodes, they have suffered from high dimensionality and did not scale well with higher resolution images. An image of 1000×1000 pixels with RGB color channels has 3 million weights, which is too high to be feasibly processed efficiently in a fully connected manner on a large scale. Also, such network architectures do not take into account the spatial structure of the data and treat input pixels that are far apart in the same way as those that are close to each other. This ignores the locality of reference in image data, both computationally and semantically. Therefore, for purposes such as image recognition that are influenced by spatially local input patterns, the full connection of neurons is wasteful.
[0106] Convolutional neural networks are biologically inspired variants of multi-layer perceptrons specifically designed to emulate the behavior of the visual cortex. These models alleviate the problems posed by the MLP architecture by leveraging the strong spatially local correlations present in natural images. The convolutional layer is the core building block of the CNN. The parameters of the layer consist of a set of learnable filters (the kernels mentioned above), which have a small receptive field but span the full depth of the input volume. During the forward pass, each filter is convolved over the width and height of the input volume, computing the dot product between the filter and the entries of the input and generating a 2D activation map for that filter. As a result, the network learns filters that activate when it detects some specific type of feature at some spatial location within the input.
[0107] Stacking the activation maps of all filters along the depth dimension forms the full output volume of the convolutional layer. Thus, all entries within the output volume can also be interpreted as the output of neurons that share neurons and parameters within the same activation map, focusing on small regions within the input. The feature map, i.e., the activation map, is the output activation of a given filter. The feature map and activation have the same meaning. In some papers, it is called the activation map because it is a mapping corresponding to the activation of different parts of the image, and it is also called the feature map because it is a mapping of where specific types of features can be found within the image. High activation means that a specific feature has been found.
[0108] Another important concept in CNNs is pooling, which is a form of non-linear downsampling. There are several non-linear functions for performing pooling, among which max pooling is the most common. It divides the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.
[0109] Intuitively, the exact location of a feature is not as important as its approximate location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. The pooling layer gradually reduces the spatial size of the representation, reduces the number of parameters, memory footprint, and computational load within the network, and thus also helps to control overfitting. In CNN architectures, it is common to periodically insert a pooling layer between consecutive convolutional layers. The pooling operation provides another form of translational invariance.
[0110] The pooling layer functions independently for each depth slice of the input and spatially resizes it. The most common form is a pooling layer where a 2×2 filter is applied with a stride of 2 in all depth slices of the input along both width and height, discarding 75% of the activations. In this case, all max operations are over four numbers. The depth dimension remains invariant. In addition to max pooling, the pooling unit can use other functions such as average pooling or l2-norm pooling. Average pooling has historically been used more often, but recently it has become less preferred compared to max pooling, which generally functions better in practice. There is a recent trend to use smaller filters or to discard the pooling layer altogether due to the aggressive reduction of the size of the representation. "Region of interest" pooling (also known as ROI pooling) is a variant of max pooling where the output size is fixed and the input rectangle is a parameter. Pooling is an important component of convolutional neural networks for object detection based on the Fast R-CNN architecture.
[0111] The ReLU mentioned above is an abbreviation of the rectified linear unit that applies a non-saturating activation function. It effectively removes negative values from the activation map by setting negative values to 0. It enhances the decision function and the non-linear characteristics of the entire network without affecting the receptive field of the convolutional layer. Other functions, such as the saturated hyperbolic tangent and sigmoid functions or LeakyReLU, are also used to enhance non-linearity. ReLU is often preferred over other functions because it can train neural networks several times faster without a significant penalty to generalization accuracy.
[0112] After several convolutional layers and max pooling layers, high-level inferences in a neural network are made by fully connected layers. Neurons in fully connected layers have connections to all activations of the previous layer, as seen in a normal (non-convolutional) artificial neural network. Thus, those activations can be computed as an affine transformation, followed by a bias offset (vector addition of a learned or fixed bias term) after matrix multiplication.
[0113] The "loss layer" (which includes the calculation of a loss function) determines how unfavorable the deviation between the predicted (output) label and the actual label is during training and is usually the final layer of a neural network. Various loss functions suitable for different tasks can be used. The softmax loss is used to predict a single class out of K mutually exclusive classes. The sigmoid cross-entropy loss is used to predict K independent probability values in [0,1]. The Euclidean loss is used for regression to real-valued labels.
[0114] In summary, Figure 1 shows the data flow in a typical convolutional neural network. First, the input image passes through convolutional layers and is abstracted into a feature map containing several channels corresponding to the number of filters within the set of learnable filters of this layer. Next, the feature map is subsampled, for example using a pooling layer, which reduces the dimensions of each channel within the feature map. Then, the data reaches another convolutional layer that can have a different number of output channels. As mentioned above, the number of input and output channels are hyperparameters of the layer. To establish the connections of the network, these parameters need to be synchronized between two connected layers such that the number of input channels of the current layer equals the number of output channels of the previous layer. For the first layer that processes input data, e.g., an image, the number of input channels usually equals the number of channels of the data representation (e.g., 3 channels for RGB or YUV representations of an image or video, or 1 channel for a grayscale image or video representation).
[0115] Autoencoders and Unsupervised Learning An autoencoder is a type of artificial neural network used to learn efficient data encoding without a teacher. A schematic diagram thereof is shown in FIG. 2. The purpose of an autoencoder is to learn the representation (encoding) of a dataset, typically for dimensionality reduction, by training the network to ignore the signal “noise”. Along with the reduction side, a reconstruction side is learned, and the autoencoder attempts to generate from the reduced encoding its original input, and thus a representation as close as possible to its name. In the simplest case, when given one hidden layer, the encoder stage of the autoencoder takes the input x and maps it to h. h = σ(Wx + b)
[0116] This image h is usually called a code, latent variable, or latent representation. Here, σ is an element-wise activation function such as the sigmoid function or the rectified linear unit. W is a weight matrix and b is a bias vector. The weights and biases are usually initialized randomly and then updated iteratively during training by backpropagation. Then, the decoder stage of the autoencoder maps h to a reconstruction x' of the same shape as x. x' = σ'(W'h' + b') Here, the σ', W', and b' of the decoder may be independent of the corresponding σ, W, and b of the encoder.
[0117] Variational autoencoder models make strong assumptions about the distribution of latent variables. They use variational methods for latent representation learning, which results in an additional loss component and a specific estimator for a training algorithm called the stochastic gradient variational Bayes (SGVB) estimator. It is generated by the directed graphical model p θ (x|h), and the encoder is an approximation q θ (h|x) for the posterior distribution p ΦAssume that we learn $(h|x)$, where $\Phi$ and $\theta$ represent the parameters of the encoder (recognition model) and decoder (generation model), respectively. The probability distribution of the latent vector of the VAE typically matches the probability distribution of the training data much closer than that of a standard autoencoder. The objective of the VAE has the following form.
Number
[0118] Here, $D$ KL represents the Kullback–Leibler divergence. The prior value for the latent variable is usually set to the centered isotropic multivariate Gaussian $p$ θ $(h) = \mathcal{N}(0, I)$. Generally, the shapes of the variational and likelihood distributions are chosen to be factorized Gaussian distributions. $q$ Φ $(h|x) = \mathcal{N}(\rho(x), \omega$ 2 (x)I)$ $p$ Φ (x|h) = \mathcal{N}(\mu(h), \sigma$ 2 (h)I)$ Here, $\rho(x)$ and $\omega$ 2 (x) are the encoder outputs, and $\mu(h)$ and $\sigma$ 2 (h) are the decoder outputs.
[0119] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, have enabled researchers' interest in applying neural network-based techniques to the tasks of image and video compression. For example, End-to-end Optimized Image Compression using a network based on variational autoencoders has been proposed.
[0120] Therefore, data compression is considered a fundamental and well-studied problem in engineering and is generally formulated with the goal of designing codes for a given discrete data ensemble with minimum entropy. This solution highly depends on the knowledge of the probabilistic structure of the data, and thus, this problem is closely related to probabilistic source modeling. However, since all practical codes must have a finite entropy, continuous-valued data (such as vectors of image pixel intensities) must be quantized into a finite set of discrete values, which introduces errors.
[0121] In this context, known as the irreversible compression problem, a trade-off must be made between two competing costs, namely the entropy (rate) of the discretized representation and the error (distortion) resulting from quantization. Different compression applications, such as data storage or transmission over channels with limited capacity, require different rate-distortion trade-offs.
[0122] The joint optimization of rate and distortion is difficult. Without additional constraints, the general problem of optimal quantization in high-dimensional spaces is intractable. For this reason, most existing image compression methods work by linearly transforming the data vector into an appropriate continuous-valued representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This approach is called transform coding due to the central role of the transformation.
[0123] For example, JPEG uses the discrete cosine transform for blocks of pixels, and JPEG2000 uses the multi-scale orthogonal wavelet decomposition. Typically, the three components of a transform coding method (the transform, the quantizer, and the entropy code) are optimized separately (often by manual parameter adjustment). The latest video compression standards such as HEVC, VVC, and EVC also use transform representations to encode the residual signal after prediction. Several transforms, such as the discrete cosine and sine transforms (DCT, DST) and the low-frequency non-separable manually optimized transform (LFNST), are used for this purpose.
[0124] Variational Image Compression The variational autoencoder (VAE) framework can be considered as a non-linear transformation encoding model. The transformation process can be mainly divided into four parts. This is illustrated in FIG. 3A showing the VAE framework.
[0125] The transformation process can be mainly divided into four parts, and FIG. 3A illustrates the VAE framework. In FIG. 3A, the encoder 101 maps the input image x to a latent representation (denoted by y) via the function y = f(x). This latent representation is also referred to below as a point or a part within the "latent space". The function f() is a transformation function that transforms the input signal x into a more compressible representation y. The quantizer 102 [Number] transforms the latent representation y into a quantized latent representation having (discrete) values [Number] where Q represents the quantizer function. The entropy model or hyper-encoder / decoder (also known as the hyper-prior) 103 estimates the distribution of the quantized latent representation [Number] of.
[0126] The latent space can be understood as a representation of compressed data where similar data points are close to each other within the latent space. The latent space is useful for learning data features and finding a simpler representation of the data for analysis. The quantized latent representation of the hyper-prior 3 [Number] and side information [Number] is (binary) included in the bit stream 2 using arithmetic encoding (AE). Further, the quantization latent representation is the reconstructed image
Number
Number
Number
Number
Number
[0127] In FIG. 3A, the component AE105 is the quantization latent representation
Number
Number
Number
Number
[0128] Arithmetic decoding (AD) 106 is a process that reverses the binarization process, and the binary number is converted back into the sample value. Arithmetic decoding is provided by the arithmetic decoding module 106.
[0129] Note that the present disclosure is not limited to this particular framework. Further, the present disclosure is not limited to image or video compression and can also be applied to object detection, image generation, and recognition systems.
[0130] In FIG. 3A, there are two sub - networks connected to each other. A sub - network in this context is a logical division of a portion of the entire network. For example, in FIG. 3A, modules 101, 102, 104, 105, and 106 are referred to as the "encoder / decoder" sub - network. The "encoder / decoder" sub - network is responsible for encoding (generating) and decoding (parsing) the first bitstream "bitstream 1". The second network in FIG. 3A includes modules 103, 108, 109, 110, and 107 and is referred to as the "hyper - encoder / decoder" sub - network. The second sub - network is responsible for generating the second bitstream "bitstream 2". The purposes of the two sub - networks are different.
[0131] The first sub - network is · the conversion 101 of the input image x into its latent representation y (which is easier to compress this x), and ·Quantize the potential representation y into a quantized potential representation
Number
Number
Number
[0132] The purpose of the second subnetwork is to obtain the statistical characteristics (e.g., the mean value, variance, and correlation between samples of Bitstream 1) of the samples of "Bitstream 1" so that the compression of Bitstream 1 by the first subnetwork becomes more efficient. The second subnetwork generates a second bitstream "Bitstream 2" including the said information (e.g., the mean value, variance, and correlation between samples of Bitstream 1).
[0133] The second network performs the conversion 103 of the quantized potential representation
Number
Number
Number
[0134] FIG. 3A illustrates an example of a VAE (Variational Autoencoder), the details of which may vary in different embodiments. For example, in certain embodiments, there may be additional components to more efficiently obtain the statistical characteristics of the samples of bitstream 1. In such an embodiment, there may be a context model aimed at extracting the mutual correlation information of bitstream 1. The statistical information provided by the second subnetwork may be used by the components of AE (Arithmetic Encoder) 105 and AD (Arithmetic Decoder) 106.
[0135] FIG. 3A shows the encoder and decoder in a single figure. As will be apparent to those skilled in the art, the encoder and decoder may be incorporated into different devices from each other, and in very many cases, are incorporated into different devices from each other.
[0136] FIG. 3B shows the encoder, and FIG. 3C separately shows the decoder component of the VAE framework. As input, the encoder receives a picture according to some embodiments. The input picture may include one or more channels such as color channels or other types of channels, such as depth channels or motion information channels. The outputs of the encoder (as shown in FIG. 3B) are bitstream 1 and bitstream 2. Bitstream 1 is the output of the first subnetwork of the encoder, and bitstream 2 is the output of the second subnetwork of the encoder.
[0137] Similarly, in FIG. 3C, two bitstreams, namely bitstream 1 and bitstream 2, are received as input, and the reconstructed (decoded) image is
Number
[0138] Specifically, as seen in FIG. 3B, the encoder comprises an encoder 121 that converts the input x into a signal y that is then provided to the quantizer 322. The quantizer 122 provides information to the arithmetic coding module 125 and the hyper-encoder 123. The hyper-encoder 123 provides the bitstream 2, which has already been described above, to the hyper-decoder 147, which then provides the information to the arithmetic coding module 105(125).
[0139] The output of the arithmetic symbolization module is bitstream 1. Bitstream 1 and bitstream 2 are the outputs of the signal encoding, which are then provided (transmitted) to the decoding process. Unit 101 (121) is called an "encoder", but it is also possible to call the complete subnetwork described in Figure 3B an "encoder". The encoding process generally means a unit (module) that converts an input into an encoded (e.g., compressed) output. From Figure 3B, it can be understood that unit 121 can actually be considered the center of the entire subnetwork in order to perform the conversion of input x to y, which is the compressed version of x. Compression in encoder 121 can be achieved, for example, by applying a neural network or generally by applying any processing network having one or more layers. In such a network, compression can be performed by cascade processing including downsampling that reduces the size and / or number of channels of the input. Therefore, the encoder can be called, for example, a neural network (NN)-based encoder.
[0140] The remaining parts of the figure (quantization unit, hyper-encoder, hyper-decoder, arithmetic encoder / decoder) are all parts responsible for improving the efficiency of the encoding process or converting the compressed output y into a series of bits (bitstream). Quantization can be provided to further compress the output of the NN encoder 121 by irreversible compression. AE125, in combination with the hyper-encoder 123 and hyper-decoder 127 used to construct it, can perform binarization that can further compress the quantized signal by reversible compression. Therefore, it is also possible to call the entire subnetwork in Figure 3B an "encoder".
[0141] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of the signal before converting the signal into binary numbers (bits). For example, in the VAE framework, the encoder, which is a non-linear transformation, maps the input image x to y, where y has a smaller width and height than x. Since y has a smaller width and height, it thus has a smaller size, the dimensionality (size) of the signal is reduced, and thus it becomes easier to compress the signal y. Note that in general, the encoder does not necessarily need to reduce the size of both (or generally all) dimensions. Rather, some exemplary embodiments may provide an encoder that reduces the size only in one (or generally a subset of) dimension.
[0142] In J. Balle, L. Valero Laparra, and E. P. Simoncelli (2015) (“Density Modeling of Images Using a Generalized Normalization Transformation”, In: arXiv e-prints, Presented at the 4th Int. Conf. for Learning Representations, 2016) (hereinafter referred to as “Balle”), the authors proposed a framework for the end-to-end optimization of an image compression model based on non-linear transformations. The authors optimize the mean squared error (MSE), but use a more flexible transformation constructed from a cascade of linear convolutions and non-linearity. Specifically, the authors use a generalized divisive normalization (GDN) coupled non-linearity, which has been proven to be effective for the gaussianization of image density, inspired by the model of neurons in the biological visual system. After this cascade transformation, it is followed by uniform scalar quantization (i.e., each element is rounded to the nearest integer) that effectively implements a parametric form of vector quantization on the original image space. The compressed image is reconstructed from these quantized values using an approximate parametric non-linear inverse transformation.
[0143] Such an example of the VAE framework is shown in Figure 4, which utilizes six downsampling layers marked 401 - 406. The network architecture includes a hyperprior model. On the left (g a , g s ) shows an image autoencoder architecture, and on the right (h a , h s ) corresponds to an autoencoder that implements the hyperprior. The factored prior model uses the same architecture for the analysis and synthesis transforms g a and g s . Q represents quantization, and AE, AD represent an arithmetic encoder and an arithmetic decoder, respectively. The encoder feeds the input image x through g a to yield a response y (latent representation) with a spatially varying standard deviation. The encoding g a includes a plurality of convolutional layers with subsampling and generalized divisive normalization (GDN) as the activation function.
[0144] The response is fed to h a that summarizes the distribution of the standard deviation at z. Next, z is quantized, compressed, and transmitted as side information. Next, the encoder uses the quantized vector
Number
Number
Number
Number
[0145] The layer including downsampling is indicated by a downward arrow in the layer description. The layer description "Conv Nx5x5 / 2↓" means that the layer is a convolutional layer having N channels and the size of the convolutional kernel is 5x5. As described, 2↓ means that 2-fold downsampling is performed in this layer. 2-fold downsampling results in one of the dimensions of the input signal being reduced by half in the output. In FIG. 4, 2↓ indicates that both the width and height of the input image are reduced by a factor of 2. Since there are six downsampling layers, if the width and height of the input image 414 (also denoted by x) are given by w and h, the output signal z^413 has a width and height equal to w / 64 and h / 64, respectively. The modules indicated by AE and AD are the arithmetic encoder and arithmetic decoder described with reference to FIGS. 3A to 3C. The arithmetic encoder and decoder are specific embodiments of entropy encoding. AE and AD can be replaced by other means of entropy encoding. In information theory, entropy encoding is a reversible data compression method that is a reversible process used to convert the values of symbols into binary representations. Also, "Q" in the figure corresponds to the quantization operation also referred to above in relation to FIG. 4 and is further described above in the "Quantization" section. Also, the quantization operation and the corresponding quantization unit do not necessarily exist as part of component 413 or 415 and / or can be replaced by another unit.
[0146] Figure 4 also shows a decoder including upsampling layers 407 to 412. Although implemented as convolutional layers, an additional layer 420 that does not provide upsampling to the received input is provided between upsampling layers 411 and 410 in the order of processing of the input. For the decoder, a corresponding convolutional layer 430 is also shown. Such layers can be provided within the NN to perform operations on the input that change certain characteristics without changing the size of the input. However, such layers are not necessary.
[0147] When viewed in the order of processing of bitstream 2 through the decoder, the upsampling layers pass in reverse order, i.e., from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide upsampling with an upsampling ratio of 2, indicated by ↑. Of course, not all upsampling layers necessarily have the same upsampling ratio, and other upsampling ratios such as 3, 4, or 8 may be used. Layers 407 to 412 are implemented as convolutional layers (conv). Specifically, since they can be intended to provide operations on the input that are inverse to those of the encoder, the upsampling layers can apply a transposed convolution operation to the received input such that the size of the received input is increased by a factor corresponding to the upsampling ratio. However, the present disclosure is not generally limited to transposed convolution, and upsampling may be performed in any other way, such as bilinear interpolation between two adjacent samples, or nearest neighbor sample copy.
[0148] In the first sub-network, after several convolutional layers (401 - 403), on the encoder side, generalized divisive normalization (GDN) continues, and on the decoder side, inverse GDN (IGDN) continues. In the second sub-network, the activation function applied is ReLu. It should be noted that the present disclosure is not limited to such embodiments, and generally, other activation functions may be used instead of GDN or ReLu.
[0149] Cloud solutions for machine tasks Video coding for machines (VCM) is another direction in computer science that is popular today. The main idea behind this approach is to transmit an encoded representation of image or video information targeted for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to conventional image and video coding targeted at human perception, the quality characteristic is not the reconstruction quality but the performance of computer vision tasks, such as object detection accuracy. This is shown in FIG. 5.
[0150] Video encoding for machines, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks across mobile cloud infrastructure. By partitioning the network between the mobile and the cloud, it is possible to distribute the computational workload so that the overall energy and / or latency of the system is minimized. In general, collaborative intelligence is a paradigm in which the processing of neural networks is distributed among two or more different computing nodes, such as devices, generally any functionally defined nodes. Here, the term "node" does not refer to the neural network nodes mentioned above. Rather, a (computing) node here refers to a separate device / module that (physically or at least logically) implements a part of the neural network. Such devices may be a mix of different servers, different end-user devices, servers and / or user devices and / or the cloud and / or processors, etc. In other words, computing nodes can be considered as nodes belonging to the same neural network that communicate with each other to transfer encoded data within / for the neural network. For example, to enable complex computations, one or more layers may be executed on a first device, and one or more layers may be executed on another device. However, the distribution may be more fine-grained, and a single layer may be executed on multiple devices. In the present disclosure, the term "plurality" refers to two or more. In some existing solutions, a part of the neural network function is executed on a device (such as a user device or an edge device) or multiple such devices, and then the output (feature map) is passed to the cloud. The cloud is a collection of processing or computing systems located outside the devices operating a part of the neural network. The concept of collaborative intelligence has also been extended to model training.In this case, data flows in both directions, i.e., from the cloud to the mobile during backpropagation in training and from the mobile to the cloud during the forward pass in training, and the same is true for inference.
[0151] Some works presented semantic image compression by encoding deep features and then reconstructing the input images from them. Compression based on uniform quantization was shown, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient to transmit the output of the hidden layer (deep feature map) from the mobile part to the cloud rather than sending the compressed natural image data to the cloud and performing object detection using the reconstructed images. Efficient compression of the feature maps benefits the compression and reconstruction of images and videos for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are a common approach for compressing deep features (i.e., feature maps).
[0152] Today, video content accounts for over 80% of Internet traffic and that proportion is expected to increase further. Therefore, it is important to build an efficient video compression system to generate higher-quality frames with a given bandwidth budget. Furthermore, most video-related computer vision tasks, such as video object detection or video object tracking, are affected by the quality of the compressed video, and efficient video compression can benefit other computer vision tasks. On the other hand, video compression techniques are also useful for action recognition and model compression. However, in the past few decades, video compression algorithms have relied on handcrafted modules, such as block-based motion estimation and discrete cosine transform (DCT), to reduce the redundancy of video sequences as mentioned above. Each module is well designed, but the overall compression system is not optimized end-to-end. It is desirable to further improve video compression performance by jointly optimizing the entire compression system.
[0153] End-to-end image or video compression DNN-based image compression methods can utilize large-scale end-to-end training and highly non-linear transformations, which are not used in conventional methods. However, it is not straightforward to directly apply these techniques to build an end-to-end learning system for video compression. First, learning how to generate and compress motion information adjusted for video compression remains an unsolved problem. Video compression methods rely heavily on motion information to reduce the temporal redundancy of video sequences.
[0154] A simple solution is to use a learning-based optical flow to represent motion information. However, current learning-based optical flow methods aim to generate the flow field as accurately as possible. An accurate optical flow is often not optimal for a specific video task. Furthermore, the data volume of optical flow increases significantly compared to motion information in conventional compression systems, and directly applying existing compression methods to compress optical flow values greatly increases the number of bits required to store motion information. Second, it is not clear how to build a DNN-based video compression system by minimizing rate-distortion-based objectives for both residual information and motion information. Rate-distortion optimization (RDO) aims to achieve a higher-quality (i.e., less distorted) reconstructed frame when a given number of bits (or bitrate) for compression is provided. RDO is important for video compression performance. To utilize the ability of end-to-end training for learning-based compression systems, an RDO strategy is required to optimize the entire system.
[0155] In "DVC: An End-to-end Deep Video Compression Framework" by Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao (Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 11006-11015), the authors proposed an end-to-end deep video compression (DVC) model that jointly learns motion estimation, motion compression, and residual coding.
[0156] Such an encoder is shown in Figure 6A. In particular, Figure 6A shows the overall structure of an end-to-end trainable video compression framework. To compress motion information, a CNN was designated to transform the optical flow into a corresponding representation more suitable for better compression. Specifically, an autoencoder-form network is used to compress the optical flow. The motion vector (MV) compression network is shown in Figure 6B. The network architecture is somewhat similar to ga / gs in Figure 4. In particular, the optical flow is fed into a series of convolutional operations and non-linear transformations including GDN and IGDN. The number of output channels for convolution (deconvolution) is 128 except for the last deconvolution layer, which is equal to 2. Given an optical flow of size M×N×2, the MV encoder generates a motion representation of size M / 16×N / 16×128. Next, the motion representation is quantized, entropy-coded, and sent into the bitstream. The MV decoder receives the quantized representation and uses the MV encoder to reconstruct the motion information.
[0157] Figure 6C shows the structure of the motion compensation unit. Here, the previous reconstructed frame x t-1Using the original frame and the reconstructed motion information, the warping unit generates a warping frame (usually with the help of an interpolation filter such as a bilinear interpolation filter). Next, a separate CNN with three inputs generates a predicted picture. The architecture of the motion compensation CNN is also shown in FIG. 6C.
[0158] The residual information between the original frame and the predicted frame is encoded by the residual encoder network. A highly non-linear neural network is used to convert the residual into the corresponding latent representation. Compared with the discrete cosine transform in conventional video compression systems, this approach can make better use of the non-linear transformation ability and achieve higher compression efficiency.
[0159] From the above overview, considering the different parts of the video frame framework including motion estimation, motion compensation, and residual encoding, it can be understood that the CNN-based architecture can be applied to both image and video compression. Entropy encoding is a common method used for data compression, widely adopted in the industry, and is also applicable to feature map compression for either human perception or computer vision tasks.
[0160] Video Encoding for Machines Video encoding for machines (VCM) is another direction in computer science that is becoming popular today. The main idea behind this approach is to transmit an encoded representation of image or video information targeted for further processing by computer vision (CV) algorithms such as object segmentation, detection, and recognition. In contrast to conventional image and video encoding targeted at human perception, the quality metric is not the reconstruction quality but the performance of computer vision tasks, such as object detection accuracy.
[0161] Recent research has proposed a new emerging paradigm called collaborative intelligence, in which deep models are split between mobile and cloud. Extensive experiments under various hardware configurations and wireless connection modes have revealed that the optimal operating point regarding energy consumption and / or computational latency usually involves splitting the model at a deep point in the network. It has been found that today's common solutions are rarely optimal (if at all) when the model is either fully in the cloud or fully on the mobile. The concept of collaborative intelligence has also been extended to model training. In this case, data flows in both directions, i.e., from the cloud to the mobile during backpropagation in training and from the mobile to the cloud during the forward pass in training, and the same is true for inference.
[0162] In the context of recent deep models for object detection, the irreversible compression of deep feature data has been studied based on HEVC intra coding. An increase in the compression level and the proposed compression extension training to minimize this loss by generating a more robust model against quantization noise in the feature values have been observed to result in a degradation of detection performance. However, this is still a sub-optimal solution as the codec used is very complex and optimized for natural scene compression rather than deep feature compression.
[0163] The problem of deep feature compression for collaborative intelligence has been addressed by a method for the object detection task using the general YOLOv2 network for the study of the trade-off between compression efficiency and recognition accuracy. Here, the term deep feature has the same meaning as feature map. The word "deep" is derived from the idea of collaborative intelligence when the output feature map of some hidden (deep) layer is captured and transferred to the cloud for inference. This seems to be more efficient than sending compressed natural image data to the cloud and performing object detection using the reconstructed image.
[0164] Efficient compression of feature maps benefits the compression and reconstruction of images and videos for both human perception and machine vision. The drawbacks of state-of-the-art autoencoder-based compression techniques also apply to machine vision tasks. …
[0165] Improvement in Encoding Efficiency As described above, variational autoencoders have become a state-of-the-art approach for learnable image compression. When they emerged in 2017, they supported only a single quality mode (the trade-off between rate and distortion) until variable-rate techniques emerged, and their models supported the selection of compression quality within a given specific range. The core of the variable-rate model is the gain unit, which is part of a neural network-based codec that provides the model with the ability to compress input data in different compression ranges. However, conventional gain units were designed to operate within a given range of compression quality. Conventional gain units cannot adapt to some specific bitstream sizes outside the given range of compression quality.
[0166] Therefore, some embodiments of the present disclosure introduce an autoencoder gain vector extrapolation solution to enable flexible representation and transmission of content without being limited to a given or pre-trained range of compression quality or bitstream size.
[0167] Several detailed embodiments and examples related to the encoder side and the decoder side are provided below.
[0168] Encoding Method FIG. 7 is a flowchart showing an exemplary method for encoding, the method including step 701 of receiving an image, step 702 of receiving a first encoding parameter of the image, where the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value, step 703 of obtaining a target gain vector based on the first encoding parameter, and step 704 of encoding the image based on the target gain vector.
[0169] In step 701, an image is received. In an embodiment of the present disclosure, the image is an image to be compressed, and the image may be an image captured by a camera or by an encoding device (the encoding device executes method 700 to encode the image), or the image may also be an image obtained from within the encoding device (e.g., an image stored in the encoding device's album, or a picture obtained by the encoding device from the cloud or another device). The image may be a single image or a video frame. Also, the image may be a part of the image, i.e., an image block. It should be understood that the image mentioned above may be an image or an image block having image compression requirements, and the present disclosure does not limit the source of the image to be processed. The image includes luminance (Y) samples and / or chrominance (UV) samples.
[0170] In step 702, a first encoding parameter is received, and the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value. The first encoding parameter (e.g., denoted as β) is a scalar input for the encoding process used to select the compression quality. The larger β is, the larger the bitstream size and the better the quality of the reconstructed data. The preset minimum value (e.g., denoted as β s ) and the preset maximum value (e.g., denoted as β t ) are the range of pre-trained values of the first encoding parameter [β s , β tis related to the preset maximum value β t is the preset minimum value β s or more. In some embodiments, the range of the pre-trained values of the first encoding parameter β may be discontinuous, and the value of the first parameter β is a finite discrete point. In the inference process, by applying interpolation, the value of the first parameter β can be extended to any value between β s and β t and.
[0171] In a possible design, the image includes both luma samples and chroma samples, and the first encoding parameter may include a first component and a second component. At least one of the first component and the second component of the first encoding parameter satisfies the condition that the value of the first encoding parameter is less than the preset minimum value or greater than the preset maximum value. Note that the preset minimum value and the preset maximum value of the luma samples and chroma samples may be different. The first component can be used to encode luma samples and thus control the compression quality of the luma samples of the image. The second component can be used to encode chroma samples and thus control the compression quality of the chroma samples of the image.
[0172] In a possible design, the image includes both luma samples and chroma samples, and a second encoding parameter is also obtained. One of the first encoding parameter and the second encoding parameter is used to encode luma samples, and the other is used to encode chroma samples. Chroma samples and luma samples are encoded with different encoding parameters β to control the compression quality separately. Note that the preset minimum value and the preset maximum value of the chroma samples and luma samples may be different or the same, which is not limited in the present disclosure.
[0173] In an embodiment of the present disclosure, the preset minimum value β s and the preset maximum value β tIt may be stored in the encoding device or may be received from another device, which is not limited here.
[0174] In an embodiment of the present disclosure, each value of the first encoding parameter β corresponds to a gain vector, that is, the model weights of the neural network model used to encode an image. A pre-trained set of possible values of β and the corresponding gain vectors is obtained during the model training phase. As described above, the possible values of β are a preset minimum value β s and a preset maximum value β t and are between them. In the present disclosure, the gain vector is a vector having a dimension of 1×d, where d is an integer greater than 1. In a possible embodiment, when encoding the luma samples of an image, d = 128. In a possible embodiment, when encoding the chroma samples of an image, d = 64. It should be noted that d may be other integers such as 32, 48, 96, 144, 160, 176, 192, 256.
[0175] Therefore, in a possible embodiment, a pre-trained set of possible values of β and the corresponding gain vectors is stored in the encoding device and the decoding device. In another possible embodiment, only the possible values of β or the corresponding gain vectors of the pre-trained set are stored in the encoding device and the decoding device, and the mapping relationship between the possible values of β and their corresponding gain vectors is also stored in the encoding device and the decoding device. The mapping relationship indicates a one-to-one correspondence between the possible values of β and the gain vectors. The mapping relationship may be a preset table or a preset objective function or any other form as long as they are the same in the encoding device and the decoding device, which is not limited here.
[0176] In a possible design, the luma samples and the chroma samples are trained together. Thus, the luma samples and the chroma samples share a list of possible values of the first encoding parameter. However, the gain vectors obtained after the training phase are different for the luma samples and the chroma samples. Table 1 is an exemplary table showing the pre-trained β and the corresponding gain vectors for the luma samples and the chroma samples. The number of pre-trained encoding parameters is N, and N is an integer greater than 0. In a possible design, N = 13. In Table 1, β0 may be the minimum value β s and β N-1 may be the maximum value β t .
[0177]
Table 1
[0178] In a possible design, the luma samples and the chroma samples are trained separately. Thus, the luma samples and the chroma samples may not share a list of possible values of the first encoding parameter. Thus, two tables are used to store the pre-trained [β, gain vector] pairs for the luma samples and the chroma samples. Table 2 is an exemplary table showing the pre-trained β and the corresponding gain vectors for the luma samples. Table 3 is an exemplary table showing the pre-trained β and the corresponding gain vectors for the chroma samples. The number of pre-trained encoding parameters for the chroma samples is N1, and N1 is an integer greater than 0. In a possible design, N1 = 13. The number of pre-trained encoding parameters for the chroma samples is N2, and N2 is an integer greater than 0. In a possible design, similarly N2 = 13. Note that N1 and N2 may be other integer values such as N1 = 10, N2 = 5, etc. It is possible that N1 is greater than or less than N2, which is not limited in this disclosure. In Table 2, β0 may be the minimum value β s and β N1-1 may be the maximum value β tIt may be. In Table 3, β0 is the minimum value β of the chroma sample s may be, and β N2-1 is the maximum value β of the chroma sample t may be.
[0179]
Table 2
[0180]
Table 3
[0181] In step 703, a target gain vector is obtained based on the first encoding parameter. As described above, in the pre-trained set of possible values of β and the corresponding gain vectors, the possible values of β are the preset minimum value β s and the preset maximum value β t in between. The drawback of this method is that the image compression quality of the pre-trained model is limited by some minimum and maximum values. Even if more than twice as many bits are spent, a quality better than the maximum quality cannot be obtained using existing methods. Even if a lower quality below the minimum quality is acceptable in some situations, compressing the image with fewer bits is not supported by existing methods. Therefore, in order to overcome the drawbacks of existing methods, an encoding method is proposed that supports encoding images with any image compression quality. When the target compression quality is given by the first encoding parameter, even when the value of the first encoding parameter is smaller than the preset minimum value or larger than the preset maximum value, the corresponding target gain vector can be obtained using the proposed method of the present disclosure.
[0182] FIG. 8 is a schematic diagram showing the relationship between the first encoding parameter (i.e., the network input parameter) β and the gain vector. During the training phase of a neural network (such as the encoder 101 in FIG. 3A), [β s , βt Some values of β between are input into the network and trained to obtain the corresponding gain vectors. In FIG. 8, [β s ,m s , [β r ,m r , and [β t ,m t are shown, and three pairs of pre-trained [β, gain vectors], where β s is the lower limit of the range of values of β, and β t is the upper limit of the range of values of β. The present disclosure is not limited to such embodiments. Generally, any other possible number of [β, gain vector] pairs may be pre-trained, and in many situations, more than three pairs of [β, gain vector] pairs are pre-trained. For example, in a possible design where 13 pairs of [β, gain vector] pairs are pre-trained, it should be noted that this is the case.
[0183] In the present disclosure, as shown in FIG. 8, an extrapolation method for obtaining a target gain vector based on a first encoding parameter is proposed. In the inference stage, when a specific value of the first encoding parameter (shown as β v ) is given, the target gain vector (shown as m v ) can be derived according to an objective function, that is, m v = f(β v , β r , β t , β s , m r , m s , m t ,…).
[0184] In a possible embodiment, only the closest pre-trained input parameter and its corresponding gain vector are used to derive the new gain vector m v of β v . For example, when β v is greater than a preset maximum value β t , only the maximum value β t and its corresponding gain vector m t are used to derive m vis used to derive this. In this situation, m v = f(β v , β t , m t ). The function m v = f(β v , β t , m t ) has one possible implementation as
Equation
[0185] Similarly, when β v is less than a preset minimum value β s , only the minimum value β s and its corresponding gain vector m s are used to derive m v . In this situation, m v = f(β v , β s , m s ). Similarly, one possible implementation of the function m v = f(β v , β s , m s ) is the same as Equation (1)
Equation
[0186] In both Equation (1) and Equation (2), the value of K is preset and may be stored in the encoding device and the decoding device. The transmission of K is not required, and thus, the size of the bitstream can be made smaller. The default value of K may be 1. In another possible embodiment, the value of K may depend on β v . Different values of K can be used in Equation (1) and Equation (2). For example, K = 1.5 in Equation (1) and K = 0.5 in Equation (2).
[0187] The above embodiments propose deriving a target gain vector outside a predetermined range, and thus, any desired image compression quality can be achieved. In situations where lower distortion is desired, higher compression quality can be achieved by extrapolating a predetermined maximum gain vector. In situations where a lower bitstream size is desired, lower compression quality can be achieved by extrapolating a predetermined minimum gain vector.
[0188] In a possible embodiment, the two closest pre-trained values of β and their corresponding gain vectors are used to obtain the target gain vector m v of β v . For example, when β v is greater than a preset maximum value β t , the maximum value β t and its corresponding gain vector m t , as well as the pre-trained value β (shown as β v in FIG. 8) next closest to β r and its corresponding gain vector m r are used to derive m v of β v . In this situation, m v = f(β v , β r , β t , m r , m t ). The function m v = f(β v , β r , β t , m r,m t )One possible embodiment is
Number
Number
[0189] Similarly, when β v is less than the preset minimum value β s , the minimum value β s and its corresponding gain vector m s , as well as β v and the next closest pre-trained value β (shown as β r in FIG. 8) and its corresponding gain vector m r are used to derive m v of β v . In this situation, m v = f(β v , β r , β s , m r , m s ). One possible embodiment of the function m v = f(β v , β r , β s , m r , m s ) is similar to Equation (3)
Number
Number
[0190] In the example of FIG. 8, only three pairs of pre-trained [β, gain vector] are shown, and therefore, β v The pre-trained values β that are next closest to are all shown as β in equations (3) and (4). Note that FIG. 8 is a schematic diagram, and in other possible embodiments, more than three pairs of [β, gain vector] are pre-trained, and β r When is smaller than β v and β s When is larger than β v and β t The pre-trained value β that is next closest to should be different. v
[0191] In both equations (3) and (4), when K is set to the default value, the transmission of K is not required, and therefore, the size of the bitstream can be made smaller. The default value of K may be 1. In another possible embodiment, the value of K may depend on β v . Different values of K can be used in equations (3) and (4), for example, for large β v K = 1, and for smaller β v then
Number
[0192] Using more pre-trained pairs of [β, gain vector] results in better performance.
[0193] In a possible embodiment, more than one pair of [β, gain vector] is used to derive a new gain vector m of β v In a possible implementation of this embodiment, the N (N > 2) pre-trained values of β closest to β and their corresponding gain vectors are used to derive the target gain vector m of β v For example, when β v is greater than a preset maximum value β v the target gain vector satisfies the following conditions. v For example, when β v is greater than a preset maximum value β t the target gain vector satisfies the following conditions.
Equation
[0194] In a possible design, the specific form of Equation (5) is as follows.
Equation
[0195] Similarly, when β v is less than a preset minimum value β s the target gain vector satisfies the following conditions.
Equation
[0196] In a possible design, the specific form of Equation (6) is as follows.
Equation
[0197] In a possible embodiment, all pre-trained pairs of [β, gain vector] are used to derive a new gain vector m of β v And each pre-trained value of β is given a weight to control the contribution of different pre-trained values of β and their corresponding gain vectors. v And each pre-trained value of β is given a weight to control the contribution of different pre-trained values of β and their corresponding gain vectors.
[0198] It should be noted that in the above possible forms of the function, a bias or offset can be further added. The form of the function is not limited in the present disclosure. All operations related to the gain vector are element-wise operations, which means that the operations are performed on the individual elements of the gain vector.
[0199] In a possible design, when the image contains only chroma samples or only luma samples, the target gain vector is obtained in step 703, and when the image contains both chroma samples and luma samples, perhaps by executing step 703 twice, the target gain vector for the luma samples is obtained in step 703, and the target gain vector for the chroma samples is also obtained in step 703. In step 704, the image is encoded based on the target gain vector. After the target gain vector is obtained in step 703, the target gain vector is used to encode the image. FIG. 9 is a flowchart of an exemplary method for encoding an image based on a target gain vector, and FIG. 10 is an exemplary VAE framework capable of executing method 900 of FIG. 9. As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible embodiments.
[0200] The encoding method 900 shown in FIG. 9 includes step 901 of obtaining a first feature map of the image using a neural network, step 902 of obtaining a second feature map based on the first feature map and the target gain vector, step 903 of quantizing the second feature map to obtain a quantized second feature map, and step 904 of encoding the quantized second feature map to obtain a bitstream.
[0201] In step 901, a neural network is used to obtain a latent representation of the image (i.e., the first feature map). In the VAE framework, the function of the neural network can be implemented by the encoder 1001 shown in FIG. 10. The encoder 1001 maps the input image x (i.e., the image) to a latent representation (denoted by y) via the function y = f(x). The function f() is a conversion function that converts the input signal x into a more compressible representation y. Usually, y is a tensor having a shape of w×h×d, for example, w×h×128 or w×h×64, where w represents the feature map width, h represents the feature map height, and d represents the number of feature map channels. The feature map width and the feature map height may or may not be equal, and it should be noted that this is not limited here. The number of feature map channels may be other integer values, for example, 32, 48, 96, 144, 160, 176, 192, 256, etc., and this is not limited in the present disclosure. In a possible design, the number of feature map channels for the luma samples and the chroma samples may be different. For example, in the case of luma samples, it is 128, and in the case of chroma samples, it is 64. The unit 1001 is called an "encoder", but it is also possible to call the complete encoding network described in FIG. 10 an "encoder". The encoding process generally means a unit (module) that converts an input into an encoded (e.g., compressed) output.
[0202] In step 902, after the first feature map of the image is obtained in step 901, further calculations can be performed on the first feature map based on the target gain vector. The target gain vector may be the target gain vector obtained in step 703 of method 700. In a possible implementation, the function of step 902 can be executed by the gain unit 1011 after the encoder 1001. The input of the gain unit 1011 (denoted by y) is the output of the encoder 1001. The gain unit 1011 further transforms the first feature map with the target gain vector. As described in step 702, the gain vector is a vector having a dimension of 1×d, where d is an integer greater than 1. In the present disclosure, d is equal to the number of feature map channels of y. Therefore, each element of the target gain vector can be mapped to the feature map channels of y. In a possible implementation, the gain unit 1011 multiplies the target gain vector and the first feature map to obtain a second feature map. The output of the gain unit 1011 is
Number
[0203] Note that step 703 needs to be executed before step 902, and the actions described in the other steps of methods 700 and 900 can be executed in different orders and still achieve the desired results. As an example, the process shown in the accompanying figures does not necessarily require the specific order or sequential order shown in order to achieve the desired results. In certain implementations, multitasking and parallel processing may be advantageous. For example, step 702 may be executed before step 901, or step 702 may be executed after step 901, or step 702 and step 901 may be executed in parallel.
[0204] In step 903, the second feature map obtained in step 902 is quantized. The quantization process can be performed by the quantizer 1002 shown in FIG. 10. The quantizer 1002 [Number] by, the latent representation [Number] is converted into a quantized latent representation having (discrete) values [Number] where Q represents the quantizer function.
[0205] In step 904, the quantized second feature map obtained in step 903 is further encoded into bitstream 1 using entropy coding. The entropy coding process can be performed by the arithmetic encoder (AE) 1005 shown in FIG. 10. The samples of the quantized latent representation [Number] (i.e., the quantized second feature map) are converted into a binary string (which is then included in a bitstream that may further include portions corresponding to the encoded image or additional side information). Note that the arithmetic encoder 1005 is a particular implementation of entropy coding. The AE can be replaced by other means of entropy coding. In information theory, entropy coding is a reversible data compression method, which is a reversible process used to convert the values of symbols into binary representations.
[0206] The output of step 904 is a bitstream, which is then provided (transmitted) to the decoding process. As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible embodiments.
[0207] In a possible design, when the image contains only chroma samples or only luma samples, a target gain vector is obtained in step 703, and method 900 is executed to encode the image based on the target gain vector. When the image contains both chroma samples and luma samples, perhaps by executing step 703 twice, a target gain vector for the luma samples is obtained in step 703, and a target gain vector for the chroma samples is also obtained in step 703. The luma samples and chroma samples are independently encoded using method 900 based on their corresponding target gain vectors. The encoded luma samples and chroma samples can be packed into a bitstream for storage or transmission.
[0208] In a possible embodiment, encoding method 700 further includes encoding a first encoding parameter into the bitstream. Depending on the situation, different images may need to be compressed with different qualities, and thus, encoding method 700 can be used to set the first encoding parameter to different values to encode different images. The proposed method 700 enables achieving different compression qualities with a single pre-trained model, especially achieving compression qualities outside the pre-trained range.
[0209] The first encoding parameter is signaled in a bitstream (e.g., as a first flag such as compression_quality_level) and transmitted to a decoding device, whereby the decoding device can decode the encoded image data based on the first encoding parameter. In a possible implementation, the first encoding parameter is encoded or signaled in a picture parameter set (PPS) of the bitstream. In a possible implementation, the first encoding parameter is encoded or signaled in a sequence parameter set (SPS) of the bitstream. In a possible implementation, the first encoding parameter is encoded or signaled in a picture header (PH) or a slice header (SH) of the bitstream.
[0210] In a possible design, the number of bits used to signal the first encoding parameter (the first flag) within the bitstream is 16 or less. For example, the first parameter can take 15 bits in the bitstream.
[0211] In some possible designs, the first encoding parameter is signaled in other forms to save bits for transmitting the first encoding parameter. For example, in a possible design, the first encoding parameter is signaled in its base-2 logarithmic form. In another possible design, the first encoding parameter is signaled in its base-2 logarithmic form using a bias. The base of the logarithm can be another value such as 10, which is not limited here. In this way, the bits used to transmit the first encoding parameter can be reduced to 4 bits or less.
[0212] Optionally, the side information z output by the hyper encoder 1003 may pass through the gain unit 1012 before quantization. The function of the gain unit 1012 is the same as that of the gain unit 1011, but the gain vectors used in the gain unit 1012 and the gain unit 1011 may be different. This embodiment enables side information to be encoded in a more flexible manner.
[0213] In a possible embodiment, the image includes only luma samples, and the method 700 is used to encode the luma samples based on the first encoding parameter to obtain a bitstream, and the first encoding parameter is signaled as the first flag (such as compression_quality_level) in the bitstream as described above.
[0214] In a possible embodiment, the image includes only chroma samples, and the method 700 is used to encode the chroma samples based on the first encoding parameter to obtain a bitstream, and the first encoding parameter is signaled as the first flag (such as compression_quality_level) in the bitstream as described above.
[0215] In a possible embodiment, the image includes both luma samples and chroma samples. When the luma samples are encoded based on a first encoding parameter using method 700, the chroma samples are encoded based on a second encoding parameter using method 700, or when the chroma samples are encoded based on a first encoding parameter using method 700, the luma samples are encoded based on a second encoding parameter using method 700. Note that some steps of method 700 may be executed only once when encoding an image using both luma samples and chroma samples, such as in step 701. Both the first encoding parameter and the second encoding parameter are signaled in the bitstream as described above. For example, the first encoding parameter is signaled as a first flag (such as compression_quality_level_luma), and the second encoding parameter is signaled as a second flag (such as compression_quality_level_chroma). Or, the first encoding parameter is signaled as a first flag (such as compression_quality_level_chroma), and the second encoding parameter is signaled as a second flag (such as compression_quality_level_luma). The second encoding parameter is signaled in the same way as the first encoding parameter is signaled. Thus, the number of bits used to transmit the encoding parameters is doubled. At least one of the values of the first encoding parameter and the second encoding parameter is less than a preset minimum value or greater than a preset maximum value.
[0216] In a possible embodiment, the image includes both luma samples and chroma samples, and the first encoding parameter includes a first component and a second component. The luma samples are encoded based on the first component using method 700, and the chroma samples are encoded based on the second component using method 700. Note that some steps of method 700 may be executed only once when encoding the image using both luma samples and chroma samples, such as in step 701. Both the first component and the second component are signaled in the bitstream as described above. For example, the first component is signaled as a first flag (such as compression_quality_level_luma), and the second component is signaled as a second flag (such as compression_quality_level_chroma). Both the first component and the second component are signaled in the same way as the first encoding parameter is signaled as described above. Thus, the number of bits used to transmit the first encoding parameter is doubled.
[0217] In a possible embodiment, different regions of the entire image may be encoded based on different encoding parameters β v and some of the encoding parameters β v are signaled in the bitstream to be transmitted to other devices. For example, the region of interest may be encoded with a larger β v to obtain better compression quality, and the background region may be encoded with a smaller β v to obtain a smaller bitstream size. The proposed method enables different compression qualities to be achieved for different regions of a single image, the encoding process is more flexible, and different compression requirements can be met for different situations.
[0218] The embodiment according to FIG. 7 may be configured to provide an output that can be easily decoded by the decoding method described with reference to FIG. 11. The method 700 of FIG. 7 is described as being executed by a neural network system of one or more computers located at one or more positions. For example, a system configured to perform image compression, such as the neural network of FIG. 1, can execute the method 700. In general, the embodiments mentioned above can be combined to provide higher flexibility.
[0219] Decoding method FIG. 11 shows an exemplary method for decoding an image based on a neural network architecture, including step 1101 of obtaining a bitstream containing encoded image data, step 1102 of analyzing the bitstream to obtain a first encoding parameter, where the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value, step 1103 of obtaining a target inverse gain vector based on the first encoding parameter, and step 1104 of obtaining an image based on the target inverse gain vector.
[0220] In step 1101, a bitstream containing encoded image data is probably obtained from an encoding device or a distribution device. The bitstream may include information on some side information (e.g., the average value or variance of the encoded samples...), and information on some encoding parameters (e.g., the encoding mode, compression quality parameter, quantization parameter...).
[0221] In step 1102, the bitstream is analyzed to obtain a first encoding parameter, where the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value. The first encoding parameter (e.g., denoted as β) is a scalar input for the decoding process used to select the compression quality.
[0222] The first flag (such as compression_quality_level) is signaled in the bitstream to specify the compression quality of the image, i.e., the first encoding parameter β. In a possible design, the first flag is the compression quality and may require 16 bits or 15 bits for storage and transmission. In this design, the first encoding parameter β is directly parsed from the bitstream. In another possible design, the first flag is in the form of the base-2 logarithm of the first encoding parameter (such as log2_compression_quality_level or log2_compression_quality_level_minus1), and the first encoding parameter can be derived as follows: the first encoding parameter β = 1 << the first flag, or the first encoding parameter β = 1 << (the first flag + bias), where "<<" means a left shift, the first flag may be an integer in the range of 0 to 16, and the bias may be 1, 2, 3, 4, 5, 8, etc.
[0223] In a possible embodiment, the first flag and the second flag are signaled in the bitstream to specify the compression quality of the image. When the first flag (such as compression_quality_level_luma) is related to the first encoding parameter (denoted as β Y as shown) that specifies the compression quality of the luma samples of the image, the second flag (such as compression_quality_level_chroma) is related to the second encoding parameter (denoted as β UV as shown) that specifies the compression quality of the chroma samples of the image. When the first flag (such as compression_quality_level_chroma) is related to the first encoding parameter (denoted as β UV as shown) that specifies the compression quality of the chroma samples of the image, the second flag (such as compression_quality_level_luma) is related to the second encoding parameter (denoted as β Yis related to (shown as). Both the first encoding parameter and the second encoding parameter can be derived using the method described in the previous paragraph. At least one of the values of the first encoding parameter and the second encoding parameter is less than a preset minimum value or greater than a preset maximum value.
[0224] In a possible embodiment, the first flag is signaled in the bitstream to specify the compression quality of the image. The first flag includes a first component (such as compression_quality_level_luma) and a second component (such as compression_quality_level_chroma). The first component is related to a first encoding parameter (β Y shown as) that specifies the compression quality of the luma samples of the image, and the second component is related to a second encoding parameter (β UV shown as) that specifies the compression quality of the chroma samples of the image. Both the first encoding component and the second encoding component can be derived using the method described in the above paragraph.
[0225] In step 1103, after obtaining the first encoding parameter, a target gain vector is obtained based on the first encoding parameter. During the network training phase, a pre-trained set of possible values of β and the corresponding gain vectors is obtained and stored in both the encoding device and the decoding device. The possible values of β are between a preset minimum value β s and a preset maximum value β t . Therefore, the target gain vector m v can be obtained using the method described with reference to FIG. 8. Next, the target inverse gain vector m v ' can be obtained based on the target gain vector, and the target inverse gain vector m v ' satisfies the following conditions.
Equation
[0226] In a possible design, a pre-trained set of possible values of β and the corresponding inverse gain vectors are obtained during the training phase and then stored in the decoding device. Thus, the target inverse gain vector m v ’ can be obtained using the method described with reference to FIG. 8.
[0227] In step 1104, an image is obtained based on the target inverse gain vector. After obtaining the target inverse gain vector in step 1103, the target inverse gain vector is used to decode the image. FIG. 12 is a flowchart of an exemplary method for decoding an image based on the target inverse gain vector, and FIG. 10 is an exemplary VAE framework capable of executing the method 1200 of FIG. 12. Generally, any of the embodiments and their possible designs or possible implementations mentioned above can be combined to provide more flexibility.
[0228] The decoding method 1200 shown in FIG. 12 includes step 1201 of analyzing a bitstream to obtain a first latent representation of the image using entropy decoding, step 1202 of obtaining a second latent representation based on the first latent representation and the target inverse gain vector, and step 1203 of decoding the second latent representation to obtain an image using a neural network.
[0229] In step 1201, the entropy decoding process re-samples the binary numbers (i.e.,
Number
Number
[0230] In step 1202, after the first latent representation of the image is obtained in step 1201, further calculations can be performed on the first latent representation based on the target inverse gain vector. The target inverse gain vector may be the target inverse gain vector obtained in step 1103 of method 1100. In a possible embodiment, the function of step 1202 can be executed by the inverse gain unit 1013 after the AD1006. The input of the gain unit 1013 (
Number
Number
Number
[0231] Step 1103 needs to be executed before step 1202, and it should be noted that the actions described in the other steps of methods 1100 and 1200 can be executed in a different order and still achieve the desired results. As an example, the process shown in the accompanying figures does not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous. For example, step 1102 may be executed before step 1201, or step 1102 may be executed after step 1201, or step 1102 and step 1201 may be executed in parallel.
[0232] In a possible embodiment, a target gain vector is obtained in step 1103' (which is not shown in FIG. 11), and then, in step 1104' (which is not shown in FIG. 11), a bitstream is decoded to obtain an image based on the target gain vector. In this embodiment, the inverse gain unit 1013 multiplies the reciprocal of the target inverse gain vector and the first latent representation to obtain a second latent representation.
[0233] In step 1203, the second latent representation obtained in step 1202 is further decoded into an image using a neural network. In the VAE framework, the function of the neural network can be implemented by the decoder 1004 shown in FIG. 10. Although unit 1004 is called a "decoder", it is also possible to call the complete decoding network described in FIG. 10 a "decoder". The decoding process generally means a unit (module) that converts the latent representation into an image output.
[0234] The output of step 1203 is an image
Number
[0235] Optionally, when the side information z output by the hyper-encoder 1003 is processed by the gain unit 1012 before quantization. The inverse gain unit 1014 is used to process the output of AD1010 in FIG. 10.
[0236] As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible implementations. Side information may be provided.
[0237] Although the operations are shown in the drawings in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together into a single software product or packaged into multiple software products.
[0238] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes shown in the accompanying figures do not necessarily require the particular order shown, or a sequential order, to achieve desirable results. In certain embodiments, multitasking and parallel processing may be advantageous.
[0239] Based on the framework shown in FIG. 10, a simple example is provided to illustrate the technical effect of the provided method. Assuming that the output of the encoder 1001 is y = 8.53, Conventional method: A gain vector is not used to process y, and then
Number
Number
Mathematics
Mathematics
Mathematics
Mathematics
[0240] As shown in the above simple example, the proposed method of the present disclosure results in much less distortion, especially when the gain vector has a large value.
[0241] Encoding device and decoding device Furthermore, as already mentioned, the present disclosure also provides a device configured to perform the steps of the methods described above. FIG. 13 shows an encoding device 1300 for encoding for processing by a neural network-based unit. The device 1300 includes a feature map acquisition module 1310 configured to acquire a first feature map from an input image, a gain unit 1330 configured to transform the first feature map based on a target gain vector to obtain a second feature map, a quantization module 1340 configured to quantize the second feature map to obtain a quantized second feature map, and an entropy encoding module 1350 configured to encode the quantized second feature map to obtain a bitstream (e.g., using entropy encoding). Optionally, the device 1300 may further include a gain vector acquisition module 1320 configured to acquire a target gain vector based on a first encoding parameter.
[0242] Corresponding to the encoding device 1300 mentioned above, a decoding device 1400 for decoding a bitstream to reconstruct an image by a neural network-based unit is shown in FIG. 14. The device may include an entropy decoding module 1410 configured to analyze the bitstream to obtain a first latent representation of the image using entropy decoding, an inverse gain unit 1430 configured to obtain a second latent representation based on the first latent representation and a target inverse gain vector, and an image reconstruction module configured to decode the second latent representation to obtain an image using a neural network. Optionally, the device 1400 may further include an inverse gain vector acquisition module 1420 configured to acquire a target inverse gain vector based on a first encoding parameter.
[0243] Note that these devices may be further configured to perform any of the additional features including the exemplary embodiments mentioned above. For example, a device for decoding a feature map for neural network processing based on a bitstream is provided, and the device comprises a processing circuit configured to perform any of the steps of the decoding methods described above. Similarly, a device for encoding a feature map for neural network processing into a bitstream is provided, and the device comprises a processing circuit configured to perform any of the steps of the encoding methods described above.
[0244] Additional devices may be provided that utilize device 1300 and / or 1400. For example, a device for image or video encoding may include encoding device 1300. Further, it may include decoding device 1400. A device for image or video decoding may include decoding device 1400 and / or encoding device 1300.
[0245] Furthermore, an encoding system may be provided that utilizes device 1300 and / or 1400. For example, the encoding system may be deployed on a server. The server receives a bitstream, then decodes the bitstream using device 1400 or another decoder to obtain an image, and then the image is encoded using device 1300 or another encoder. After encoding, the newly obtained bitstream is stored and / or transmitted to other devices.
[0246] Some exemplary embodiments in hardware and software A corresponding system that can deploy the encoder-decoder processing chain referred to above is shown in FIG. 15. FIG. 15 is a schematic block diagram showing an exemplary encoding system that can utilize the technology of the present application, for example, a video, image, audio, and / or other encoding system (or short encoding system). The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video encoding system 10 represent examples of devices that can be configured to execute the techniques according to various examples described in the present disclosure. For example, video encoding and decoding can use a distributed neural network that can be distributed and apply the bitstream analysis and / or bitstream generation referred to above to transmit feature maps between distributed computing nodes (two or more).
[0247] As shown in FIG. 15, the encoding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 to decode the encoded picture data 13.
[0248] The source device 12 includes an encoder 20 and additionally, i.e., optionally, may include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.
[0249] The picture source 16 may include any kind of picture capture device for capturing real-world pictures, such as a camera, and / or any kind of picture generation device, such as a computer graphics processor for generating computer animation pictures, or any kind of other device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures), or it may be these. The picture source may be any kind of memory or storage for storing any of the aforementioned pictures.
[0250] Distinguished from the processing executed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be called raw picture or raw picture data 17.
[0251] The preprocessor 18 is configured to receive the (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain the preprocessed picture 19 or preprocessed picture data 19. The preprocessing executed by the preprocessor 18 may include, for example, trimming, color format conversion (e.g., from RGB to YCbCr), color correction, or noise removal. It can be understood that the preprocessing unit 18 may be an optional component. Note that the preprocessing may also use a neural network (such as those in any of FIGS. 1 to 7) that uses presence indicator signaling.
[0252] The video encoder 20 is configured to receive the preprocessed picture data 19 and provide the encoded picture data 21.
[0253] The communication interface 22 of the source device 12 receives the encoded picture data 21 and is configured to transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 for storage or direct reconstruction to another device, such as the destination device 14 or any other device.
[0254] The destination device 14 comprises a decoder 30 (e.g., a video decoder 30) and additionally, i.e., optionally, may comprise a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.
[0255] The communication interface 28 of the destination device 14 is configured to receive the encoded picture data 21 (or any further processed version thereof), e.g., directly from the source device 12 or from any other source, such as a storage device, e.g., an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.
[0256] The communication interface 22 and the communication interface 28 are configured to transmit or receive the encoded picture data 21 or the encoded data 13 via a direct communication link between the source device 12 and the destination device 14, e.g., via a direct wired or wireless connection, or via any kind of network, e.g., a wired or wireless network or any combination thereof, or via any kind of private network and public network or any combination thereof.
[0257] The communication interface 22 may be configured to, e.g., package the encoded picture data 21 in a suitable format, e.g., packets, and / or process the encoded picture data using any kind of transmission encoding or processing for transmission via a communication link or communication network.
[0258] The communication interface 28 that forms the counterpart of the communication interface 22 may be configured to process the transmitted data, for example, by receiving the transmitted data and using any kind of corresponding transmission decoding or processing and / or depackaging to obtain the encoded picture data 21.
[0259] Both the communication interface 22 and the communication interface 28 may be configured as a unidirectional communication interface or a bidirectional communication interface as indicated by the arrow of the communication channel 13 in FIG. 15, which points from the source device 12 to the destination device 14. For example, it may be configured to send and receive messages, for example, to set up a connection, and to check and exchange any other information related to the communication link and / or data transmission, for example, the encoded picture data transmission. The decoder 30 is configured to receive the encoded picture data 21 and provide the decoded picture data 31 or the decoded picture 31.
[0260] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also referred to as reconstructed picture data), for example, the decoded picture 31, in order to obtain the post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, trimming, or resampling, or any other processing for preparing the decoded picture data 31 for display by, for example, the display device 34.
[0261] The display device 34 of the destination device 14 is configured to receive the post-processed picture data 33 in order to display a picture, for example, to a user or a viewer. The display device 34 may be any kind of display for representing the reconstructed picture, such as an integrated or external display or monitor, or may include this. The display may include, for example, a liquid crystal display (LCD), an organic light emitting diode (OLED) display, a plasma display, a projector, a micro LED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other kind of display.
[0262] FIG. 15 shows the source device 12 and the destination device 14 as separate devices, but embodiments of the device may also include both or both functions, the source device 12 or corresponding functions as well as the destination device 14 or corresponding functions. In such embodiments, the source device 12 or corresponding functions as well as the destination device 14 or corresponding functions may be implemented using the same hardware and / or software, or by separate hardware and / or software or any combination thereof.
[0263] As will be apparent to those skilled in the art based on the description, the presence and (exact) division of the functions of the different units or functions within the source device 12 and / or the destination device 14 as shown in FIG. 15 may be changed according to the actual device and application.
[0264] The encoder 20 (e.g., video encoder 20) or decoder 30 (e.g., video decoder 30), or both the encoder 20 and decoder 30, may be implemented by processing circuitry such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, hardware, video encoding specific, or any combination thereof. The encoder 20 may be implemented by processing circuitry 46 to embody various modules including a neural network or portions thereof. The decoder 30 may be implemented by processing circuitry 46 to embody any of the encoding systems or subsystems described herein. The processing circuitry may be configured to perform various operations as described hereinafter. If the technology is implemented partially in software, the device may store instructions for the software in a suitable non-transitory computer-readable storage medium and execute the instructions in hardware using one or more processors to perform the technology of the present disclosure. Either the video encoder 20 or the video decoder 30 may be integrated, for example, as part of a combined encoder / decoder (CODEC) in a single device, as shown in FIG. 16.
[0265] The source device 12 and the destination device 14 can be any of a wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video gaming console, a video streaming device (content service server or content delivery server), a broadcast receiver device, or a broadcast transmitter device, etc. It may not use an operating system, or may use any type of operating system. In some cases, the source device 12 and the destination device 14 may include wireless communication. Thus, the source device 12 and the destination device 14 may be wireless communication devices.
[0266] In some cases, the video encoding system 10 shown in FIG. 15 is merely an example, and the technology of the present disclosure may be applied to video encoding settings (e.g., video encoding or video decoding) that do not necessarily include data communication between an encoding device and a decoding device. In other examples, data is retrieved from local memory and streamed over a network, etc. The video encoding device may encode data and store it in memory, and / or the video decoding device may retrieve data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other, which simply encode data into memory and / or retrieve and decode data from memory.
[0267] FIG. 17 is a schematic diagram of a video encoding device 8000 according to an embodiment of the present disclosure. The video encoding device 8000 is suitable for implementing the disclosed embodiments described herein. In one embodiment, the video encoding device 8000 may be a decoder, such as the video decoder 30 of FIG. 15, or an encoder, such as the video encoder 20 of FIG. 15.
[0268] The video encoding device 8000 includes an ingress port 8010 (or input port 8010) and a receiver unit (Rx) 8020 for receiving data, a processor, logic unit, or central processing unit (CPU) 8030 for processing data, a transmitter unit (Tx) 8040 and an egress port 8050 (or output port 8050) for transmitting data, and a memory 8060 for storing data. The video encoding device 8000 may also include optical - electrical (OE) components and electrical - optical (EO) components coupled to the ingress port 8010, the receiver unit 8020, the transmitter unit 8040, and the egress port 8050 for the output or input of optical or electrical signals.
[0269] The processor 8030 is implemented by hardware and software. The processor 8030 may be implemented as one or more CPU chips, cores (e.g., as a multi - core processor), FPGAs, ASICs, and DSPs. The processor 8030 communicates with the ingress port 8010, the receiver unit 8020, the transmitter unit 8040, the egress port 8050, and the memory 8060. The processor 8030 includes a neural - network - based codec 8070. The neural - network - based codec 8070 implements the disclosed embodiments described above. For example, the neural - network - based codec 8070 performs, processes, prepares, or provides various encoding operations. Thus, the inclusion of the neural - network - based codec 8070 brings a substantial improvement to the functionality of the video encoding device 8000 and results in the conversion of the video encoding device 8000 to different states. Alternatively, the neural - network - based codec 8070 is stored in the memory 8060 and implemented as instructions executed by the processor 8030.
[0270] Memory 8060 may include one or more disks, tape drives, and solid state drives, and may be used as an overflow data storage device to store a program when the program is selected for execution and to store instructions and data read during program execution. Memory 8060 may be, for example, volatile and / or non-volatile, and may be a read only memory (ROM), random access memory (RAM), ternary content addressable memory (TCAM), and / or static random access memory (SRAM).
[0271] FIG. 18 is a simplified block diagram of an apparatus that can be used as either or both of the source device 12 and the destination device 14 from FIG. 15, according to an exemplary embodiment.
[0272] The processor 9002 within the apparatus 9000 can be a central processing unit. Alternatively, the processor 9002 can be any other type of device, or multiple devices, capable of manipulating or processing information that exists currently or will be developed in the future. Although the disclosed embodiments can be implemented with a single processor, such as processor 9002 as illustrated, advantages in terms of speed and efficiency can be achieved using more than one processor.
[0273] In one embodiment, the memory 9004 in the apparatus 9000 can be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device can be used as the memory 9004. The memory 9004 can include code and data 9006 that are accessed by the processor 9002 using the bus 9012. The memory 9004 can further include an operating system 9008 and an application program 9010, and the application program 9010 includes at least one program that enables the processor 9002 to execute the methods described herein. For example, the application program 9010 can include applications 1 to N, and the applications 1 to N further include a video encoding application that executes the methods described herein.
[0274] The apparatus 9000 can also include one or more output devices such as a display 9018. In one example, the display 9018 can be a touch-sensitive display that combines a display and a touch-sensitive element operable to sense touch input. The display 9018 can be coupled to the processor 9002 via the bus 9012.
[0275] Although shown here as a single bus, the bus 9012 of the apparatus 9000 can be composed of multiple buses. Further, secondary storage can be directly coupled to other components of the apparatus 9000 or accessed via a network, and can include a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, the apparatus 9000 can be implemented in a wide variety of configurations.
[0276] In one possible embodiment, an exemplary method for storing a bitstream, the method comprising: obtaining a bitstream according to any one of the encoding methods shown previously; Storing a bitstream in a memory medium A method is provided that includes this.
[0277] Optionally, the method Performing an encryption process on the bitstream to obtain an encrypted bitstream, and Storing the encrypted bitstream in a memory medium further includes this.
[0278] It should be understood that any of the known encryption methods may be used.
[0279] Optionally, the method further includes performing a segmentation process on the bitstream to obtain a plurality of bitstream segments, and storing the plurality of bitstream segments in a memory medium.
[0280] Optionally, the method further includes obtaining at least one backup of the bitstream and storing the at least one backup in a memory medium. It should be understood that the at least one backup of the bitstream can be stored in a memory medium different from the memory medium storing the original bitstream.
[0281] Optionally, the method Receiving a plurality of bitstreams generated according to any one of the encoding methods shown above, and Separately assigning address information or identification information to the plurality of bitstreams, and Storing the bitstream at the corresponding position according to the address information or identification information corresponding to the plurality of bitstreams further includes this.
[0282] Optionally, the method Classify the bitstream to obtain at least two bitstreams, where the at least two bitstreams include a first bitstream and a second bitstream, and Store the first bitstream in a first storage space and store the second bitstream in a second storage space, and Further include.
[0283] Optionally, the method Further includes transmitting the bitstream to the terminal device by a video streaming device, where the video streaming device can be a content server or a content delivery server.
[0284] In one possible embodiment, an exemplary system for storing a bitstream, the system includes A receiver configured to receive a bitstream generated by any one of the previous encoding methods, A processor configured to perform an encryption process on the bitstream to obtain an encrypted bitstream, A computer-readable storage medium configured to store the encrypted bitstream, and A system is provided that includes.
[0285] Optionally, the system includes several storage media, and the several storage media can be deployed at different positions. Also, multiple bitstreams may be stored in different storage media in a distributed manner. For example, the several storage media include a first storage media configured to store a first bitstream and a second storage media configured to store a second bitstream.
[0286] Optionally, the system includes a video streaming device, which can be a content server or a content delivery server. The video streaming device is configured to obtain a bitstream from one of the storage media and transmit the bitstream to a terminal device.
[0287] In one possible embodiment, an exemplary method for converting the format of a bitstream is provided, and the method includes: Receiving a bitstream in a first format generated by any one of the previously shown encoding methods; Converting the bitstream in the first format into a bitstream in a second format; Storing the bitstream in the second format in a storage media. The method is provided.
[0288] Optionally, the method further includes: Responding to an access request from a terminal-side device and transmitting the stored bitstream in the second format to the terminal-side device. The method further includes this step.
[0289] In one possible embodiment, an exemplary system for converting a bitstream format is provided. The system includes: A receiver configured to receive a bitstream in a first format generated by any one of the previously shown encoding methods; A processor configured to convert the bitstream in the first format into a bitstream in a second format, where the processor is further configured to store the bitstream in the second format in a storage media; The storage media is configured to store the bitstream in the second format. The system includes the processor and the storage media. A transmitter configured to transmit a stored bitstream in a second format to a terminal-side device in response to an access request from the terminal-side device A system is provided that includes
[0290] In one possible embodiment, an exemplary method for processing a bitstream, the method comprising Receiving a transport stream including a video stream and an audio stream, the video stream being generated by any one of the encoding methods shown above Demultiplexing the transport stream to separate the video stream and the audio stream Decoding the video stream using a video decoder to obtain video data Decoding the audio stream using an audio decoder to obtain audio data A method is provided that includes
[0291] Optionally, the method further comprises Synchronizing the audio data and the video data Outputting the synchronization result to a player for playback and further includes
[0292] Optionally, the method further comprises Decoding the bitstream to obtain video data or image data Performing at least one of luminance mapping, chroma mapping, resolution adjustment, or format conversion on the video data or image data Transmitting the video data or image data to a display and further includes
[0293] In one possible embodiment, an exemplary method for transmitting a bitstream based on a user operation request, the method comprising Receive a first operation request from an end-side device, where the first operation request is used to request the playback of a target video, and In response to the first operation request, determine a bitstream corresponding to the target video in a storage medium, where the bitstream corresponding to the target video is a bitstream generated according to any one of the previously shown encoding methods, and Transmit the target bitstream to the end-side device A method is provided that includes the above steps.
[0294] Optionally, the method further includes Encapsulating the bitstream to obtain a transport stream in a first format, and Transmitting the transport stream in the first format to a terminal-side device for display, or Transmitting the transport stream in the first format to a storage space for storage And further includes the above steps.
[0295] In one possible embodiment, an exemplary system for transmitting a bitstream based on a user operation request, the system includes A storage medium configured to store a bitstream, where the bitstream is a bitstream generated according to any one of the previously shown encoding methods, A receiver configured to receive a first operation request, A processor configured to determine a target bitstream in the storage medium in response to the first operation request, A transmitter configured to transmit the target bitstream to a terminal-side device A system is provided that includes the above components.
[0296] Optionally, the processor is further configured to Encapsulate the bitstream to obtain a transport stream in a first format And the system further includes the above steps. Transmit a transport stream of the first format to a terminal-side device for display, or Transmit the transport stream of the first format to a storage space for storage Further include a transmitter further configured as described above.
[0297] In one possible embodiment, an exemplary method for downloading a bitstream, the method comprising: Obtain a bitstream from a storage medium, the bitstream being generated according to any one of the encoding methods shown above; Decode the bitstream to obtain a streaming media file; Split the streaming media file into a plurality of streaming media segments; Download the plurality of streaming media segments separately A method is provided that includes.
[0298] In one possible embodiment, an exemplary system for downloading a bitstream, the system comprising: An acquisition unit configured to obtain a bitstream from a storage medium, the bitstream being generated according to any one of the encoding methods shown above; A decoder configured to decode the bitstream to obtain a streaming media file; A processor configured to split the streaming media file into a plurality of streaming media segments, The processor being configured to download the plurality of streaming media segments separately, and a processor A system is provided that includes. However, the present invention is not limited to any of these exemplary embodiments.
[0299] In summary, the present disclosure relates to a method and apparatus for encoding data (into a bitstream for still image or video processing). In particular, the data is processed by a network including a gain unit. In this process, a feature map is generated by an encoder layer. Before quantization, the feature map is processed by a gain unit that further transforms the feature map with a target gain vector. The target gain vector corresponds to a first encoding parameter used to control the compression quality of an image. A user can set different values of the first encoding parameter to control the different compression qualities of different images. A user can set different values of the first encoding parameter to control the different compression qualities of luma samples and chroma samples. The first encoding parameter is also encoded into the bitstream. By this approach, flexible processing that can function with different bitstream sizes is provided. Therefore, depending on the first encoding parameter that can vary according to the content of the encoded picture data, the data can be efficiently encoded within the bitstream.
[0300] The present disclosure further relates to a method and apparatus for decoding data (from a bitstream for still image or video processing). In particular, the data is processed by a network including an inverse gain unit. In this process, a feature map is generated by an entropy decoder. Before further decoding, the feature map is processed by an inverse gain unit that further transforms the feature map with a target inverse gain vector. The target inverse gain vector is obtained based on a first encoding parameter that can be parsed from the bitstream.
[0301] The present disclosure provides an efficient method for encoding images or videos with different qualities and bitstream sizes without training more models.
Description of Signs
[0302] 10 video encoding system, 12 source device, 13 communication channel, 14 destination device, 16 picture source, 17 picture data, 18 pre-processor, 19 pre-processed picture data, 20 encoder, 21 encoded picture data, 22 communication interface, 28 communication interface, 30 decoder, 31 decoded picture data, 32 post-processor, 33 post-processed picture data, 34 display device, 40 video encoding system, 41 imaging device, 42 antenna, 43 processor, 44 memory store, 45 display device, 46 processing circuit, 101 encoder, 102 quantizer, 103 hyper-encoder, 104 decoder, 105 arithmetic encoder, 106 arithmetic decoder, 107 hyper-decoder, 108 quantizer, 109 arithmetic encoder, 110 arithmetic decoder, 121 encoder, 122 quantizer, 123 hyper-encoder, 125 arithmetic encoder, 127 hyper-decoder, 144 decoder, 146 arithmetic decoder, 147 hyper-decoder 401 Downsampling layer, 402 Downsampling layer, 403 Downsampling layer, 404 Downsampling layer, 405 Downsampling layer, 406 Downsampling layer, 407 Upsampling layer, 408 Upsampling layer, 409 Upsampling layer, 410 Upsampling layer, 411 Upsampling layer, 412 Upsampling layer, 413 Quantizer, 414 Input image, 415 Quantizer, 420 Convolution layer, 430 Convolution layer, 1001 Encoder, 1002 Quantizer, 1003 Hyper-encoder, 1004 Decoder, 1005 Arithmetic encoder, 1006 Arithmetic decoder, 1007 Hyper-decoder, 1008 Quantizer, 1009 Arithmetic encoder, 1010 Arithmetic decoder, 1011 Gain unit, 1012 Gain unit, 1013 Inverse gain unit, 1014 Inverse gain unit, 1300 Encoding device, 1310 Feature map acquisition module, 1320 Gain vector acquisition module, 1330 Gain unit, 1340 Quantization module, 1350 Entropy encoding module, 1400 Decoding device, 1410 Entropy decoding module, 1420 Inverse gain vector acquisition module, 1430 Inverse gain unit, 1440 Image reconstruction module, 8000 Video encoding device, 8010 Ingress port, 8020 Receiver unit, 8030 Processor, 8040 Transmitter unit, 8050 Egress port, 8060 Memory, 8070 Neural network-based codec, 9000 Device, 9002 Processor, 9004 Memory, 9006 Data, 9008 Operating system, 9010 Application program, 9012 Bus, 9018 Display
Claims
1. An image encoding method, comprising: a step of acquiring an image; a step of acquiring a first encoding parameter of the image, wherein a value of the first encoding parameter is smaller than a preset minimum value or larger than a preset maximum value; a step of acquiring a target gain vector based on the first encoding parameter; and a step of encoding the image based on the target gain vector. The method according to claim 1, further comprising:
2. When the value of the first encoding parameter is smaller than the preset minimum value, the step of acquiring the target gain vector based on the first encoding parameter comprises: a step of acquiring the target gain vector based on the first encoding parameter, the preset minimum value, and a first gain vector corresponding to the preset minimum value. The method according to claim 1, further comprising:
3. The step of acquiring the target gain vector based on the first encoding parameter, the preset minimum value, and the first gain vector comprises: a step of acquiring a first ratio of the first encoding parameter to the preset minimum value; and a step of acquiring the target gain vector based on the first ratio and the first gain vector. The method according to claim 2, further comprising:
4. The step of acquiring the target gain vector based on the first ratio and the first gain vector comprises: a step of multiplying the first ratio by the first gain vector to acquire the target gain vector. The method according to claim 3, further comprising:
5. The target gain vector satisfies the following condition: 【Number 1】 Here, m v is the target gain vector, β s is the preset minimum value, β v is the first encoding parameter, m s is the first gain vector corresponding to the preset minimum value, and K is a preset value The method according to any one of claims 2 to 4.
6. When the value of the first encoding parameter is larger than the preset maximum value, the step of acquiring the target gain vector based on the first encoding parameter comprises: a step of acquiring the target gain vector based on the first encoding parameter, the preset maximum value, and a second gain vector corresponding to the preset maximum value. The method according to claim 1, further comprising:
7. The step of acquiring the target gain vector based on the first encoding parameter, the preset maximum value, and the second gain vector comprises: Obtaining a second ratio of the first encoding parameter to the preset maximum value; Obtaining the target gain vector based on the second ratio and the second gain vector The method according to claim 6, comprising:
8. The step of obtaining the target gain vector based on the second ratio and the second gain vector comprises: Multiplying the second ratio by the second gain vector to obtain the target gain vector The method according to claim 7, comprising:
9. The target gain vector satisfies the following conditions 【Number 2】 Here, m v is the target gain vector, β t is the preset maximum value, β v is the first encoding parameter, m t is the second gain vector corresponding to the preset maximum value, and K is a preset value. The method according to any one of claims 6 to 8.
10. When the value of the first encoding parameter is smaller than the preset minimum value, the step of obtaining the target gain vector based on the first encoding parameter comprises: Obtaining the target gain vector based on the first encoding parameter, the preset minimum value, the first gain vector corresponding to the preset minimum value, the third preset value closest to the preset minimum value, and the third gain vector corresponding to the third preset value The method according to claim 1, comprising:
11. When the value of the first encoding parameter is greater than the preset maximum value, the step of obtaining the target gain vector based on the first encoding parameter comprises: Obtaining the target gain vector based on the first encoding parameter, the preset maximum value, the second gain vector corresponding to the preset maximum value, the fourth preset value closest to the preset maximum value, and the fourth gain vector corresponding to the fourth preset value The method according to claim 1, comprising:
12. The step of obtaining the target gain vector based on the first encoding parameter comprises: Obtaining the target gain vector based on the first encoding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values include the preset minimum value and / or the preset maximum value The method according to claim 1, comprising:
13. The step of encoding the image based on the target gain vector is Obtaining a first feature map of the image using a neural network; Obtaining a second feature map based on the first feature map and the target gain vector; Quantizing the second feature map to obtain a quantized second feature map; Encoding the quantized second feature map to obtain a bitstream The method according to any one of claims 1 to 12, comprising:
14. The step of obtaining a second feature map based on the first feature map and the target gain vector comprises: Multiplying the target gain vector by the first feature map The method according to claim 13, comprising:
15. The first feature map is a tensor having a shape of w×h×d, the target gain vector is a vector having a dimension of 1×d, w and h represent the width and height of the first feature map, and d represents the number of channels of the first feature map. The method according to claim 13 or 14.
16. When the first feature map is a feature map of the luma samples of the image, d is equal to 128, or when the first feature map is a feature map of the chroma samples of the image, d is equal to 64. The method according to claim 15.
17. The method further comprises obtaining a second encoding parameter of the image, When the first encoding parameter is used to encode the luma samples of the image, the second encoding parameter is used to encode the chroma samples of the image, or When the first encoding parameter is used to encode the chroma samples of the image, the second encoding parameter is used to encode the luma samples of the image, The method according to any one of claims 1 to 16.
18. The method comprises: Encoding the first encoding parameter into a bitstream The method according to any one of claims 1 to 17, further comprising:
19. The first encoding parameter is encoded in a picture parameter set (PPS) of the bitstream. The method according to claim 18.
20. The method according to claim 18 or 19, wherein the number of bits used to signal the first encoding parameter in the bitstream is 16 or less.
21. An image decoding method, comprising: obtaining a bitstream including encoded image data; analyzing the bitstream to obtain a first encoding parameter, wherein the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value; obtaining a target inverse gain vector based on the first encoding parameter; obtaining an image based on the target inverse gain vector The method includes.
22. When the value of the first encoding parameter is less than the preset minimum value, the step of obtaining a target inverse gain vector based on the first encoding parameter is: obtaining the target inverse gain vector based on the first encoding parameter, the preset minimum value, and a first gain vector corresponding to the preset minimum value The method according to claim 21, including.
23. The step of obtaining the target inverse gain vector based on the first encoding parameter, the preset minimum value, and the first gain vector includes: obtaining a first ratio of the first encoding parameter to the preset minimum value; obtaining the target inverse gain vector based on the first ratio and the first gain vector The method according to claim 22, including.
24. The step of obtaining the target inverse gain vector based on the first ratio and the first gain vector includes: multiplying the first ratio and the first gain vector to obtain a target gain vector; obtaining the target inverse gain vector based on the target gain vector The method according to claim 23, including.
25. The target gain vector satisfies the following conditions, [Number 3] 、 Here, m v is the target gain vector, β s is the preset minimum value, β v is the first coding parameter, m s is the first gain vector corresponding to the preset minimum value, and K is a preset value The method according to claim 24.
26. When the value of the first encoding parameter is greater than a preset maximum value, the step of obtaining a target inverse gain vector based on the first encoding parameter is: obtaining the target inverse gain vector based on the first encoding parameter, the preset maximum value, and a second gain vector corresponding to the preset maximum value The method according to claim 21, comprising: **Claim 27** The step of obtaining the target inverse gain vector based on the first encoding parameter, the preset maximum value, and the second gain vector comprises: obtaining a second ratio of the first encoding parameter to the preset maximum value; and obtaining the target inverse gain vector based on the second ratio and the second gain vector The method according to claim 26, comprising: **Claim 28** The step of obtaining the target inverse gain vector based on the second ratio and the second gain vector comprises: multiplying the second ratio and the second gain vector to obtain a target gain vector; and obtaining the target inverse gain vector based on the target gain vector The method according to claim 27, comprising: **Claim 29** The target gain vector satisfies the following condition: 【Number 4】 Here, m v is the target gain vector, β t is the preset maximum value, β v is the first coding parameter, m t is the second gain vector corresponding to the preset maximum value, and K is a preset value. The method according to claim 27. **Claim 30** The target inverse gain vector satisfies the following condition: 【Number 5】 wherein 【Number 6】 is the target inverse gain vector, m v is the target gain vector, C is a vector whose elements are all constants, and * means element-wise multiplication operation The method according to any one of claims 24 to 25 or any one of claims 28 to 29. **Claim 31** When the value of the first encoding parameter is smaller than a preset minimum value, the step of obtaining a target inverse gain vector based on the first encoding parameter comprises: obtaining the target inverse gain vector based on the first encoding parameter, the preset minimum value, a first gain vector corresponding to the preset minimum value, a third preset value closest to the preset minimum value, and a third gain vector corresponding to the third preset value The method according to claim 21, comprising: **Claim 32** When the value of the first encoding parameter is greater than a preset maximum value, the step of obtaining a target inverse gain vector based on the first encoding parameter comprises: a step of obtaining the target inverse gain vector based on the first encoding parameter, the preset maximum value, the second gain vector corresponding to the preset maximum value, the fourth preset value closest to the preset maximum value, and the fourth gain vector corresponding to the fourth preset value The method according to claim 21, comprising this step.
33. The step of obtaining a target inverse gain vector based on the first encoding parameter is a step of obtaining the target inverse gain vector based on the first encoding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values include the preset minimum value and / or the preset maximum value The method according to claim 21, comprising this step.
34. The step of decoding the bitstream to obtain an image based on the target inverse gain vector is a step of analyzing the bitstream to obtain a first latent representation of the image using entropy decoding; a step of obtaining a second latent representation based on the first latent representation and the target inverse gain vector; a step of decoding the second latent representation to obtain the image using a neural network The method according to any one of claims 21 to 33, comprising these steps.
35. The step of obtaining a second latent representation based on the first latent representation and the target inverse gain vector is a step of multiplying the target inverse gain vector and the first latent representation The method according to claim 34, comprising this step.
36. The method according to claim 34 or 35, wherein the first latent representation is a tensor having a shape of w×h×d, the target inverse gain vector is a vector having a dimension of 1×d, w and h represent the width and height of the first feature map, and d represents the number of channels of the first feature map.
37. The method according to claim 36, wherein d is equal to 128 when the first latent representation is a feature map of the luminance samples of the image, or d is equal to 64 when the first latent representation is a feature map of the chrominance samples of the image.
38. The method further includes analyzing the bitstream to obtain a second encoding parameter, when the first encoding parameter is used to decode the luma samples of the image, the second encoding parameter is used to decode the chroma samples of the image, or when the first encoding parameter is used to decode the chroma samples of the image, the second encoding parameter is used to decode the luma samples of the image, The method according to any one of claims 21 to 37.
39. A computer program product including program code for performing the method according to any one of claims 1 to 38 when executed on one or more processors.
40. A computer-readable storage medium storing a computer program executable by one or more processors, wherein when the computer program is executed by the one or more processors, the one or more processors perform the method according to any one of claims 1 to 38.
41. An encoder comprising a processing circuit for performing the method according to any one of claims 1 to 38.
42. A storage medium storing a bitstream obtained using the method according to any one of claims 1 to 20.
43. An encoded bitstream including encoded image data and a plurality of syntax elements, the plurality of syntax elements including a flag used to derive a value of an encoding parameter indicating the compression quality of the encoded image data, the value of the encoding parameter being less than a preset minimum value or greater than a preset maximum value.
44. An encoded bitstream comprising encoded image data and a plurality of syntax elements, wherein the plurality of syntax elements include a first flag and a second flag, the first flag being used to derive a value of a first encoding parameter indicating a compression quality of luma samples of the encoded image data, and the second flag being used to derive a value of a second encoding parameter indicating a compression quality of chroma samples of the encoded image data.
45. An encoding device, a memory including instructions, a processor coupled to the memory, the processor being configured to execute the instructions to cause the encoding device to perform the encoding method according to any one of claims 1 to 20. An encoding device comprising the same.
46. A decoding device, a memory including instructions, a processor coupled to the memory, the processor being configured to execute the instructions to cause the decoding device to perform the decoding method according to any one of claims 21 to 38. A decoding device comprising the same.
47. An encoding apparatus, a receiver unit configured to receive a picture to be encoded or a bitstream to be decoded, a transmitter unit coupled to the receiver unit, the transmitter unit being configured to transmit the bitstream to a decoder or to transmit a decoded image to a display, a memory coupled to at least one of the receiver unit or the transmitter unit, the memory being configured to store instructions, a processor coupled to the memory, the processor being configured to execute the instructions stored in the memory to perform the method according to any one of claims 1 to 20 or any one of claims 21 to 38. An encoding apparatus comprising the same.
48. A feature map acquisition module configured to acquire a first feature map from an input image, a gain unit configured to transform the first feature map based on a target gain vector to acquire a second feature map. A quantization module configured to quantize the second feature map to obtain a quantized second feature map; An entropy encoding module configured to encode the quantized second feature map to obtain a bitstream; An encoding device comprising the same.
49. An entropy decoding module configured to analyze a bitstream to obtain a first latent representation of an image using entropy decoding; An inverse gain unit configured to obtain a second latent representation based on the first latent representation and a target inverse gain vector; An image reconstruction module configured to decode the second latent representation to obtain the image using a neural network; A decoding device comprising the same.
50. An encoder; A decoder communicating with the encoder, wherein the encoder or the decoder includes the decoding device, encoding device, or encoding apparatus according to any one of claims 45 to 49; a decoder; An encoding system comprising the same.