Methods and apparatus for image coding and decoding

JP7901190B2Active Publication Date: 2026-08-05HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-06-30
Publication Date
2026-08-05

Smart Images

  • Figure 0007901190000077
    Figure 0007901190000077
  • Figure 0007901190000078
    Figure 0007901190000078
  • Figure 0007901190000079
    Figure 0007901190000079
Patent Text Reader

Abstract

The present disclosure relates to an image encoding method and device for a system using a neural network. A first encoding parameter is used to control the image compression quality when encoding an image, and the value of the first encoding parameter is smaller than a preset minimum value or larger than a preset maximum value. In particular, a target gain vector is obtained based on the first encoding parameter, and then the target gain vector is used to encode the image to obtain a bitstream. The first encoding parameter is signaled in the bitstream and transmitted to the decoding side so that the bitstream can be correctly decoded. Therefore, the method enables flexible setting of the compression quality or the bitstream size without training a new encoding model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of this disclosure generally relate to the field of encoding and decoding that are database-driven on neural network architectures. In particular, some embodiments relate to methods and apparatus for such encoding and decoding of images and / or videos from bitstreams using multiple processing layers. [Background technology]

[0002] Hybrid image and video codecs have been used for decades to compress image and video data. In such codecs, signals are typically encoded in blocks by predicting blocks and further encoding only the difference between the original blocks and their predictions. In particular, such encoding may involve transformation, quantization, and generation of the bitstream, usually including some form of entropy coding. Typically, the three components of a hybrid encoding method (transformation, quantization, and entropy coding) are optimized separately. Modern video compression standards such as High Efficiency Video Coding (HEVC), Versatile Video Coding (VVC), and Essential Video Coding (EVC) also use transform representations to encode the residual signal after prediction.

[0003] In recent years, neural network architectures have been applied to image and / or video coding. Generally, these neural network (NN)-based methods can be applied to image and video coding in a variety of different ways. For example, several end-to-end optimized image or video coding frameworks have been considered. Furthermore, deep learning has been used to determine or optimize some parts of the end-to-end coding framework, such as the selection or compression of prediction parameters. Also, several neural network-based methods have been considered for use in hybrid image and video coding frameworks, for example, as trained deep learning models for intra or inter-prediction in image or video coding.

[0004] The end-to-end optimized image or video coding applications described above have in common that they generate some feature map data that should be transmitted between the encoder and the decoder.

[0005] A neural network is a machine learning model that uses one or more layers of nonlinear units, and based on this, the machine learning model can predict the output of an incoming input. Some neural networks include one or more hidden layers in addition to the output layer. A corresponding feature map may be provided as the output of each hidden layer. Such a corresponding feature map of each hidden layer can be used as input to subsequent layers in the network, i.e., subsequent hidden layers or output layers. Each layer of the network generates an output from an incoming input according to the current values ​​of its respective set of parameters. In a neural network that is partitioned between devices, for example between an encoder and a decoder, between a device and the cloud, or between different devices, the feature map at the partition location (e.g., the first device) is compressed and transmitted to the remaining layers of the neural network (e.g., the second device).

[0006] Typically, a model is trained for a desired compression quality or bitstream size, and if a different compression quality is required, a new model needs to be trained, which requires a significant amount of time and computational cost. Furthermore, the amount of storage required increases with the number of models.

[0007] Further improvements in encoding and decoding using trained network architectures may be desirable. [Overview of the project]

[0008] This disclosure provides methods and apparatus for improving the flexibility of pre-trained image or video encoding models. Thus, the storage and transmission costs of encoded images or videos are reduced without sacrificing image quality, or distortion of reconstructed images is reduced without using more bits.

[0009] The aforementioned and other objectives are achieved by the subject matter of the independent claims. Further embodiments will become apparent from the dependent claims, description, and drawings.

[0010] Certain embodiments are outlined in the attached independent claims, and other embodiments are outlined in the dependent claims. [Means for solving the problem]

[0011] According to a first aspect, the disclosure relates to a method for image coding using a neural network. The method may be performed by an coding device. The method includes acquiring an image, acquiring a first coding parameter of the image such that the value of the first coding parameter is less than a preset minimum value or greater than a preset maximum value, acquiring a target gain vector based on the first coding parameter, and coding the image based on the target gain vector. The first coding parameter (denoted as β) is used to select a compression quality. A larger β results in a larger bitstream size and better quality of reconstructed data.

[0012] Such a method can give a pre-trained coding model more flexibility, as it allows encoding images at any desired compression quality or bitstream size without training a new coding model, especially when the compression quality indicated by the first coding parameter is outside the pre-training range. Using this method, a pre-trained model can be used to encode images at any desired quality, which can be flexibly deployed to various scenarios.

[0013] Each of the above-described and other embodiments may, individually or in combination, optionally include one or more of the following features:

[0014] In possible embodiments, a preset minimum value β s and a preset maximum value β t This is stored in the encoding device. During the training phase of the encoding model (e.g., encoder 101 in Figure 3A), [β s ,β t Several values ​​of β between [ ] are input to the model and trained to obtain the corresponding gain vector. Next, a pre-trained set of possible values ​​of β and the corresponding gain vectors is acquired and stored in both the encoding and decoding devices. Pre-set minimum value β s This is the lower limit of the range of β values, and the preset maximum value β. t This is the upper limit of the range of β values.

[0015] In possible embodiments, when the value of the first coding parameter is smaller than a preset minimum, obtaining a target gain vector based on the first coding parameter includes obtaining a target gain vector based on the first coding parameter, a preset minimum, and a first gain vector corresponding to the preset minimum.

[0016] Such a method enables encoding an image with a smaller bitstream size and a lower compression quality, which can meet the encoding requirements in poor communication network situations.

[0017] In a possible embodiment, a preset minimum value and a first gain vector corresponding to the preset minimum value are stored as a table in an encoding device or a remote device.

[0018] In a possible embodiment, the relationship between the preset minimum value and the first gain vector is stored in an encoding device or a remote device. The relationship may be in the form of a function or a network model.

[0019] In a possible embodiment, obtaining a target gain vector based on a first encoding parameter, a preset minimum value, and a first gain vector includes obtaining a first ratio of the first encoding parameter to the preset minimum value and obtaining a target gain vector based on the first ratio and the first gain vector.

[0020] In a possible embodiment, obtaining a target gain vector based on the first ratio and the first gain vector includes multiplying the first ratio and the first gain vector to obtain the target gain vector. The multiplication may be an element-wise multiplication operation.

[0021] In a possible embodiment, the target gain vector satisfies the following conditions:

Equation

[0022] In possible embodiments, when the value of the first coding parameter is greater than a preset maximum value, obtaining a target gain vector based on the first coding parameter includes obtaining a target gain vector based on the first coding parameter, a preset maximum value, and a second gain vector corresponding to the preset maximum value.

[0023] This method allows for encoding images with higher compression quality. This can meet the encoding requirements of some high-encoding-quality situations, such as encoding high-resolution movies or real-time transmission of sporting events.

[0024] In possible embodiments, a preset maximum value and a second gain vector corresponding to the preset maximum value are stored as a table in the encoding device or remote device.

[0025] In possible embodiments, a pre-set maximum value and its relationship to a second gain vector are stored in the encoding device or remote device. The relationship may take the form of a function or network model.

[0026] In possible embodiments, obtaining a target gain vector based on a first coding parameter, a preset maximum value, and a second gain vector includes obtaining a second ratio of the first coding parameter to the preset maximum value, and obtaining a target gain vector based on the second ratio and the second gain vector.

[0027] In possible embodiments, obtaining a target gain vector based on a second ratio and a second gain vector involves multiplying the second ratio and the second gain vector to obtain the target gain vector. The multiplication may be an element-wise multiplication operation.

[0028] In possible embodiments, the target gain vector satisfies the following conditions:

number

[0029] In possible embodiments, when the value of the first coding parameter is less than a preset minimum, obtaining a target gain vector based on the first coding parameter includes obtaining a target gain vector based on the first coding parameter, a preset minimum, a first gain vector corresponding to the preset minimum, a third preset value closest to the preset minimum, and a third gain vector corresponding to the third preset value. The third preset value is a pre-trained value from a pre-trained set of several values ​​of β that is closest to the preset minimum.

[0030] This method allows for encoding images with smaller bitstream sizes and relatively good compression quality by using two closest preset values ​​of encoding parameters and their corresponding gain vectors.

[0031] In possible embodiments, when the value of the first coding parameter is greater than a preset maximum value, obtaining a target gain vector based on the first coding parameter includes obtaining a target gain vector based on the first coding parameter, a preset maximum value, a second gain vector corresponding to the preset maximum value, a fourth preset value closest to the preset maximum value, and a fourth gain vector corresponding to the fourth preset value. The fourth preset value is a pre-trained value from a pre-trained set of several values ​​of β that is closest to the preset maximum value.

[0032] This method allows for encoding images with even higher compression quality by using two closest preset values ​​of the encoding parameters and their corresponding gain vectors.

[0033] In a possible embodiment, obtaining a target gain vector based on a first coding parameter includes obtaining a target gain vector based on a first coding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values ​​include a preset minimum and / or preset maximum.

[0034] Such methods can provide even greater coding performance by using more pre-trained β values ​​and their corresponding gain vectors.

[0035] In a possible embodiment of the first aspect and any one of its possible embodiments, encoding an image based on a target gain vector includes: obtaining a first feature map of the image using a neural network; obtaining a second feature map based on the first feature map and the target gain vector; quantizing the second feature map to obtain a quantized second feature map; and encoding the quantized second feature map to obtain a bitstream. For example, the quantized second feature map may be encoded using entropy coding.

[0036] In possible embodiments, obtaining a second feature map based on a first feature map and a target gain vector involves multiplying the target gain vector by the first feature map. The multiplication may be an element-wise multiplication operation.

[0037] In possible embodiments, the first feature map is a tensor having the shape w × h × d, and the target gain vector is a vector of dimension 1 × d, where w and h represent the width and height of the first feature map, and d represents the number of channels in the first feature map. The gain vector can be viewed as part of the model weights.

[0038] In a possible embodiment, when the first feature map is a feature map of a luma sample of the image, d is equal to 128.

[0039] In a possible embodiment, when the first feature map is a feature map of a chroma sample of an image, d is equal to 64.

[0040] In possible embodiments, the method further includes obtaining a second encoding parameter of an image, wherein when the first encoding parameter is used to encode a luma sample of the image, the second encoding parameter is used to encode a chroma sample of the image, or when the first encoding parameter is used to encode a chroma sample of the image, the second encoding parameter is used to encode a luma sample of the image. At least one of the values ​​of the first and second encoding parameters is less than a preset minimum value or greater than a preset maximum value.

[0041] In possible embodiments, the method further includes encoding a first encoding parameter into a bitstream. In possible designs, the first encoding parameter is encoded directly into the bitstream. In another possible design, the first encoding parameter is encoded into the bitstream as a first flag (e.g., the base-2 logarithm of the first encoding parameter) to conserve bits, and the first flag can be used to derive the first encoding parameter.

[0042] In possible embodiments, the first encoding parameter is encoded in the picture parameter set (PPS) of the bitstream.

[0043] In possible embodiments, the number of bits used to signal the first coding parameter in the bitstream is 16 or less. When the first flag is used to derive the first coding parameter, the number of bits used to signal the first coding parameter in the bitstream can be reduced to as few as 4.

[0044] In possible embodiments, the image includes both luminous and chroma samples, and a second encoding parameter is also signaled in the bitstream, such that when the first encoding parameter is used to encode the luminous sample of the image, the second encoding parameter is used to encode the chroma sample of the image, or when the first encoding parameter is used to encode the chroma sample of the image, the second encoding parameter is used to encode the luminous sample of the image. At least one of the values ​​of the first and second encoding parameters is less than a preset minimum value or greater than a preset maximum value.

[0045] According to a second aspect, the Disclosure relates to a method for decoding a bitstream to obtain an image. The method is performed by a decoding device. The method includes obtaining a bitstream containing encoded image data; analyzing the bitstream to obtain a first encoding parameter, wherein the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value; obtaining a target inverse gain vector based on the first encoding parameter; and obtaining an image based on the target inverse gain vector.

[0046] Such a method can give a pre-trained decoding model more flexibility, as it allows decoding of images at any desired compression quality or bitstream size without training a new decoding model, especially when the compression quality indicated by the first encoding parameter is outside the pre-training range. Using this method, a pre-trained model can be used to encode or decode images at any desired quality, which can be flexibly deployed to a variety of scenarios.

[0047] Each of the above-described and other embodiments may, individually or in combination, optionally include one or more of the following features:

[0048] In possible embodiments, a preset minimum value β s and a preset maximum value β t This is stored in the decoding device. During the training phase of the decoding model (e.g., decoder 104 in Figure 3A), [β s ,β t Several values ​​of β between [ ] are input to the model and trained to obtain the corresponding gain vector. Next, a pre-trained set of possible values ​​of β and the corresponding gain vectors is acquired and stored in both the encoding and decoding devices. Pre-set minimum value β s This is the lower limit of the range of β values, and the preset maximum value β. t This is the upper limit of the range of β values.

[0049] In possible embodiments, when the value of the first coding parameter is less than a preset minimum, obtaining a target inverse gain vector based on the first coding parameter includes obtaining a target inverse gain vector based on the first coding parameter, a preset minimum, and a first gain vector corresponding to the preset minimum.

[0050] This method allows for the decoding of images with smaller bitstream sizes and lower compression quality. This can meet encoding requirements in poor communication network conditions. Using this method, a pre-trained model can be used to encode or decode images at any desired quality, and this can be flexibly deployed to various scenarios.

[0051] In possible embodiments, a preset minimum value and a first gain vector corresponding to the preset minimum value are stored as a table in the decoding device or remote device.

[0052] In possible embodiments, a pre-set minimum value and its relationship to a first gain vector are stored in a decoding device or remote device. The relationship may take the form of a function or network model.

[0053] In possible embodiments, obtaining a target inverse gain vector based on a first coding parameter, a preset minimum value, and a first gain vector includes obtaining a first ratio of the first coding parameter to the preset minimum value, and obtaining a target inverse gain vector based on the first ratio and the first gain vector.

[0054] In possible embodiments, obtaining a target inverse gain vector based on a first ratio and a first gain vector includes multiplying the first ratio and the first gain vector to obtain the target gain vector, and obtaining a target inverse gain vector based on the target gain vector.

[0055] In possible embodiments, the target gain vector satisfies the following conditions:

number

[0056] In possible embodiments, when the value of the first coding parameter is greater than a preset maximum value, obtaining a target inverse gain vector based on the first coding parameter includes obtaining a target inverse gain vector based on the first coding parameter, a preset maximum value, and a second gain vector corresponding to the preset maximum value.

[0057] This method allows for decoding images encoded with even higher compression quality using two closest preset values ​​of the encoding parameters and their corresponding gain vectors. Using this method, a pre-trained model can be used to encode or decode images at any desired quality, and this can be flexibly deployed to various scenarios.

[0058] In possible embodiments, obtaining a target inverse gain vector based on a first coding parameter, a preset maximum value, and a second gain vector includes obtaining a second ratio of the first coding parameter to the preset maximum value, and obtaining a target inverse gain vector based on the second ratio and the second gain vector.

[0059] In possible embodiments, obtaining a target inverse gain vector based on a second ratio and a second gain vector includes multiplying the second ratio and the second gain vector to obtain the target gain vector, and obtaining a target inverse gain vector based on the target gain vector.

[0060] In possible embodiments, the target gain vector satisfies the following conditions:

number

[0061] In possible embodiments, the target inverse gain vector satisfies the following conditions:

number

[0062] In possible embodiments, when the value of the first coding parameter is less than a preset minimum, obtaining a target inverse gain vector based on the first coding parameter includes obtaining a target inverse gain vector based on the first coding parameter, a preset minimum, a first gain vector corresponding to the preset minimum, a third preset value closest to the preset minimum, and a third gain vector corresponding to the third preset value.

[0063] In possible embodiments, when the value of the first coding parameter is greater than a preset maximum value, obtaining a target inverse gain vector based on the first coding parameter includes obtaining a target inverse gain vector based on the first coding parameter, a preset maximum value, a second gain vector corresponding to the preset maximum value, a fourth preset value closest to the preset maximum value, and a fourth gain vector corresponding to the fourth preset value.

[0064] In possible embodiments, obtaining a target inverse gain vector based on a first coding parameter includes obtaining a target inverse gain vector based on a first coding parameter, N preset values, and N gain vectors corresponding to the N preset values, where N is an integer greater than 2, and the N preset values ​​include a preset minimum and / or preset maximum.

[0065] In possible embodiments, decoding a bitstream to acquire an image based on a target inverse gain vector includes analyzing the bitstream to acquire a first latent representation of the image using entropy decoding, acquiring a second latent representation based on the first latent representation and the target inverse gain vector, and decoding the second latent representation to acquire an image using a neural network.

[0066] In possible embodiments, obtaining a second latent representation based on a first latent representation and a target inverse gain vector involves multiplying the target inverse gain vector by the first latent representation. The multiplication may be an element-wise multiplication operation.

[0067] In possible embodiments, the first latent representation is a tensor of shape w × h × d, and the target inverse gain vector is a vector of dimension 1 × d, where w and h represent the width and height of the first feature map, and d represents the number of channels in the first feature map.

[0068] In a possible embodiment, when the first latent representation is a feature map of a luma sample of an image, d is equal to 128.

[0069] In possible embodiments, when the first latent representation is a feature map of a chroma sample of an image, d is equal to 64.

[0070] In possible embodiments, the method further comprises parsing a bitstream to obtain a second encoding parameter, the second encoding parameter being used to decode a chroma sample of an image when the first encoding parameter is used to decode a chroma sample of an image, or the second encoding parameter being used to decode a chroma sample of an image when the first encoding parameter is used to decode a chroma sample of an image.

[0071] The proposed method allows for the decoding of images using luminal and chromal samples compressed at different quality levels.

[0072] According to a third aspect, the disclosure relates to an apparatus / device for decoding an image or video. Such an apparatus for decoding may refer to the same advantageous effects as the method for decoding according to the second aspect. Further details are not described here. The decoding apparatus provides technical means for performing the actions in the method defined according to the second aspect. Its functions may be performed by hardware or by hardware running corresponding software. In possible embodiments, the decoding apparatus / device includes an entropy decoding module configured to analyze a bitstream to obtain a first latent representation of an image using entropy decoding; an inverse gain unit configured to obtain a second latent representation based on the first latent representation and a target inverse gain vector; and an image reconstruction module configured to decode the second latent representation to obtain an image using a neural network. These modules may be adapted to provide their respective functions corresponding to the method example according to the second aspect. Further details are referred to the detailed description in the method example. Further details are not described here.

[0073] In possible embodiments, the decoding device further comprises an inverse gain vector acquisition module configured to acquire a target inverse gain vector based on a first coding parameter.

[0074] According to a fourth aspect, the disclosure relates to an apparatus / device for encoding an image or video. Such an apparatus for encoding may refer to the same advantageous effects as the method for encoding according to the first aspect. Further details are not described here. The encoding apparatus provides technical means for performing the actions in the method defined according to the first aspect. Its functions may be performed by hardware or by hardware running corresponding software. In possible embodiments, the encoding apparatus includes a feature map acquisition module configured to acquire a first feature map from an input image; a gain unit configured to transform the first feature map based on a target gain vector to acquire a second feature map; a quantization module configured to quantize the second feature map to acquire a quantized second feature map; and an entropy coding module configured to encode the quantized second feature map to acquire a bitstream (e.g., using entropy coding). These modules may be adapted to provide the respective functions corresponding to the method example according to the first aspect. Further details are referred to the detailed description in the method example. Further details are not described here.

[0075] In possible embodiments, the encoding device may further include a gain vector acquisition module configured to acquire a target gain vector based on a first encoding parameter.

[0076] A method according to a first aspect of this disclosure may be performed by an apparatus according to a fourth aspect of this disclosure. Further features and embodiments of the method according to a first aspect of this disclosure correspond to the respective features and embodiments of the apparatus according to a fourth aspect of this disclosure. The advantages of the method according to a first aspect may be the same as the advantages of the corresponding embodiments of the apparatus according to a fourth aspect.

[0077] A method according to a second aspect of this disclosure may be performed by an apparatus according to a third aspect of this disclosure. Further features and embodiments of the method according to a second aspect of this disclosure correspond to the respective features and embodiments of the apparatus according to a third aspect of this disclosure. The advantages of the method according to a second aspect may be the same as the advantages of the corresponding embodiments of the apparatus according to a third aspect.

[0078] According to a fifth aspect, the disclosure relates to a video stream or image decoding device including a processor and memory. The memory stores instructions causing the processor to perform the method according to a second aspect.

[0079] According to a sixth aspect, the disclosure relates to a video stream or image encoding device including a processor and memory. The memory stores instructions causing the processor to perform the method according to the first aspect.

[0080] According to the seventh aspect, a computer-readable storage medium is proposed that stores instructions causing one or more processors to encode video or image data when executed. The instructions cause one or more processors to perform a method according to the first or second aspect or any possible embodiment of the first or second aspect.

[0081] According to the eighth aspect, the disclosure relates to a computer program product which includes program code for performing a method according to the first or second aspect or any possible embodiment of the first or second aspect when executed on a computer.

[0082] According to the ninth aspect, the present disclosure relates to a coder comprising a processing circuit for performing a method according to the first or second aspect or any possible embodiment of the first or second aspect.

[0083] According to the tenth aspect, the disclosure relates to a storage medium that stores a bitstream obtained using a method according to the first aspect or any possible embodiment of the first aspect.

[0084] According to the eleventh aspect, the present disclosure relates to a storage medium that stores a bitstream which can be decoded using a method according to the second aspect or any possible embodiment of the second aspect.

[0085] According to a twelfth aspect, the Disclosure relates to an encoded bitstream comprising encoded image data and a plurality of syntax elements, wherein the plurality of syntax elements comprises a first flag (such as compression_quality_level), the first flag indicating the compression quality of the encoded image data. The encoded bitstream may be obtained by performing a method according to a first aspect of the Disclosure or any possible embodiment of the first aspect. The encoded bitstream may be decoded by performing a method according to a second aspect of the Disclosure or any possible embodiment of the second aspect.

[0086] According to a thirteenth aspect, the Disclosure relates to an encoded bitstream comprising encoded image data and a plurality of syntax elements, wherein the plurality of syntax elements comprises a first flag (e.g., compression_quality_level_luma) and a second flag (e.g., compression_quality_level_chroma), the first flag indicating the compression quality of the luma sample of the encoded image data, and the second flag indicating the compression quality of the chroma sample of the encoded image data. The encoded bitstream may be obtained by performing a method according to a first aspect of the Disclosure or any possible embodiment of a first aspect. The encoded bitstream may be decoded by performing a method according to a second aspect of the Disclosure or any possible embodiment of a second aspect.

[0087] According to a 14th aspect, the disclosure relates to an encoding device comprising: a receiver unit configured to receive a picture to be encoded or a bitstream to be decoded; a transmitter unit coupled to the receiver unit, the transmitter unit configured to transmit a bitstream to a decoder or a decoded image to a display; a memory coupled to at least one of the receiver unit or the transmitter unit, the memory configured to store instructions; and a processor coupled to the memory, the processor configured to execute instructions stored in the memory in order to perform a method according to a first or second aspect or any possible embodiment of the first or second aspect.

[0088] According to the 15th aspect, the Disclosure relates to an encoding system comprising an encoder and a decoder that communicates with the encoder, wherein the encoder or decoder includes a decoding device according to the 3rd or 5th aspect of the Disclosure, an encoding device according to the 4th or 6th aspect of the Disclosure, or an encoding apparatus according to the 14th aspect of the Disclosure.

[0089] Details of one or more embodiments are described in the accompanying drawings and the following description. Other features, purposes, and advantages will become apparent from the description, drawings, and claims.

[0090] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying figures and drawings. [Brief explanation of the drawing]

[0091] [Figure 1] This is a schematic diagram showing the channels processed by the layers of a neural network. [Figure 2] This is a schematic diagram illustrating the autoencoder type of neural network. [Figure 3A]This is a schematic diagram illustrating an exemplary network architecture for the encoder and decoder sides, including a hyperplier model. [Figure 3B] This is a schematic diagram showing a typical network architecture on the encoder side, including a hyperplier model. [Figure 3C] This is a schematic diagram showing a typical network architecture on the decoder side, including a hyperplier model. [Figure 4] This is a schematic diagram illustrating an exemplary network architecture for the encoder and decoder sides, including a hyperplier model. [Figure 5] This is a block diagram showing the structure of a cloud-based solution for machine-based tasks such as machine vision tasks. [Figure 6A] A block diagram showing a neural network-based end-to-end video compression framework. [Figure 6B] This block diagram shows some illustrative details of the application of neural networks for motion field compression. [Figure 6C] This block diagram shows some illustrative details of the application of neural networks for motion compensation. [Figure 7] This is a flowchart illustrating an exemplary method for coding. [Figure 8] This is a schematic diagram showing the relationship between the first coding parameter and the gain vector. [Figure 9] This is a flowchart of an exemplary method for encoding an image based on a target gain vector. [Figure 10] This is a schematic diagram showing an exemplary network architecture for the encoder and decoder side, including the gain unit. [Figure 11] This is a flowchart illustrating an exemplary method for decoding. [Figure 12] This is a flowchart illustrating an exemplary method for decoding a bitstream based on a target gain vector. [Figure 13]This block diagram shows an example of a video encoding device configured to implement an encoding embodiment of the present disclosure. [Figure 14] This block diagram shows an example of a video decoding device configured to carry out the decoding embodiments of the present disclosure. [Figure 15] This block diagram shows an example of a video encoding system configured to implement embodiments of the present disclosure. [Figure 16] This block diagram shows another example of a video encoding system configured to implement embodiments of the present disclosure. [Figure 17] This is a block diagram showing examples of encoding or decoding devices. [Figure 18] This is a block diagram showing another example of an encoding or decoding device. [Modes for carrying out the invention]

[0092] The same reference number and symbol in different drawings may refer to the same element.

[0093] The following description refers to accompanying drawings that form part of the Disclosure and illustrate specific aspects of the embodiments of the Disclosure or specific aspects in which the embodiments of the Disclosure may be used. It is understood that the embodiments of the Disclosure may be used in other aspects and may include structural or logical modifications not shown in the drawings. Therefore, the following detailed description should not be constrained, and the scope of the Disclosure is defined by the accompanying claims.

[0094] For example, disclosures relating to a described method may also apply to a corresponding device or system configured to perform that method, and vice versa. For example, if one or more specific method steps are described, a corresponding device may include one or more units, e.g., functional units, to perform the described method step, even if one or more units are not explicitly described or illustrated (e.g., one unit performs one or more steps, or multiple units each perform one or more of the steps). Conversely, if a particular device is described based on one or more units, e.g., functional units, a corresponding method may include one step to perform the function of one or more units, even if one or more steps are not explicitly described or illustrated (e.g., one step performs the function of one or more units, or multiple steps each perform one or more of the functions of the units). Furthermore, it is understood that the various exemplary embodiments and / or features described herein may be combined with each other unless otherwise specified.

[0095] In the specification, claims, and accompanying drawings of this application, terms such as “first” and “second” are intended to distinguish similar subjects, but do not necessarily indicate a specific order or sequence. Terms used in this manner are interchangeable in appropriate contexts, and it should be understood that this is merely a method of distinction used when subjects having the same attributes are described in embodiments of this application. Furthermore, the terms “includes,” “having,” and any other variations thereof constitute non-exclusive inclusion, meaning that a process, method, system, product, or device including a set of units may include other units that are not necessarily limited to those units, are not expressly enumerated, or are specific to such a process, method, product, or device.

[0096] In the specification, claims, and accompanying drawings of this application, the term "and / or" is merely a relational relationship used to describe the related subjects. The term "and / or" indicates that three relationships may exist. For example, A and / or B may represent three cases: A alone exists, both A and B exist, and B alone exists.

[0097] The following provides an overview of some of the technical terms and frameworks used in embodiments of this disclosure.

[0098] Artificial neural networks Artificial neural networks (ANNs), or connectionist systems, are computing systems vaguely inspired by the biological neural networks that make up animal brains. Such systems "learn" to perform tasks by considering examples, without generally being programmed with task-specific rules. For example, in image recognition, they may learn to identify images containing cats by analyzing exemplary images manually labeled "cat" or "no cat," and using the results to identify cats in other images. They do this without any prior knowledge of cats, for example, that cats have fur, tails, whiskers, and cat-like faces. Instead, they automatically generate discriminative characteristics from the examples they process.

[0099] ANNs are based on a collection of connected units or nodes called artificial neurons, which roughly model the neurons of the biological brain. Each connection, like a synapse in the biological brain, can transmit signals to other neurons. The artificial neuron that receives the signal can then process it and send signals to the neurons it is connected to.

[0100] In an ANN embodiment, the "signals" in a connection are real numbers, and the output of each neuron is calculated by some nonlinear function of the sum of its inputs. These connections are called edges. Neurons and edges typically have weights that adjust as learning progresses. The weights increase or decrease the strength of the signal in the connection. Neurons may have thresholds such that they transmit a signal only if the aggregated signal exceeds that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (input layer) to the last layer (output layer), sometimes traversing multiple layers.

[0101] The initial goal of ANN methods was to solve problems in the same way the human brain solves them. Over time, attention shifted to performing specific tasks, leading to deviations from biology. ANNs are used in a variety of tasks, including computer vision, speech recognition, machine translation, social network filtering, playing board and video games, medical diagnosis, and even in activities traditionally thought to be left to humans, such as drawing.

[0102] The name "Convolutional Neural Network" (CNN) indicates that the network uses a mathematical operation called convolution. Convolution is a special type of linear operation. A convolutional network is a neural network that uses convolution instead of general matrix multiplication in at least one of its layers.

[0103] Figure 1 schematically illustrates the general concept of processing by neural networks such as CNNs. A convolutional neural network consists of an input layer, an output layer, and several hidden layers. The input layer is the layer to which the input (such as a portion of an image as shown in Figure 1) is provided for processing. The hidden layers of a CNN typically consist of a series of convolutional layers that convolve by multiplication or other dot products. The result of the layers is one or more feature maps (f.maps in Figure 1), sometimes called channels. Subsampling may exist in some or all of the layers. As a result, the feature maps may be small, as shown in Figure 1. The activation function in a CNN is usually a ReLU (Normalized Linear Unit) layer, followed by further convolutions such as pooling layers, fully connected layers, and normalization layers, which are called hidden layers because their inputs and outputs are masked by the activation function and the final convolution. These layers are colloquially called convolutions, but this is merely a convention. Mathematically, it is technically a sliding dot product or cross-correlation. This is important for matrix indices in that it affects how weights are determined at specific index points.

[0104] When programming a CNN to process images, the input is a tensor of the shape (number of images) × (image width) × (image height) × (image depth), as shown in Figure 1. It should be known that the image depth can be composed of the channels of the image. After passing through the convolutional layer, the image is abstracted into a feature map of the shape (number of images) × (feature map width) × (feature map height) × (feature map channels). The convolutional layer in the neural network should have the following attributes: a convolution kernel (hyperparameter) defined by width and height; the number of input and output channels (hyperparameters); and the depth of the convolutional filter (input channels) must be equal to the number of channels (depth) of the input feature map.

[0105] In the past, conventional multilayer perceptron (MLP) models have been used for image recognition. However, due to the full connectivity between nodes, they suffered from high dimensionality and did not scale well with higher resolution images. A 1000x1000 pixel image with RGB color channels has 3 million weights, which is too high to be feasible for large-scale, efficient processing with full connectivity. Furthermore, such network architectures do not take into account the spatial structure of the data, treating distant input pixels as equally as nearby pixels. This ignores the locality of reference in image data, both computationally and semantically. Therefore, for purposes such as image recognition, which depend on spatially local input patterns, full connectivity of neurons is wasteful.

[0106] Convolutional neural networks (CNNs) are a biologically inspired variation of the multilayer perceptron, specifically designed to emulate the behavior of the visual cortex. These models mitigate the challenges posed by MLP architectures by leveraging the strong spatially local correlations present in natural images. The convolutional layer is the core building block of a CNN. The layer's parameters consist of a set of learnable filters (kernels mentioned above), which have small receptive fields but extend to the full depth of the input volume. During the forward pass, each filter is convolved across the width and height of the input volume, calculating the dot product between the filter and the input entries to generate a two-dimensional activation map of that filter. As a result, the network learns which filters are activated when it detects some particular type of feature at some spatial location in the input.

[0107] Stacking the activation maps of all filters along the depth dimension forms the entire output volume of the convolutional layer. Therefore, all entries in the output volume can also be interpreted as the outputs of neurons that share parameters with neurons in the same activation map, focusing on small regions within the input. The feature map, or activation map, is the output activation of a given filter. Feature map and activation have the same meaning. In some papers, it is called an activation map because it is a mapping corresponding to the activations of different parts of an image, and also called a feature map because it is a mapping of where certain types of features can be found in the image. High activation means that a particular feature has been found.

[0108] Another important concept in CNNs is pooling, which is a form of nonlinear downsampling. There are several nonlinear functions for performing pooling, of which max pooling is the most common. It divides the input image into a set of non-overlapping rectangles and outputs the maximum value for each such sub-region.

[0109] Intuitively, the precise location of a feature is less important than its approximate location relative to other features. This is the idea behind the use of pooling in convolutional neural networks. Pooling layers progressively reduce the spatial size of representations, and also help reduce the number of parameters in the network, the memory footprint, and the computational complexity, and thus control overfitting. In CNN architectures, it is common to periodically insert pooling layers between consecutive convolutional layers. Pooling operations provide another form of translation invariance.

[0110] A pooling layer operates independently for each depth slice of the input, spatially resizing it. The most common form is a pooling layer where a 2x2 filter is applied with a stride of 2 across all depth slices of the input by 2 along both width and height, discarding 75% of the activations. In this case, all max operations exceed 4. The depth dimension remains invariant. In addition to max pooling, pooling units can use other functions such as mean pooling or l2 norm pooling. Mean pooling has historically been frequently used, but recently it has become less preferred compared to max pooling, which often performs better in practice. There is a recent trend to use smaller filters or discard the pooling layer altogether due to the aggressive reduction of representation size. "Region of Interest" pooling (also known as ROI pooling) is a variation of max pooling where the output size is fixed and the input rectangle is parameterized. Pooling is a key component of convolutional neural networks for object detection based on fast R-CNN architectures.

[0111] The ReLU mentioned above is an abbreviation for Normalized Linear Unit, which applies a non-saturated activation function. It effectively removes negative values ​​from the activation map by setting negative values ​​to 0. It enhances the nonlinear properties of the decision function and the entire network without affecting the receptive field of the convolutional layer. Other functions, such as the saturated hyperbolic tangent and sigmoid functions or LeakyReLU, are also used to enhance nonlinearity. ReLU is often preferred over other functions because it trains neural networks several times faster without a significant penalty to generalized accuracy.

[0112] After several convolutional and max pooling layers, high-level inference in the neural network is performed by fully connected layers. The neurons in the fully connected layers have connections to all activations of the previous layer, as seen in typical (non-convolutional) artificial neural networks. Thus, their activations can be computed as affine transformations, followed by matrix multiplication and then bias offsets (vector addition of learning or fixed bias terms).

[0113] The "loss layer" (which includes the calculation of the loss function) determines how unfavorable the training should be in the deviation between the predicted (output) label and the actual label, and is usually the final layer of a neural network. Various loss functions can be used that are suitable for different tasks. Softmax loss is used to predict a single class from K mutually exclusive classes. Sigmoid cross-entropy loss is used to predict K independent probability values ​​in [0,1]. Euclidean loss is used for regression to real-valued labels.

[0114] In summary, Figure 1 illustrates the data flow in a typical convolutional neural network. First, the input image passes through a convolutional layer and is abstracted into a feature map containing several channels corresponding to the number of filters in the set of learnable filters in this layer. Next, the feature map is subsampled, for example, using a pooling layer, which reduces the dimensionality of each channel in the feature map. The data then reaches another convolutional layer, which may have a different number of output channels. As mentioned above, the number of input and output channels are hyperparameters of the layer. To establish network connectivity, these parameters need to be synchronized between two connected layers so that the number of input channels in the current layer equals the number of output channels in the previous layer. For the first layer processing input data, e.g., an image, the number of input channels is usually equal to the number of channels in the data representation (e.g., 3 channels for an RGB or YUV representation of an image or video, or 1 channel for a grayscale image or video representation).

[0115] Autoencoders and unsupervised learning An autoencoder is a type of artificial neural network used to learn efficient data coding in an unsupervised manner. A schematic diagram of it is shown in Figure 2. The purpose of an autoencoder is to learn a representation (encode) of a dataset, typically for dimensionality reduction, by training the network to ignore signal "noise". Along with the reduction side, the reconstruction side is learned, and the autoencoder attempts to generate a representation from the reduced encoding that is as close as possible to its original input, and therefore its name. In its simplest case, given one hidden layer, the encoder stage of the autoencoder takes an input x and maps it to h. h = σ(Wx + b)

[0116] This image h is typically called the sign, latent variable, or latent representation. Here, σ is an element-wise activation function, such as a sigmoid function or normalized linear unit. W is the weight matrix, and b is the bias vector. The weights and biases are typically initialized randomly and then updated iteratively during training by backpropagation. The decoder stage of the autoencoder then maps h to a reconstructed x' with the same shape as x. x'=σ'(W'h'+b') Here, the decoder's σ', W', and b' may be independent of the corresponding σ, W, and b of the encoder.

[0117] Variational autoencoder models make strong assumptions about the distribution of latent variables. They use variational methods for latent representation learning, resulting in a further loss component and a specific estimator for the training algorithm called the Stochastic Gradient Variational Bayes (SGVB) estimator. It is used when the data is directed to a graphical model p θ (x|h) generates the posterior distribution p θ Approximation q for (h|x) ΦAssuming that (h|x) is learned, where Φ and θ are the parameters of the encoder (recognition model) and decoder (generative model), respectively, the probability distribution of the latent vector of the VAE typically matches the probability distribution of the training data much more closely than a standard autoencoder. The purpose of the VAE takes the following forms:

number

[0118] Here, D KL This represents the Kullback-Leibler divergence. The prior values ​​for the latent variables are typically central isotropic multivariate Gaussian p. θ (h) is set to N(0,I). Generally, the shapes of the variational and likelihood distributions are chosen so that they are factored Gaussian distributions. q Φ (h|x)=N(ρ(x),ω 2 (x)I) p Φ (x|h)=N(μ(h),σ 2 (h)I) Here, ρ(x) and ω 2 (x) is the encoder output, and μ(h) and σ 2 (h) is the decoder output.

[0119] Recent advances in the field of artificial neural networks, particularly convolutional neural networks, have enabled researchers to pursue their interest in applying neural network-based techniques to image and video compression tasks. For example, End-to-end Optimized Image Compression using a network based on variational autoencoders has been proposed.

[0120] Therefore, data compression is considered a fundamental and well-studied problem in engineering, generally formulated with the objective of designing a code for a given discrete data ensemble with minimum entropy. This solution relies heavily on knowledge of the probabilistic structure of the data, and thus the problem is closely related to probabilistic source modeling. However, since all actual codes must have finite entropy, continuous-value data (such as vectors of image pixel intensity) must be quantized into a finite set of discrete values, which introduces errors.

[0121] In this context, known as the lossy compression problem, two competing costs must be traded off: the entropy (rate) of the discretized representation and the error (distortion) resulting from quantization. Different compression applications, such as data storage or transmission over channels with limited capacity, require different rate-distortion tradeoffs.

[0122] Joint optimization of rate and distortion is difficult. Without further constraints, the general problem of optimal quantization in high-dimensional spaces is unwieldy. For this reason, most existing image compression methods work by linearly transforming the data vector into a suitable continuous-value representation, independently quantizing its elements, and then encoding the resulting discrete representation using a reversible entropy code. This method is called transformative coding due to the central role of the transformation.

[0123] For example, JPEG uses the discrete cosine transform for blocks of pixels, while JPEG2000 uses multiscale orthogonal wavelet decomposition. Typically, the three components of a transform coding method (transform, quantizer, and entropy code) are optimized separately (often by manual parameter tuning). Modern video compression standards such as HEVC, VVC, and EVC also use transform representations to encode the residual signal after prediction. Several transforms are used for this purpose, including the discrete cosine and sine transforms (DCT, DST) and the low-frequency unseparated manual-optimized transform (LFNST).

[0124] Variational image compression The Variable Autoencoder (VAE) framework can be considered a nonlinear transform coding model. The transform process can be divided into four main parts, which are illustrated in Figure 3A, which shows the VAE framework.

[0125] The transformation process can be divided into four main parts, and Figure 3A illustrates the VAE framework. In Figure 3A, the encoder 101 maps the input image x to a latent representation (denoted by y) via the function y=f(x). This latent representation will also be referred to below as a part or point in the “latent space”. The function f() is a transformation function that converts the input signal x to a more compressible representation y. The quantizer 102,

number

number

number

[0126] The latent space can be understood as a compressed representation of data where similar data points are close to each other within the latent space. The latent space is useful for learning data features and finding simpler representations of the data for analysis. (Hyperplier 3 Quantized Latent Representation)

number

number

number

number

number

number

number

[0127] In Figure 3A, component AE105 is a quantized latent representation

number

number

number

number

[0128] Arithmetic decoding (AD) 106 is the reverse of the binarization process, where the binary numbers are converted back to sample values. Arithmetic decoding is provided by the arithmetic decoding module 106.

[0129] Please note that this disclosure is not limited to this particular framework. Furthermore, this disclosure is not limited to image or video compression and can also be applied to object detection, image generation, and recognition systems.

[0130] In Figure 3A, there are two interconnected subnetworks. In this context, a subnetwork is a logical division of a part of the overall network. For example, in Figure 3A, modules 101, 102, 104, 105, and 106 are called the "encoder / decoder" subnetwork. The "encoder / decoder" subnetwork is responsible for encoding (generating) and decoding (analyzing) the first bitstream, "bitstream 1". The second network in Figure 3A, comprising modules 103, 108, 109, 110, and 107, is called the "hyperencoder / decoder" subnetwork. The second subnetwork is responsible for generating the second bitstream, "bitstream 2". The two subnetworks have different purposes.

[0131] The first subnetwork is, • The transformation of the input image x to its latent representation y (this x is easier to compress) 101, • Quantizing latent expression y

number

number

number

[0132] The purpose of the second subnetwork is to obtain the statistical properties of the samples of "bitstream 1" (e.g., mean, variance, and correlation between samples of bitstream 1) so that the compression of bitstream 1 by the first subnetwork becomes more efficient. The second subnetwork generates a second bitstream, "bitstream 2," which contains the aforementioned information (e.g., mean, variance, and correlation between samples of bitstream 1).

[0133] The second network is a quantized latent representation

number

number

number

number

number

number

number

number

number

number

number

number

number

[0134] Figure 3A illustrates an example of a VAE (Variational Autoencoder), the details of which may differ in different embodiments. For example, in certain embodiments, additional components may exist to more efficiently obtain the statistical characteristics of samples in bitstream 1. In such one embodiment, there may be a context modeler that aims to extract cross-correlation information of bitstream 1. The statistical information provided by the second subnetwork may be used by components of the AE (Arithmetic Encoder) 105 and AD (Arithmetic Decoder) 106.

[0135] Figure 3A shows an encoder and decoder in a single diagram. As will be obvious to those skilled in the art, encoders and decoders may be incorporated into different devices, and very often they are incorporated into different devices.

[0136] Figure 3B shows the encoder, and Figure 3C shows the decoder components of the VAE framework separated. As input, the encoder receives a picture, according to some embodiments. The input picture may include one or more channels, such as a color channel or other types of channels, such as a depth channel or a motion information channel. The outputs of the encoder (as shown in Figure 3B) are bitstream 1 and bitstream 2. Bitstream 1 is the output of the encoder's first subnetwork, and bitstream 2 is the output of the encoder's second subnetwork.

[0137] Similarly, in Figure 3C, two bitstreams, namely bitstream 1 and bitstream 2, are received as input, and the reconstructed (decoded) image is

number

[0138] Specifically, as shown in Figure 3B, the encoder includes an encoder 121 that converts the input x into a signal y which is then provided to the quantizer 322. The quantizer 122 provides information to the arithmetic coding module 125 and the hyperencoder 123. The hyperencoder 123 provides the bitstream 2 already described above to the hyperdecoder 147, which then provides information to the arithmetic coding module 105(125).

[0139] The output of the arithmetic coding module is bitstream 1. Bitstreams 1 and 2 are the outputs of the coded signals, which are then supplied (transmitted) to the decoding process. Unit 101 (121) is called the “encoder,” but it is also possible to call the entire subnetwork described in Figure 3B the “encoder.” The coding process generally refers to a unit (module) that converts an input into a coded (e.g., compressed) output. From Figure 3B, it can be seen that unit 121 can actually be considered the center of the entire subnetwork, as it performs the conversion of input x to y, which is a compressed version of x. Compression in encoder 121 can be achieved, for example, by applying a neural network, or in general by applying any processing network having one or more layers. In such a network, compression can be performed by cascading processes that include downsampling, which reduces the size and / or number of channels of the input. Thus, the encoder can be called, for example, a neural network (NN) based encoder.

[0140] The rest of the diagram (quantization unit, hyperencoder, hyperdecoder, arithmetic encoder / decoder) are all responsible for improving the efficiency of the encoding process or converting the compressed output y into a sequence of bits (bitstream). Quantization may be provided to further compress the output of the NN encoder 121 by lossy compression. Combined with the hyperencoder 123 and hyperdecoder 127 used to construct the AE125, the AE125 can perform binarization, which can further compress the quantized signal by lossless compression. Thus, it is also possible to refer to the entire subnetwork in Figure 3B as the "encoder".

[0141] Most deep learning (DL)-based image / video compression systems reduce the dimensionality of a signal before converting it to binary (bits). For example, in the VAE framework, the encoder, which is a nonlinear transformation, maps the input image x to y, where y has smaller width and height than x. Because y has smaller width and height, it is therefore smaller in size, reducing the dimensionality (size) of the signal and thus making it easier to compress the signal y. It should be noted that, in general, encoders do not necessarily have to reduce the size of both (or generally all) dimensions. Rather, some exemplary embodiments may provide encoders that reduce the size in only one (or generally a subset) dimension.

[0142] In J. Balle, L. Valero Laparra, and EP Simoncelli (2015) ("Density Modeling of Images Using a Generalized Normalization Transformation", In: arXiv e-prints, Presented at the 4th Int.Conf. for Learning Representations, 2016) (hereinafter referred to as "Balle"), the authors proposed a framework for end-to-end optimization of image compression models based on nonlinear transformations. The authors optimize the mean squared error (MSE) but use a more flexible transformation constructed from a cascade of linear convolution and nonlinearity. Specifically, the authors use a generalized division-normalization (GDN) coupled nonlinearity, inspired by a model of neurons in the biological visual system, which has been proven effective for Gaussianization of image density. This cascaded transformation is followed by uniform scalar quantization (i.e., each element is rounded to the nearest integer), which effectively implements a parametric form of vector quantization in the original image space. The compressed image is reconstructed from these quantized values ​​using an approximate parametric nonlinear inverse transform.

[0143] Such an example of the VAE framework is shown in Figure 4, which utilizes six downsampling layers marked 401 - 406. The network architecture includes a hyperprior model. On the left (g a , g s ) represents an image autoencoder architecture, and on the right (h a , h s ) corresponds to an autoencoder that implements the hyperprior. The factored prior model uses the same architecture for the analysis and synthesis transforms g a and g s . Q represents quantization, and AE, AD represent an arithmetic encoder and an arithmetic decoder, respectively. The encoder feeds the input image x through g a to produce a response y (latent representation) with a spatially varying standard deviation. The encoding g a includes multiple convolutional layers with subsampling and generalized divisive normalization (GDN) as the activation function.

[0144] The response is fed to h a which aggregates the standard deviation distribution at z. Next, z is quantized, compressed, and transmitted as side information. Next, the encoder uses the quantized vector

Number

Number

Number

Number

number

number

number

[0145] Layers that include downsampling are indicated by downward arrows in the layer descriptions. The layer description "Conv Nx5x5 / 2↓" means that the layer is a convolutional layer with N channels and a convolutional kernel size of 5x5. As stated, 2↓ means that 2x downsampling is performed in this layer. 2x downsampling results in one of the dimensions of the input signal being reduced by half in the output. In Figure 4, 2↓ indicates that both the width and height of the input image are reduced by 2x. Because there are six downsampling layers, if the width and height of the input image 4¹⁴ (also indicated by x) are given by w and h, the output signal z^4¹⁴ will have a width and height equal to w / 64 and h / 64, respectively. The modules indicated by AE and AD are arithmetic encoders and arithmetic decoders, which are described with reference to Figures 3A-3C. Arithmetic encoders and decoders are specific embodiments of entropy coding. AE and AD can be replaced by other means of entropy coding. In information theory, entropy coding is a lossless data compression method, a lossless process used to convert the values ​​of symbols into binary representations. The "Q" in the figure corresponds to the quantization operation, also referenced above in relation to Figure 4, and is further explained above in the "Quantization" section. Furthermore, the quantization operation and its corresponding quantization unit are not necessarily part of component 413 or 415, and / or can be replaced by another unit.

[0146] Figure 4 also shows a decoder with upsampling layers 407-412. A further layer 420, implemented as a convolutional layer but not providing upsampling to the received input, is placed between the upsampling layers 411 and 410 in the order of input processing. A corresponding convolutional layer 430 is also shown for the decoder. Such layers can be placed in the NN to perform operations on the input that modify specific characteristics without changing the size of the input. However, such layers are not required.

[0147] When viewed in the processing order of bitstream 2 passing through the decoder, the upsampling layers pass in the reverse order, i.e., from upsampling layer 412 to upsampling layer 407. Each upsampling layer is shown here to provide upsampling with an upsampling ratio of 2, indicated by ↑. Of course, not all upsampling layers necessarily have the same upsampling ratio, and other upsampling ratios such as 3, 4, or 8 may be used. Layers 407-412 are implemented as convolutional layers (conv). Specifically, they may be intended to provide an operation on the input that is the inverse of the encoder operation, so the upsampling layers may apply an inverse convolution operation to the received input such that the size of the received input is increased by a coefficient corresponding to the upsampling ratio. However, this disclosure is not generally limited to inverse convolution, and upsampling may be performed in any other way, such as by bilinear interpolation between two adjacent samples or nearest neighbor sample copying.

[0148] In the first subnetwork, several convolutional layers (401-403) are followed by generalized division normalization (GDN) on the encoder side and inverse GDN (IGDN) on the decoder side. In the second subnetwork, the activation function applied is ReLU. It should be noted that this disclosure is not limited to such embodiments, and in general, other activation functions may be used instead of GDN or ReLU.

[0149] Cloud solutions for machine tasks Video coding for machines (VCM) is another direction in computer science that is gaining popularity today. The main idea behind this technique is to transmit coded representations of image or video information that are intended for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. In contrast to conventional image and video coding aimed at human perception, the quality characteristic is not reconstruction quality, but rather performance on computer vision tasks, such as object detection accuracy. This is illustrated in Figure 5.

[0150] Video coding for machines, also known as collaborative intelligence, is a relatively new paradigm for the efficient deployment of deep neural networks across mobile and cloud infrastructure. By splitting the network between mobile and the cloud, it is possible to distribute the computational workload so that the overall energy and / or latency of the system is minimized. In general, collaborative intelligence is a paradigm in which the processing of a neural network is distributed among two or more different computing nodes, e.g., devices, generally any functionally defined nodes. Here, the term “node” does not refer to the neural network nodes mentioned above. Rather, a (computational) node here refers to a separate device / module that implements a part of the neural network (physically or at least logically). Such devices may be a mixture of different servers, different end-user devices, servers and / or user devices and / or the cloud and / or processors. In other words, computing nodes can be thought of as nodes belonging to the same neural network that communicate with each other to transmit encoded data within / for the neural network. For example, to enable the execution of complex calculations, one or more layers may run on a first device, and one or more layers may run on another device. However, the distribution may be finer, and a single layer may run on multiple devices. In this disclosure, the term “multiple” refers to two or more. In some existing solutions, a portion of the neural network functionality runs on a device (such as a user device or edge device) or multiple such devices, and the output (feature map) is then passed to the cloud. The cloud is a collection of processing or computing systems located outside the devices running the portion of the neural network. The concept of collaborative intelligence has also been extended to model training.In this case, data flows in both directions: from the cloud to mobile during backpropagation in training, and from mobile to the cloud during the forward pass in training, and the same applies to inference.

[0151] Some works have presented semantic image compression by encoding deep features and then reconstructing the input image from them. Compression based on uniform quantization has been shown, followed by context-based adaptive arithmetic coding (CABAC) from H.264. In some scenarios, it may be more efficient to transmit the output of the hidden layer (deep feature map) from the mobile unit to the cloud than to send compressed natural image data to the cloud and perform object detection using the reconstructed image. Efficient compression of feature maps benefits the compression and reconstruction of images and videos for both human perception and machine vision. Entropy coding methods, such as arithmetic coding, are common techniques for compressing deep features (i.e., feature maps).

[0152] Today, video content accounts for over 80% of internet traffic, and this percentage is expected to increase further. Therefore, it is crucial to build efficient video compression systems that produce higher-quality frames within a given bandwidth budget. Furthermore, most video-related computer vision tasks, such as video object detection or video object tracking, are affected by the quality of compressed video, and efficient video compression can benefit other computer vision tasks. Meanwhile, video compression techniques are also useful for action recognition and model compression. However, over the past few decades, video compression algorithms have relied on handcrafted modules, such as block-based motion estimation and discrete cosine transform (DCT), to reduce the redundancy of video sequences, as mentioned above. While each module may be well-designed, the overall compression system is not end-to-end optimized. It is desirable to further improve video compression performance by collectively optimizing the entire compression system.

[0153] End-to-end image or video compression DNN-based image compression methods can utilize large-scale end-to-end training and highly non-linear transformations, which are not used in conventional methods. However, it is not straightforward to directly apply these techniques to build an end-to-end learning system for video compression. First, learning how to generate and compress motion information adjusted for video compression remains an unsolved problem. Video compression methods rely heavily on motion information to reduce the temporal redundancy of video sequences.

[0154] A simple solution is to use a learning-based optical flow to represent motion information. However, current learning-based optical flow methods aim to generate the flow field as accurately as possible. An accurate optical flow is often not optimal for specific video tasks. Additionally, the data volume of optical flow increases significantly compared to motion information in conventional compression systems, and directly applying existing compression techniques to compress optical flow values greatly increases the number of bits required to store motion information. Second, it is not clear how to build a DNN-based video compression system by minimizing rate-distortion-based objectives for both residual information and motion information. Rate-distortion optimization (RDO) aims to achieve a higher-quality (i.e., less distortion) reconstructed frame when a given number of bits (or bitrate) for compression is provided. RDO is important for video compression performance. To utilize the ability of end-to-end training for learning-based compression systems, an RDO strategy is required to optimize the entire system.

[0155] In "DVC: An End-to-end Deep Video Compression Framework". Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp.11006-11015, the authors proposed an end-to-end deep video compression (DVC) model that collaboratively learns motion estimation, motion compression, and residual coding.

[0156] Such an encoder is shown in Figure 6A. In particular, Figure 6A shows the overall structure of an end-to-end trainable video compression framework. To compress motion information, a CNN is specified to transform the optical flow into a corresponding representation suitable for better compression. Specifically, an autoencoder-type network is used to compress the optical flow. The motion vector (MV) compression network is shown in Figure 6B. The network architecture is somewhat similar to ga / gs in Figure 4. In particular, the optical flow is fed into a series of convolutional operations and nonlinear transformations, including GDN and IGDN. The number of output channels for convolution (deconvolution) is 128, except for the last deconvolutional layer, which is equal to 2. Given an optical flow of size M × N × 2, the MV encoder generates a motion representation of size M / 16 × N / 16 × 128. The motion representation is then quantized, entropy coded, and sent to a bitstream. The MV decoder receives the quantized representation and uses the MV encoder to reconstruct the motion information.

[0157] Figure 6C shows the structure of the motion compensation unit. Here, the previous reconstructed frame x t-1Using the reconstructed motion information, the warping unit generates warping frames (usually with the help of interpolation filters such as bilinear interpolation filters). Then, a separate CNN with three inputs generates the predicted picture. The architecture of the motion-compensated CNN is also shown in Figure 6C.

[0158] The residual information between the original frame and the predicted frame is encoded by a residual encoder network. A highly nonlinear neural network is used to convert the residuals into their corresponding latent representations. Compared to the discrete cosine transform in conventional video compression systems, this method can better utilize the capabilities of nonlinear transformations and achieve higher compression efficiency.

[0159] From the above overview, considering the different parts of the video framework, including motion estimation, motion compensation, and residual coding, it can be understood that CNN-based architectures can be applied to both image and video compression. Entropy coding is a common method used for data compression, widely adopted in the industry, and is also applicable to feature map compression for either human perception or computer vision tasks.

[0160] Video coding for machines Video coding for machines (VCM) is another direction in computer science that is gaining popularity today. The main idea behind this technique is to transmit coded representations of image or video information that are intended for further processing by computer vision (CV) algorithms, such as object segmentation, detection, and recognition. In contrast to traditional image and video coding aimed at human perception, the quality characteristic is not reconstruction quality, but rather performance on computer vision tasks, such as object detection accuracy.

[0161] Recent research has proposed a new deployment paradigm called collaborative intelligence, in which deep models are partitioned between mobile and cloud. Extensive experimentation under various hardware configurations and wireless connectivity modes has revealed that the optimal operating point in terms of energy consumption and / or computational latency typically involves partitioning the model at a point deep in the network. It has been found that current common solutions are rarely optimal when (if ever) the model resides entirely in the cloud or entirely on mobile. The concept of collaborative intelligence has also been extended to model training. In this case, data flows in both directions: from cloud to mobile during backpropagation in training, and from mobile to cloud during the forward pass in training, and similarly in inference.

[0162] In the context of recent deep models for object detection, lossy compression of deep feature data has been studied based on HEVC intra coding. Increasing the compression level and proposed compression extension training to minimize this loss by generating a model more robust to quantization noise in feature values ​​has resulted in a decrease in detection performance. However, this remains a suboptimal solution because the codecs used are highly complex and optimized for natural scene compression rather than deep feature compression.

[0163] The problem of deep feature compression for collaborative intelligence has been addressed by techniques for object detection tasks using the common YOLOv2 network for studying the trade-off between compression efficiency and recognition accuracy. Here, the term "deep feature" is synonymous with "feature map." The word "deep" comes from the idea of ​​collaborative intelligence where output feature maps from some hidden (deep) layer are captured and transferred to the cloud for inference. This is seen as more efficient than sending compressed natural image data to the cloud and performing object detection using the reconstructed images.

[0164] Efficient feature map compression benefits image and video compression and reconstruction for both human perception and machine vision. The shortcomings of state-of-the-art autoencoder-based compression techniques also apply to machine vision tasks. ...

[0165] Improved coding efficiency As mentioned above, variational autoencoders have become a state-of-the-art technique for learnable image compression. When they emerged in 2017, they only supported a single-quality mode (a trade-off between rate and distortion) until variable-rate methods appeared, and their models supported the selection of compression quality within a given specific range. At the heart of variable-rate models is the gain unit, which is part of a neural network-based codec that gives the model the ability to compress input data with different compression ranges. However, conventional gain units were designed to operate within a given range of compression quality. Conventional gain units cannot adapt to some specific bitstream sizes outside that given range of compression quality.

[0166] Accordingly, some embodiments of this disclosure introduce gain vector extrapolation solutions for autoencoders to enable flexible representation and transmission of content without being limited to a predetermined or pre-trained range of compression quality or bitstream size.

[0167] The following provides several detailed embodiments and examples related to the encoder and decoder sides.

[0168] Encoding method Figure 7 is a flowchart illustrating an exemplary method for encoding, which includes the steps of: receiving an image 701; receiving a first encoding parameter of the image 702, such that the value of the first encoding parameter is less than a preset minimum or greater than a preset maximum; obtaining a target gain vector based on the first encoding parameter 703; and encoding the image based on the target gain vector 704.

[0169] In step 701, an image is received. In embodiments of the Disclosure, the image is an image to be compressed, and the image may be an image captured by a camera by an encoding device (the encoding device performs method 700 to encode the image), or the image may also be an image acquired from within the encoding device (e.g., an image stored in the encoding device's album, or a picture acquired by the encoding device from the cloud or another device). The image may be a single image or a frame of a video. The image may also be a part of an image, i.e., an image block. The images referred to above may be images or image blocks having image compression requirements, and it should be understood that the Disclosure does not limit the source of the image to be processed. The image includes a lumina (Y) sample and / or a chroma (UV) sample.

[0170] In step 702, a first encoding parameter is received, the value of which is either less than a preset minimum or greater than a preset maximum. The first encoding parameter (e.g., denoted as β) is a scalar input for the encoding process used to select the compression quality. A larger β results in a larger bitstream size and better quality of reconstructed data. s (shown as) and a preset maximum value (for example, β t (shown as) is the range of pre-trained values ​​of the first coding parameter [β s ,β tRelated to ]. Pre-set maximum value β t This is the pre-set minimum value β s That concludes the explanation. In some embodiments, the range of pre-trained values ​​of the first coding parameter β may be discontinuous, and the values ​​of the first parameter β are finite discrete points. In the inference process, by applying interpolation, the values ​​of the first parameter β are β s and β t It can be extended to any value between [value] and [value].

[0171] In possible designs, the image includes both luma samples and chroma samples, and the first encoding parameter may include a first component and a second component. At least one of the first and second components of the first encoding parameter satisfies the condition that the value of the first encoding parameter is less than a preset minimum or greater than a preset maximum. Note that the preset minimum and preset maximum values ​​of the luma samples and chroma samples may be different. The first component may be used to encode the luma samples and thus control the compression quality of the luma samples in the image. The second component may be used to encode the chroma samples and thus control the compression quality of the chroma samples in the image.

[0172] In possible designs, the image includes both luminous and chroma samples, and a second encoding parameter is also obtained. One of the first and second encoding parameters is used to encode the luminous sample, and the other is used to encode the chroma sample. The chroma and luminous samples are encoded with different encoding parameters β to control the compression quality separately. The preset minimum and maximum values ​​for the luminous and chroma samples may be different or the same, and this is not limited to the present disclosure.

[0173] In embodiments of this disclosure, a preset minimum value β s and a preset maximum value β tThis may be stored in an encoding device or received from another device, and is not limited to this.

[0174] In embodiments of this disclosure, each value of the first coding parameter β corresponds to a gain vector, i.e., the model weights of the neural network model used to code the image. A pre-trained set of possible values ​​of β and the corresponding gain vectors is obtained during the model training phase. As described above, the possible values ​​of β are a preset minimum value β s and the pre-set maximum value β t It lies between [the two points]. In this disclosure, the gain vector is a vector having dimension 1 × d, where d is an integer greater than 1. In possible embodiments, when encoding lumens of an image, d = 128. In possible embodiments, when encoding chromens of an image, d = 64. Note that d may be any other integer, such as 32, 48, 96, 144, 160, 176, 192, 256, etc.

[0175] Therefore, in a possible embodiment, a pre-trained set of possible values ​​of β and their corresponding gain vectors is stored in the encoding and decoding devices. In another possible embodiment, only the pre-trained set of possible values ​​of β or their corresponding gain vectors is stored in the encoding and decoding devices, and the mapping relationships between the possible values ​​of β and their corresponding gain vectors are also stored in the encoding and decoding devices. The mapping relationships represent a one-to-one correspondence between possible values ​​of β and their gain vectors. The mapping relationships may be a pre-configured table, a pre-configured objective function, or any other form, as long as they are the same in the encoding and decoding devices, and this is not limited thereto.

[0176] In possible designs, luma samples and chroma samples are trained together. Therefore, luma samples and chroma samples share a list of possible values ​​for the first coding parameter. However, the gain vectors obtained after the training phase differ between luma samples and chroma samples. Table 1 is an exemplary table showing the pre-trained β and corresponding gain vectors for luma samples and chroma samples. The number of pre-trained coding parameters is N, where N is an integer greater than 0. In possible designs, N=13. In Table 1, β0 is the minimum value β. s It may be, β N-1 The maximum value is β t That's fine.

[0177] [Table 1]

[0178] In possible designs, luma samples and chroma samples are trained separately. Therefore, luma samples and chroma samples do not need to share a list of possible values ​​for the first coding parameter. Thus, two tables are used to store the pre-trained [β, gain vector] pairs for luma samples and chroma samples. Table 2 is an exemplary table showing the pre-trained β and corresponding gain vector for luma samples. Table 3 is an exemplary table showing the pre-trained β and corresponding gain vector for chroma samples. The number of pre-trained coding parameters for chroma samples is N1, where N1 is an integer greater than 0. In possible designs, N1 = 13. The number of pre-trained coding parameters for chroma samples is N2, where N2 is an integer greater than 0. In possible designs, N2 = 13. Note that N1 and N2 may be other integer values, such as N1 = 10, N2 = 5, etc. N1 may be greater than or equal to N2 or less than N2, and this is not limited to this disclosure. In Table 2, β0 is the minimum value β of the Luma sample. s It may be, β N1-1 This is the maximum value β of the luma sample. tThis may also be the case. In Table 3, β0 is the minimum value β of the chromatic sample. s It may be, β N2-1 This is the maximum value β of the chroma sample. t That's fine.

[0179] [Table 2]

[0180] [Table 3]

[0181] In step 703, the target gain vector is obtained based on the first coding parameter. As described above, in the pre-trained set of possible values ​​of β and the corresponding gain vectors, the possible values ​​of β are the preset minimum value β. s and the pre-set maximum value β t It lies between [the two points]. The drawback of this method is that the image compression quality of the pre-trained model is limited by some minimum and maximum values. Even if more than twice the bits are used, it is not possible to obtain a quality better than the maximum quality using existing methods. Even if a lower quality below the minimum quality is acceptable in some situations, compressing an image with fewer bits is not supported by existing methods. Therefore, in order to overcome the shortcomings of existing methods, an encoding method is proposed that supports encoding an image at an arbitrary image compression quality. When the target compression quality is given by a first encoding parameter, the corresponding target gain vector can be obtained using the proposed method of this disclosure even when the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value.

[0182] Figure 8 is a schematic diagram showing the relationship between the first coding parameter (i.e., network input parameter) β and the gain vector. During the training phase of a neural network (such as encoder 101 in Figure 3A), [β s ,βt Some values of β between are input into the network and trained to obtain the corresponding gain vectors. In FIG. 8, [β s ,m s , [β r ,m r , and [β t ,m t are shown, where three pairs of pre-trained [β, gain vectors] are shown, β s is the lower limit of the range of values of β, and β t is the upper limit of the range of values of β. The present disclosure is not limited to such embodiments. Generally, any other possible number of [β, gain vector] pairs may be pre-trained, and in many situations, more than three pairs of [β, gain vector] pairs are pre-trained. For example, in a possible design where 13 pairs of [β, gain vector] pairs are pre-trained, it should be noted that this is the case.

[0183] In the present disclosure, as shown in FIG. 8, an extrapolation method for obtaining a target gain vector based on a first encoding parameter is proposed. In the inference stage, when a specific value of the first encoding parameter (shown as β v ) is given, the target gain vector (shown as m v ) can be derived according to an objective function, that is, m v = f(β v ,β r ,β t ,β s ,m r ,m s ,m t ,…).

[0184] In a possible embodiment, only the closest pre-trained input parameter and its corresponding gain vector are used to derive the new gain vector m v of β v . For example, when β v is greater than a preset maximum value β t , only the maximum value β t and its corresponding gain vector m t are used to derive m vIt is used to derive m. v =f(β v ,β t ,m t ) is the function m v =f(β v ,β t ,m t One possible embodiment of ) is,

number

[0185] Similarly, β v The minimum value β is set in advance. s When it is smaller than, the minimum value β s and its corresponding gain vector m s Only, m v It is used to derive m. v =f(β v ,β s ,m s Similarly, the function m v =f(β v ,β s ,m s One possible embodiment of ) is similar to formula (1).

number

[0186] In both equations (1) and (2), the value of K may be predetermined and stored in the encoding and decoding devices, so that transmission of K is not required, and therefore the size of the bitstream can be smaller. The default value of K may be 1. In another possible embodiment, the value of K is β v It may depend on the following: Different values ​​of K can be used in equation (1) and equation (2), for example, K=1.5 in equation (1) and K=0.5 in equation (2).

[0187] The above embodiment proposes deriving a target gain vector outside a predetermined range, and thus any desired image compression quality can be achieved. In situations where lower distortion is desired, higher compression quality can be achieved by extrapolating a predetermined maximum gain vector. In situations where a lower bitstream size is desired, lower compression quality can be achieved by extrapolating a predetermined minimum gain vector.

[0188] In a possible embodiment, the two closest pre-trained values ​​of β and their corresponding gain vectors are β v The target gain vector m v Used to obtain β. v The maximum value β is set in advance. t When it is greater than, the maximum value β t and its corresponding gain vector m t , and β v The next closest pre-trained value β (in Figure 8, β) r (as shown) and its corresponding gain vector m r However, β v no m v It is used to derive m. v =f(β v ,β r ,β t ,m r ,m t ) is the function m v =f(β v ,β r ,β t ,m r,m t One possible embodiment of ) is,

number

number

[0189] Similarly, β v The minimum value β is set in advance. s When it is smaller than, the minimum value β s and its corresponding gain vector m s , and β v The next closest pre-trained value β (in Figure 8, β) r (as shown) and its corresponding gain vector m r However, β v no m v It is used to derive m. v =f(β v ,β r ,β s ,m r ,m s ) is the function m v =f(β v ,β r ,β s ,m r ,m s One possible embodiment of ) is similar to equation (3).

number

number

[0190] In the example in Figure 8, only three pre-trained pairs of [β, gain vector] are shown, and therefore β v The next closest pre-trained values ​​β are all β in equations (3) and (4). r Please note that this is shown as follows. Figure 8 is a schematic diagram, and in other possible embodiments, more than three pairs of [β, gain vector] are pre-trained, and β v ga β s When it is smaller than β, and β v ga β t When it is greater than β v The next closest pre-trained value β should be different.

[0191] In both equations (3) and (4), when K is set to its default value, transmission of K is not required, and therefore the size of the bitstream can be smaller. The default value of K may be 1. In another possible embodiment, the value of K is β v It may depend on the following: Different values ​​of K can be used in equations (3) and (4), for example, a large β v Then K=1, and a smaller β v So

number

[0192] Using more pre-trained [β, gain vector] pairs results in better performance.

[0193] In possible embodiments, there are more than one pairs of [β, gain vector], β v The new gain vector m v Used to derive . In possible embodiments of this embodiment, β v The N (N>2) pre-trained values ​​β closest to β, and their corresponding gain vectors, can be obtained by polynomial extrapolation. v The target gain vector m v It is used to derive β. v The maximum value β is set in advance. t When it is greater than this, the target gain vector satisfies the following condition.

number

[0194] In possible designs, the specific form of equation (5) is as follows:

number

[0195] Similarly, β v The minimum value β is set in advance. s When it is smaller than this, the target gain vector satisfies the following condition.

number

[0196] In possible designs, the specific form of equation (6) is as follows:

number

[0197] In possible embodiments, all pre-trained pairs of [β, gain vector] are β v The new gain vector m v This is used to derive the following. Then, each pre-trained value β is given weights to control the contributions of different pre-trained values ​​of β and their corresponding gain vectors.

[0198] It should be noted that in the possible forms of the function described above, additional biases or offsets may be added. The forms of the function are not limited in this disclosure. All operations related to the gain vector are element-wise operations, meaning that the operations are performed on each individual element of the gain vector.

[0199] In possible designs, when the image contains only chroma samples or only luma samples, the target gain vector is obtained in step 703; when the image contains both chroma and luma samples, the target gain vector for the luma samples is obtained in step 703, and the target gain vector for the chroma samples is also obtained in step 703, possibly by performing step 703 twice. In step 704, the image is encoded based on the target gain vector. After obtaining the target gain vector in step 703, the target gain vector is used to encode the image. Figure 9 is a flowchart of an exemplary method for encoding an image based on a target gain vector, and Figure 10 is an exemplary VAE framework that can perform the method 900 of Figure 9. As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible embodiments.

[0200] The encoding method 900 shown in Figure 9 includes the steps of: 901 obtaining a first feature map of an image using a neural network; 902 obtaining a second feature map based on the first feature map and a target gain vector; 903 quantizing the second feature map to obtain a quantized second feature map; and 904 encoding the quantized second feature map to obtain a bitstream.

[0201] In step 901, a neural network is used to obtain a latent representation of the image (i.e., a first feature map). In the VAE framework, the function of the neural network may be performed by an encoder 1001 shown in Figure 10. The encoder 1001 maps the input image x (i.e., the image) to a latent representation (denoted by y) via the function y = f(x). The function f() is a transformation function that converts the input signal x to a more compressible representation y. Typically, y is a tensor of shape w × h × d, for example w × h × 128 or w × h × 64, where w represents the feature map width, h represents the feature map height, and d represents the number of feature map channels. Note that the feature map width and feature map height may or may not be equal, and this is not limited here. The number of feature map channels may be other integer values, for example 32, 48, 96, 144, 160, 176, 192, 256, etc., and this is not limited here. In possible designs, the number of feature map channels for luma samples and chroma samples may differ, for example, 128 for luma samples and 64 for chroma samples. Unit 1001 is called an "encoder," but the complete encoding network described in Figure 10 can also be called an "encoder." The encoding process generally refers to a unit (module) that transforms an input into an encoded (e.g., compressed) output.

[0202] In step 902, after the first feature map of the image has been obtained in step 901, further calculations can be performed on the first feature map based on the target gain vector. The target gain vector may be the target gain vector obtained in step 703 of method 700. In a possible embodiment, the function of step 902 can be performed by a gain unit 1011 after the encoder 1001. The input to the gain unit 1011 (denoted by y) is the output of the encoder 1001. The gain unit 1011 further transforms the first feature map with the target gain vector. As stated in step 702, the gain vector is a vector with dimension 1 × d, where d is an integer greater than 1. In this disclosure, d is equal to the number of feature map channels in y. Thus, each element of the target gain vector can be mapped to a feature map channel in y. In a possible embodiment, the gain unit 1011 multiplies the target gain vector and the first feature map to obtain a second feature map. The output of the gain unit 1011 is

number

[0203] It should be noted that step 703 must be performed before step 902, and that the actions described in the other steps of methods 700 and 900 can be performed in a different order and still achieve the desired results. As an example, the process shown in the attached diagram does not necessarily require the specific order shown or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous. For example, step 702 may be performed before step 901, or step 702 may be performed after step 901, or step 702 and step 901 may be performed in parallel.

[0204] In step 903, the second feature map obtained in step 902 is quantized. The quantization process can be performed by the quantizer 1002 shown in Figure 10. The quantizer 1002 is

number

number

number

[0205] In step 904, the quantized second feature map obtained in step 903 is further encoded into bitstream 1 using entropy coding. The entropy coding process can be performed by the arithmetic encoder (AE) 1005 shown in Figure 10. Quantized latent representation

number

[0206] The output of step 904 is a bitstream, which is then provided (transmitted) to the decoding process. As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible embodiments.

[0207] In a possible design, when the image contains only chroma samples or only luma samples, a target gain vector is obtained in step 703, and method 900 is performed to encode the image based on the target gain vector; when the image contains both chroma samples and luma samples, the target gain vector for the luma samples is obtained in step 703, possibly by performing step 703 twice, and the target gain vector for the chroma samples is also obtained in step 703, and the luma samples and chroma samples are encoded independently using method 900 based on their corresponding target gain vectors. The encoded luma samples and chroma samples can be packed into a bitstream for storage or transmission.

[0208] In possible embodiments, the encoding method 700 further includes encoding a first encoding parameter into a bitstream. Depending on the circumstances, different images may need to be compressed with different qualities, and therefore the encoding method 700 can be used to encode different images by setting the first encoding parameter to different values. The proposed method 700 makes it possible to achieve different compression qualities with a single pre-trained model, in particular to achieve compression qualities outside the pre-trained range.

[0209] A first encoding parameter is signaled in the bitstream (for example, as a first flag such as compression_quality_level) and transmitted to a decoding device, which then decodes the image data encoded based on the first encoding parameter. In possible embodiments, the first encoding parameter is encoded or signaled in the picture parameter set (PPS) of the bitstream. In possible embodiments, the first encoding parameter is encoded or signaled in the sequence parameter set (SPS) of the bitstream. In possible embodiments, the first encoding parameter is encoded or signaled in the picture header (PH) or slice header (SH) of the bitstream.

[0210] In a possible design, the number of bits used to signal the first encoded parameter (first flag) in the bitstream is 16 or less. For example, the first parameter could take 15 bits in the bitstream.

[0211] In some possible designs, the first coding parameter is signaled in other forms to conserve bits used to transmit the first coding parameter. For example, in a possible design, the first coding parameter is signaled in its base-2 logarithmic form. In another possible design, the first coding parameter is signaled in its base-2 logarithmic form with a bias. The base of the logarithm can be any other value, such as 10, and is not limited here. In this way, the bits used to transmit the first coding parameter can be reduced to four bits or less.

[0212] Optionally, the side information z output by the hyperencoder 1003 may pass through the gain unit 1012 before quantization. The function of the gain unit 1012 is the same as that of the gain unit 1011, but the gain vectors used in the gain unit 1012 and the gain unit 1011 may be different. This embodiment allows for a more flexible way of encoding the side information.

[0213] In possible embodiments, the image contains only luma samples, and method 700 is used to encode the luma samples based on a first encoding parameter to obtain a bitstream, the first encoding parameter is signaled as a first flag (such as compression_quality_level) in the bitstream as described above.

[0214] In possible embodiments, the image contains only chroma samples, and method 700 is used to encode the chroma samples based on a first encoding parameter to obtain a bitstream, the first encoding parameter being signaled as a first flag (such as compression_quality_level) in the bitstream as described above.

[0215] In possible embodiments, the image includes both a luma sample and a chroma sample, where the luma sample is encoded using Method 700 based on a first encoding parameter, the chroma sample is encoded using Method 700 based on a second encoding parameter, or vice versa. Some steps of Method 700 may be performed only once when encoding an image using both a luma sample and a chroma sample, as in step 701. Both the first and second encoding parameters are signaled in the bitstream as described above, for example, the first encoding parameter is signaled as a first flag (e.g., compression_quality_level_luma) and the second encoding parameter is signaled as a second flag (e.g., compression_quality_level_chroma). Alternatively, the first coded parameter is signaled as a first flag (e.g., compression_quality_level_chroma), and the second coded parameter is signaled as a second flag (e.g., compression_quality_level_luma). The second coded parameter is signaled in the same way as the first coded parameter. Thus, the number of bits used to transmit the coded parameter is doubled. At least one of the values ​​of the first coded parameter and the second coded parameter is either less than a preset minimum or greater than a preset maximum.

[0216] In possible embodiments, the image includes both luma samples and chroma samples, and the first encoding parameter includes a first component and a second component. The luma samples are encoded based on the first component using Method 700, and the chroma samples are encoded based on the second component using Method 700. It should be noted that some steps of Method 700 may be performed only once when encoding the image using both luma and chroma samples, as in step 701. Both the first and second components are signaled in the bitstream as described above, for example, the first component is signaled as a first flag (e.g., compression_quality_level_luma), and the second component is signaled as a second flag (e.g., compression_quality_level_chroma). Both the first and second components are signaled in the same way that the first encoding parameter is signaled as described above. Therefore, the number of bits used to transmit the first encoding parameter is doubled.

[0217] In possible embodiments, different regions of the entire image are encoded with different encoding parameters β v It may also be encoded based on several encoding parameters β v This is signaled in the bitstream so that it can be transmitted to other devices. For example, to obtain better compression quality, the region of interest is larger beta. v It may also be encoded with a smaller background region to obtain a smaller bitstream size. v It may also be encoded using [a specific encoding method]. The proposed method allows for different compression qualities to be achieved for different regions of a single image, and the encoding process is more flexible and can meet different compression requirements for different situations.

[0218] The embodiment shown in Figure 7 may be configured to provide an output that is readily decoded by the decoding method described with reference to Figure 11. Method 700 in Figure 7 is described as being performed by one or more computer neural network systems located at one or more locations. For example, a system configured to perform image compression, e.g., the neural network in Figure 1, can perform Method 700. In general, the embodiments mentioned above can be combined to provide greater flexibility.

[0219] Decryption method Figure 11 is a flowchart illustrating an exemplary method for decoding an image based on a neural network architecture, which includes the steps of: 1101 acquiring a bitstream containing encoded image data; 1102 analyzing the bitstream to acquire a first encoding parameter such that the value of the first encoding parameter is less than a preset minimum or greater than a preset maximum; 1103 acquiring a target inverse gain vector based on the first encoding parameter; and 1104 acquiring an image based on the target inverse gain vector.

[0220] In step 1101, a bitstream containing the encoded image data is obtained, presumably from an encoding device or a distribution device. The bitstream may contain some side information (e.g., mean or variance of encoded samples…) and some encoding parameter information (e.g., encoding mode, compression quality parameters, quantization parameters…).

[0221] In step 1102, the bitstream is parsed to obtain a first coding parameter, the value of which is either less than a preset minimum or greater than a preset maximum. The first coding parameter (e.g., denoted as β) is a scalar input for the decoding process, used to select the compression quality.

[0222] A first flag (e.g., compression_quality_level) is signaled in the bitstream to specify the image compression quality (i.e., a first coding parameter β). In a possible design, the first flag is the compression quality and may require 16 or 15 bits for storage and transmission. In this design, the first coding parameter β is parsed directly from the bitstream. In another possible design, the first flag is the base-2 logarithmic form of the first coding parameter (e.g., log2_compression_quality_level or log2_compression_quality_level_minus1), and the first coding parameter can be derived as follows: first coding parameter β = 1 << first flag, or first coding parameter β = 1 << (first flag + bias), where "<<" means left shift, and the first flag may be an integer in the range of 0 to 16, and the bias may be 1, 2, 3, 4, 5, 8, etc.

[0223] In possible embodiments, a first flag and a second flag are signaled in the bitstream to specify the compression quality of the image. The first flag (e.g., compression_quality_level_luma) is a first encoding parameter (β) that specifies the compression quality of the luma samples of the image. Y When related to (as shown), the second flag (e.g., compression_quality_level_chroma) is a second encoding parameter (β) that specifies the compression quality of the chroma samples in the image. UV Related to (shown as). The first flag (e.g., compression_quality_level_chroma) is related to the first encoding parameter (β) which specifies the compression quality of the chroma samples in the image. UV When related to (as shown), the second flag (e.g., compression_quality_level_luma) is a second encoding parameter (β) that specifies the compression quality of the image's luma samples. YThis relates to (as shown). Both the first and second coding parameters can be derived using the method described in the previous paragraph. At least one of the values ​​of the first and second coding parameters is less than a preset minimum value or greater than a preset maximum value.

[0224] In possible embodiments, a first flag is signaled in the bitstream to specify the compression quality of the image. The first flag includes a first component (e.g., compression_quality_level_luma) and a second component (e.g., compression_quality_level_chroma). The first component is a first encoding parameter (β) that specifies the compression quality of the luma samples of the image. Y The second component is related to (shown as), and the second component is a second encoding parameter (β) that specifies the compression quality of the chroma samples of the image. UV This relates to (as shown above). Both the first and second coding components can be derived using the method described in the paragraph above.

[0225] In step 1103, after obtaining the first coding parameter, the target inverse gain vector is obtained based on the first coding parameter. During the network training phase, a pre-trained set of possible values ​​of β and the corresponding gain vectors is obtained and stored in both the coding device and the decoding device, and the possible values ​​of β are the preset minimum value β s and the pre-set maximum value β t It lies between [the two points]. Therefore, the target gain vector m v This can be obtained using the method described with reference to Figure 8. Next, the target inverse gain vector m v ' can be obtained based on the target gain vector and the target inverse gain vector m v The following conditions must be met.

number

[0226] In a possible design, a pre-trained set of possible values ​​for β and their corresponding inverse gain vectors is acquired during the training phase and then stored in the decoding device. Thus, the target inverse gain vector m v ' can be obtained using the method described with reference to Figure 8.

[0227] In step 1104, an image is acquired based on the target inverse gain vector. After acquiring the target inverse gain vector in step 1103, the target inverse gain vector is used to decode the image. Figure 12 is a flowchart of an exemplary method for decoding an image based on the target inverse gain vector, and Figure 10 is an exemplary VAE framework that can perform the method 1200 of Figure 12. In general, any of the embodiments mentioned above and their possible designs or possible embodiments can be combined to provide greater flexibility.

[0228] The decoding method 1200 shown in Figure 12 includes the steps of: 1201 analyzing a bitstream to obtain a first latent representation of an image using entropy decoding; 1202 obtaining a second latent representation based on the first latent representation and a target inverse gain vector; and 1203 decoding the second latent representation to obtain an image using a neural network.

[0229] In step 1201, the entropy decoding process returns the binary number to the sampled value (i.e.,

number

number

[0230] In step 1202, after the first latent representation of the image is obtained in step 1201, further calculations can be performed on the first latent representation based on the target inverse gain vector. The target inverse gain vector may be the target inverse gain vector obtained in step 1103 of method 1100. In possible embodiments, the function of step 1202 can be performed by the inverse gain unit 1013 after AD 1006. Input of gain unit 1013 (

number

number

number

[0231] It should be noted that step 1103 must be performed before step 1202, and that the actions described in the other steps of methods 1100 and 1200 can be performed in a different order and still achieve the desired results. For example, the process shown in the attached diagram does not necessarily require the specific order shown or a consecutive order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous. For example, step 1102 may be performed before step 1201, or step 1102 may be performed after step 1201, or step 1102 and step 1201 may be performed in parallel.

[0232] In a possible embodiment, a target gain vector is obtained in step 1103' (not shown in Figure 11), and then in step 1104' (not shown in Figure 11), the bitstream is decoded to obtain an image based on the target gain vector. In this embodiment, the inverse gain unit 1013 multiplies the first latent representation by the reciprocal of the target inverse gain vector to obtain a second latent representation.

[0233] In step 1203, the second latent representation obtained in step 1202 is further decoded into an image using a neural network. In the VAE framework, the function of the neural network may be performed by decoder 1004 shown in Figure 10. Unit 1004 is called a “decoder,” but it is also possible to refer to the complete decoding network described in Figure 10 as a “decoder.” The decoding process generally refers to a unit (module) that converts a latent representation into an image output.

[0234] The output of step 1203 is an image

number

[0235] Optionally, the side information z output by the hyperencoder 1003 is processed by the gain unit 1012 before quantization. The inverse gain unit 1014 is used to process the output of AD1010 in Figure 10.

[0236] As will be apparent to those skilled in the art, this embodiment can be combined with any of the embodiments mentioned above and any of their possible designs or possible embodiments. Further information may be provided.

[0237] While the drawings show operations in a specific order, this should not be understood as meaning that such operations must be performed in a specific or sequential order shown, or that all illustrated operations must be performed, in order to achieve the desired result. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as meaning that such separation is necessary in all embodiments, and the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0238] Specific embodiments of the subject matter are described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims can be performed in a different order and still achieve the desired results. As an example, the process shown in the accompanying figures does not necessarily require the specific illustrated order or a sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.

[0239] Based on the framework shown in Figure 10, a simple example is provided to illustrate the technical effect of the method provided. Assuming that the output of encoder 1001 is y = 8.53, Conventional method: The gain vector is not used to process y, then

number

number

number

number

number

number

[0240] As shown in the simple example above, the method proposed in this disclosure yields far less distortion, especially when the gain vector has a large value.

[0241] Encoding devices and decoding devices Furthermore, as already mentioned, the Disclosure also provides a device configured to perform the steps of the method described above. Figure 13 shows an encoding device 1300 for encoding for processing by a neural network-based unit. The device 1300 comprises a feature map acquisition module 1310 configured to acquire a first feature map from an input image; a gain unit 1330 configured to transform the first feature map based on a target gain vector to acquire a second feature map; a quantization module 1340 configured to quantize the second feature map to acquire a quantized second feature map; and an entropy coding module 1350 configured to encode the quantized second feature map to acquire a bitstream (e.g., using entropy coding). Optionally, the device 1300 may further comprise a gain vector acquisition module 1320 configured to acquire a target gain vector based on a first coding parameter.

[0242] Corresponding to the encoding device 1300 mentioned above, Figure 14 shows a decoding device 1400 for decoding a bitstream to reconstruct an image by a neural network-based unit. This device may comprise an entropy decoding module 1410 configured to analyze a bitstream to obtain a first latent representation of an image using entropy decoding, an inverse gain unit 1430 configured to obtain a second latent representation based on the first latent representation and a target inverse gain vector, and an image reconstruction module configured to decode the second latent representation to obtain an image using a neural network. Optionally, device 1400 may further comprise an inverse gain vector acquisition module 1420 configured to obtain a target inverse gain vector based on first encoding parameters.

[0243] It should be noted that these devices may be further configured to perform any additional features, including the exemplary embodiments mentioned above. For example, a device is provided for decoding a feature map for processing by a neural network based on a bitstream, the device comprising processing circuitry configured to perform any step of the decoding methods described above. Similarly, a device is provided for encoding a feature map for processing by a neural network into a bitstream, the device comprising processing circuitry configured to perform any step of the encoding methods described above.

[0244] Further devices utilizing devices 1300 and / or 1400 may be provided. For example, a device for image or video encoding may include encoding device 1300. Furthermore, it may include decoding device 1400. A device for image or video decoding may include decoding device 1400 and / or encoding device 1300.

[0245] Furthermore, an encoding system utilizing devices 1300 and / or 1400 may be provided. For example, the encoding system may be deployed on a server. The server receives a bitstream, then decodes the bitstream using device 1400 or another decoder to obtain an image, and then the image is encoded using device 1300 or another encoder. After encoding, the newly obtained bitstream is stored and / or transmitted to another device.

[0246] Some exemplary embodiments in hardware and software A corresponding system capable of deploying the encoder-decoder processing chain mentioned above is shown in Figure 15. Figure 15 is a schematic block diagram showing exemplary encoding systems, e.g., video, image, audio, and / or other encoding systems (or short encoding systems) that may utilize the technology of this application. The video encoder 20 (or short encoder 20) and video decoder 30 (or short decoder 30) of the video encoding system 10 represent examples of devices that may be configured to perform the technology described in the various examples of this disclosure. For example, video encoding and decoding may use a distributed neural network to which the bitstream analysis and / or bitstream generation mentioned above can be applied to transmit feature maps between distributed computing nodes (two or more).

[0247] As shown in Figure 15, the encoding system 10 includes a source device 12 configured to provide encoded picture data 21 to, for example, a destination device 14 for decoding the encoded picture data 13.

[0248] The source device 12 includes an encoder 20 and may additionally, or optionally, include a picture source 16, a preprocessor (or preprocessing unit) 18, for example, a picture preprocessor 18, and a communication interface or communication unit 22.

[0249] The picture source 16 may comprise, or may comprise, any kind of picture capture device for capturing real-world pictures, e.g., a camera, and / or any kind of picture generation device, e.g., a computer graphics processor for generating computer-animated pictures, or any other kind of device for acquiring and / or providing real-world pictures, computer-generated pictures (e.g., screen content, virtual reality (VR) pictures), and / or any combination thereof (e.g., augmented reality (AR) pictures). The picture source may also comprise any kind of memory or storage for storing any of the aforementioned pictures.

[0250] To distinguish it from the processing performed by the preprocessor 18 and the preprocessing unit 18, the picture or picture data 17 may also be called the raw picture or raw picture data 17.

[0251] The preprocessor 18 is configured to receive (raw) picture data 17 and perform preprocessing on the picture data 17 to obtain a preprocessed picture 19 or preprocessed picture data 19. Preprocessing performed by the preprocessor 18 may include, for example, cropping, color format conversion (e.g., from RGB to YCbCr), color correction, or denoising. It can be understood that the preprocessing unit 18 may consist of optional components. It should also be noted that preprocessing may also use a neural network that uses presence indicator signaling (such as one of the examples in Figures 1 to 7).

[0252] The video encoder 20 is configured to receive pre-processed picture data 19 and provide encoded picture data 21.

[0253] The communication interface 22 of the source device 12 may be configured to receive the encoded picture data 21 and transmit the encoded picture data 21 (or any further processed version thereof) via the communication channel 13 to another device, such as the destination device 14 or any other device, for storage or direct reconstruction.

[0254] The destination device 14 includes a decoder 30 (e.g., a video decoder 30) and may additionally, or optionally, include a communication interface or communication unit 28, a post-processor 32 (or post-processing unit 32), and a display device 34.

[0255] The communication interface 28 of the destination device 14 is configured to receive encoded picture data 21 (or any further processed version thereof) from, for example, directly from the source device 12, or from any other source, such as a storage device, such as an encoded picture data storage device, and to provide the encoded picture data 21 to the decoder 30.

[0256] Communication interfaces 22 and 28 may be configured to transmit or receive encoded picture data 21 or encoded data 13 via a direct communication link between the source device 12 and the destination device 14, for example, via a direct wired or wireless connection, or via any type of network, for example, a wired or wireless network or any combination thereof, or any type of private network and public network or any type of combination thereof.

[0257] The communication interface 22 may be configured, for example, to package the encoded picture data 21 into an appropriate format, such as a packet, and / or to process the encoded picture data using any kind of transmission encoding or processing for transmission over a communication link or communication network.

[0258] The communication interface 28 that forms the counterpart to the communication interface 22 may be configured, for example, to receive the transmitted data and process the transmitted data using any kind of corresponding transmission decoding or processing and / or depackaging to obtain the encoded picture data 21.

[0259] Both communication interfaces 22 and 28 may be configured as unidirectional or bidirectional communication interfaces, as indicated by the arrows of the communication channel 13 in Figure 15, pointing from the source device 12 to the destination device 14, and may be configured to confirm and exchange any other information relating to the communication link and / or data transmission, such as encoded picture data transmission, for example, to send and receive messages, for example, to set up a connection. The decoder 30 is configured to receive encoded picture data 21 and provide decoded picture data 31 or decoded picture 31.

[0260] The post-processor 32 of the destination device 14 is configured to post-process the decoded picture data 31 (also called reconstructed picture data), for example, the decoded picture 31, in order to obtain the post-processed picture data 33, for example, the post-processed picture 33. The post-processing performed by the post-processing unit 32 may include, for example, color format conversion (e.g., from YCbCr to RGB), color correction, cropping, or resampling, or any other processing to prepare the decoded picture data 31 for display by, for example, the display device 34.

[0261] The display device 34 of the destination device 14 is configured to receive post-processed picture data 33 for displaying the picture to, for example, a user or viewer. The display device 34 may be any type of display for representing the reconstructed picture, such as an integrated or external display or monitor, or may include such a display. The display may include, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a plasma display, a projector, a microLED display, a liquid crystal on silicon (LCoS), a digital light processor (DLP), or any other type of display.

[0262] Although Figure 15 shows the source device 12 and destination device 14 as separate devices, the device embodiments may also include both or both functions, such as the source device 12 or its corresponding function and the destination device 14 or its corresponding function. In such embodiments, the source device 12 or its corresponding function and the destination device 14 or its corresponding function may be implemented using the same hardware and / or software, or by separate hardware and / or software, or any combination thereof.

[0263] As will be apparent to those skilled in the art based on the description, the presence and (exact) division of different units or functions within the source device 12 and / or destination device 14, as shown in Figure 15, may be modified depending on the actual device and application.

[0264] The encoder 20 (e.g., video encoder 20) or the decoder 30 (e.g., video decoder 30), or both the encoder 20 and the decoder 30, may be implemented by processing circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, hardware, dedicated to video coding, or any combination thereof. The encoder 20 may be implemented by processing circuitry 46 to implement various modules, including a neural network or a portion thereof. The decoder 30 may be implemented by processing circuitry 46 to implement any coding system or subsystem described herein. The processing circuitry may be configured to perform various operations, such as those described below. Where the technology is partially implemented in software, the device may store instructions for the software in a suitable non-temporary computer-readable storage medium, or it may use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Either the video encoder 20 or the video decoder 30 may be integrated as part of a combined encoder / decoder (CODEC) in a single device, for example, as shown in Figure 16.

[0265] The source device 12 and destination device 14 may include any wide range of devices, including any type of handheld or fixed device, such as a notebook or laptop computer, a mobile phone, a smartphone, a tablet or tablet computer, a camera, a desktop computer, a set-top box, a television, a display device, a digital media player, a video gaming console, a video streaming device (content service server or content distribution server), a broadcast receiver device, or a broadcast transmitter device, and may or may not use an operating system. In some cases, the source device 12 and destination device 14 may be equipped with wireless communication. Therefore, the source device 12 and destination device 14 may be wireless communication devices.

[0266] In some cases, the video encoding system 10 shown in Figure 15 is merely an example, and the techniques of this disclosure may apply to video encoding configurations (e.g., video encoding or video decoding) that do not necessarily involve data communication between an encoding device and a decoding device. In other examples, data may be retrieved from local memory and streamed over a network. The video encoding device may encode the data and store it in memory, and / or the video decoding device may retrieve the data from memory and decode it. In some examples, encoding and decoding are performed by devices that do not communicate with each other, simply encoding the data into memory and / or retrieving and decoding the data from memory.

[0267] Figure 17 is a schematic diagram of a video encoding device 8000 according to one embodiment of the present disclosure. The video encoding device 8000 is suitable for carrying out the disclosed embodiments described herein. In one embodiment, the video encoding device 8000 may be a decoder, such as the video decoder 30 in Figure 15, or an encoder, such as the video encoder 20 in Figure 15.

[0268] The video coding device 8000 comprises an ingress port 8010 (or input port 8010) and a receiver unit (Rx) 8020 for receiving data, a processor, logic unit, or central processing unit (CPU) 8030 for processing data, a transmitter unit (Tx) 8040 and an egress port 8050 (or output port 8050) for transmitting data, and memory 8060 for storing data. The video coding device 8000 may also comprise optical-electrical (OE) and electrical-optical (EO) components coupled to the ingress port 8010, receiver unit 8020, transmitter unit 8040, and egress port 8050 for outputting or inputting optical or electrical signals.

[0269] The processor 8030 is implemented by hardware and software. The processor 8030 may be implemented as one or more CPU chips, cores (e.g., as a multi-core processor), FPGAs, ASICs, and DSPs. The processor 8030 communicates with the ingress port 8010, the receiver unit 8020, the transmitter unit 8040, the egress port 8050, and the memory 8060. The processor 8030 includes a neural network-based codec 8070. The neural network-based codec 8070 implements the disclosed embodiments described above. For example, the neural network-based codec 8070 performs, processes, prepares, or provides various encoding operations. Thus, the inclusion of the neural network-based codec 8070 results in a substantial improvement to the functionality of the video encoding device 8000 and results in the transformation of the video encoding device 8000 to different states. Alternatively, the neural network-based codec 8070 is stored in memory 8060 and implemented as instructions executed by processor 8030.

[0270] The memory 8060 may comprise one or more disks, tape drives, and solid-state drives, and may be used as an overflow data storage device to store the program when it is selected for execution, and to store instructions and data read during program execution. The memory 8060 may be, for example, volatile and / or non-volatile, and may be read-only memory (ROM), random access memory (RAM), tri-level associative memory (TCAM), and / or static random access memory (SRAM).

[0271] Figure 18 is a simplified block diagram of a device that can be used as either or both of the source device 12 and destination device 14 from Figure 15, according to an exemplary embodiment.

[0272] The processor 9002 within the apparatus 9000 may be a central processing unit. Alternatively, the processor 9002 may be any other type of device, or multiple devices, capable of manipulating or processing information that currently exists or will be developed in the future. The disclosed embodiments may be implemented with a single processor, e.g., processor 9002, as shown in the figures, but advantages in speed and efficiency may be achieved using one or more processors.

[0273] In one embodiment, the memory 9004 within the device 9000 may be a read-only memory (ROM) device or a random access memory (RAM) device. Any other suitable type of storage device may be used as memory 9004. Memory 9004 may contain code and data 9006 accessed by the processor 9002 using the bus 9012. Memory 9004 may further contain an operating system 9008 and an application program 9010, the application program 9010 including at least one program that enables the processor 9002 to perform the method described herein. For example, the application program 9010 may include applications 1 to N, the applications 1 to N further including a video encoding application that performs the method described herein.

[0274] The device 9000 may also include one or more output devices, such as a display 9018. The display 9018 may, in one example, be a touch-sensitive display combining a display with a touch-sensitive element capable of sensing touch input. The display 9018 can be coupled to the processor 9002 via the bus 9012.

[0275] Although shown here as a single bus, the bus 9012 of device 9000 can consist of multiple buses. Furthermore, secondary storage can be directly coupled to other components of device 9000 or accessed via a network, and can comprise a single integrated unit such as a memory card or multiple units such as multiple memory cards. Thus, device 9000 can be implemented in a wide variety of configurations.

[0276] In one possible embodiment, an exemplary method for storing a bitstream, the method is: Obtaining a bitstream according to one of the previously shown encoding methods, Storing a bitstream in a storage medium and A method is provided that includes this.

[0277] This method is optional, To obtain an encrypted bitstream, one must perform an encryption process on the bitstream, Storing an encrypted bitstream on a storage medium and It also includes.

[0278] Please understand that any known encryption method may be used.

[0279] This method is optional, This further includes performing a segmentation process on a bitstream to obtain multiple bitstream segments, and storing the multiple bitstream segments in a storage medium.

[0280] This method is optional, This further includes taking at least one backup of the bitstream and storing at least one backup on a storage medium. It should be understood that at least one backup of the bitstream may be stored on a different storage medium than the one storing the original bitstream.

[0281] This method is optional, Receiving multiple bitstreams generated according to one of the previously shown encoding methods, Assigning address information or identification information separately to multiple bitstreams, To store bitstreams in their corresponding locations according to address information or identification information corresponding to multiple bitstreams. It also includes.

[0282] This method is optional, Classify the bitstream to obtain at least two bitstreams, and at least two bitstreams must include a first bitstream and a second bitstream. The first bitstream is stored in the first memory space, and the second bitstream is stored in the second memory space. It also includes.

[0283] This method is optional, The video streaming device transmits a bitstream to a terminal device, and the video streaming device can be a content server or content delivery server.

[0284] In one possible embodiment, an exemplary system for storing a bitstream, the system is: A receiver configured to receive a bitstream generated by one of the previous encoding methods, A processor configured to perform an encryption process on a bitstream in order to obtain an encrypted bitstream, A computer-readable storage medium configured to store an encrypted bitstream and A system including this is provided.

[0285] Optionally, the system may include several storage media, some of which may be deployed in different locations. Furthermore, multiple bitstreams may be stored on different storage media in a distributed manner. For example, some storage media may include a first storage medium configured to store a first bitstream and a second storage medium configured to store a second bitstream.

[0286] Optionally, this system includes a video streaming device, which can be a content server or a content distribution server, and is configured to retrieve a bitstream from one of the storage media and transmit the bitstream to a terminal device.

[0287] In one possible embodiment, an exemplary method for converting the format of a bitstream, the method is: Receiving a bitstream of a first format generated by one of the previously described encoding methods, Converting a bitstream of the first format to a bitstream of the second format, The second format bitstream is stored on a storage medium. A method is provided that includes this.

[0288] This method is optional, Steps to send a stored bitstream in a second format to the terminal device in response to an access request from the terminal device. It also includes.

[0289] In one possible embodiment, an exemplary system for converting bitstream formats, the system is: A receiver configured to receive a bitstream of a first format generated by one of the previously described encoding methods, A processor configured to convert a bitstream of a first format to a bitstream of a second format, The processor is further configured to store a bitstream of a second format in a storage medium. The storage medium is configured to store a bitstream of a second format, and includes a processor. A transmitter configured to send a stored bitstream in a second format to a terminal device in response to an access request from the terminal device, A system including this is provided.

[0290] In one possible embodiment, an exemplary method for processing a bitstream, the method is: A transport stream containing a video stream and an audio stream is received, and the video stream is generated by one of the encoding methods previously shown. Demultiplexing the transport stream to separate the video stream and the audio stream, Decoding a video stream using a video decoder to obtain video data, To obtain audio data, an audio decoder is used to decode the audio stream. A method is provided that includes this.

[0291] This method is optional, Synchronizing audio data and video data, Outputting the synchronization results to the playback device for playback It also includes.

[0292] This method is optional, Decoding a bitstream to obtain video or image data, Performing at least one of the following on video or image data: luminance mapping, chroma mapping, resolution adjustment, or format conversion. Transmitting video data or image data to a display It also includes.

[0293] In one possible embodiment, an exemplary method for transmitting a bitstream based on a user action request, the method is: The end-side device receives a first operation request, which is used to request playback of the target video. In response to the first operation request, the storage medium determines the bitstream corresponding to the target video, and the bitstream corresponding to the target video is a bitstream generated according to one of the encoding methods previously shown. To transmit the target bitstream to the end-side device and A method is provided that includes this.

[0294] This method is optional, Encapsulating the bitstream to obtain a transport stream in the first format, and Sending a transport stream of the first format to a terminal device for display, or Sending a transport stream of the first format to storage space for storage and It also includes.

[0295] In one possible embodiment, an exemplary system for transmitting a bitstream based on a user action request, the system is: A storage medium configured to store a bitstream, the bitstream being a bitstream generated according to one of the previously shown encoding methods, A receiver configured to receive a first operation request, A processor configured to determine a target bitstream in a storage medium in response to a first operation request, A transmitter configured to send the target bitstream to the terminal device and A system including this is provided.

[0296] Optionally, the processor is: Encapsulate the bitstream to obtain the transport stream of the first format. Further configured in this way, this system, A transport stream in the first format is sent to the terminal device for display, or, Send the transport stream of the first format to storage space for storage. Further includes a transmitter configured in such a manner.

[0297] In one possible embodiment, an exemplary method for downloading a bitstream, the method is: The bitstream is obtained from the storage medium, and the bitstream is generated according to one of the encoding methods previously shown. Decrypting the bitstream to obtain the streaming media file, Dividing a streaming media file into multiple streaming media segments, Downloading multiple streaming media segments separately and A method is provided that includes this.

[0298] In one possible embodiment, an exemplary system for downloading a bitstream, the system is: An acquisition unit is configured to acquire a bitstream from a storage medium, and the bitstream is generated according to one of the encoding methods previously shown. A decoder configured to decode a bitstream in order to obtain a streaming media file, A processor configured to split a streaming media file into multiple streaming media segments, The processor is configured to download multiple streaming media segments separately, and A system including the above is provided. However, the present invention is not limited to any of these exemplary embodiments.

[0299] In summary, this disclosure relates to a method and apparatus for encoding data (into a bitstream for still image or video processing). In particular, the data is processed by a network including a gain unit, in which a feature map is generated by an encoder layer. Prior to quantization, the feature map is processed by a gain unit that further transforms the feature map with a target gain vector. The target gain vector corresponds to a first encoding parameter used to control the compression quality of the image. The user can set different values ​​of the first encoding parameter to control different compression qualities for different images. The user can set different values ​​of the first encoding parameter to control different compression qualities for luminous and chroma samples. The first encoding parameter is also encoded into the bitstream. This technique provides flexible processing that can work with different bitstream sizes. Thus, data can be efficiently encoded in a bitstream depending on a first encoding parameter that can vary depending on the content of the encoded picture data.

[0300] This disclosure further relates to a method and apparatus for decoding data (into a bitstream for still image or video processing). In particular, the data is processed by a network including an inverse gain unit. In this process, a feature map is generated by an entropy decoder. Prior to further decoding, the feature map is processed by an inverse gain unit that further transforms the feature map with a target inverse gain vector. The target inverse gain vector is obtained based on a first encoding parameter that can be analyzed from the bitstream.

[0301] This disclosure provides an efficient method for encoding images or videos of different qualities and bitstream sizes without training more models. [Explanation of Symbols]

[0302] 10 Video encoding system, 12 Source device, 13 Communication channel, 14 Destination device, 16 Picture source, 17 Picture data, 18 Preprocessor, 19 Preprocessed picture data, 20 Encoder, 21 Encoded picture data, 22 Communication interface, 28 Communication interface, 30 Decoder, 31 Decoded picture data, 32 Postprocessor, 33 Postprocessed picture data, 34 Display device, 40 Video encoding system, 41 Imaging device, 42 Antenna, 43 Processor, 44 Memory store, 45 Display device, 46 Processing circuit, 101 Encoder, 102 Quantizer, 103 Hyperencoder, 104 Decoder, 105 Arithmetic encoder, 106 Arithmetic decoder, 107 Hyperdecoder, 108 Quantizer, 109 Arithmetic encoder, 110 Arithmetic decoder, 121 Encoder, 122 Quantizer, 123 Hyperencoder, 125 Arithmetic encoder, 127 Hyperdecoder, 144 Decoder, 146 Arithmetic decoder, 147 Hyperdecoder 401 Downsampling layer, 402 Downsampling layer, 403 Downsampling layer, 404 Downsampling layer, 405 Downsampling layer, 406 Downsampling layer, 407 Upsampling layer, 408 Upsampling layer, 409 Upsampling layer, 410 Upsampling layer, 411 Upsampling layer, 412 Upsampling layer, 413 Quantizer, 414 Input image, 415 Quantizer, 420 Convolutional layer, 430 Convolutional layer, 1001 Encoder, 1002 Quantizer, 1003 Hyperencoder, 1004 Decoder, 1005 Arithmetic encoder, 1006 Arithmetic decoder, 1007 Hyperdecoder, 1008 Quantizer, 1009 Arithmetic encoder, 1010 Arithmetic decoder, 1011 Gain unit, 1012 Gain unit, 1013 Inverse gain unit, 1014 Inverse Gain Unit, 1300 Encoding Device, 1310 Feature Map Acquisition Module, 1320 Gain Vector Acquisition Module, 1330 Gain Unit, 1340 Quantization Module, 1350 Entropy Encoding Module, 1400 Decoding Device, 1410 Entropy Decoding Module, 1420 Inverse Gain Vector Acquisition Module, 1430 Inverse Gain Unit, 1440 Image Reconstruction Module, 8000 Video Encoding Device, 8010 Ingress Port, 8020 Receiver Unit, 8030 Processor, 8040 Transmitter Unit, 8050 Egress Port, 8060 Memory, 8070 Neural Network-Based Codec, 9000 Device, 9002 Processor, 9004 Memory, 9006 Data, 9008 Operating System, 9010 Application Program, 9012 Bus, 9018 Display

Claims

1. An image encoding method using a neural network, Steps to acquire an image, A step of obtaining a first encoding parameter of the aforementioned image, wherein the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value. A step of obtaining a target gain vector based on the first coding parameter, The steps of encoding the image based on the target gain vector Includes, When the value of the first coding parameter is smaller than the preset minimum value, the step of obtaining a target gain vector based on the first coding parameter is: A step of obtaining the target gain vector based on the first coding parameter, the preset minimum value, and the first gain vector corresponding to the preset minimum value, A step of obtaining a first ratio of the first coding parameter to the preset minimum value, A step of obtaining the target gain vector based on the first ratio and the first gain vector, This includes, or further includes, the steps to obtain, When the value of the first coding parameter is greater than a preset maximum value, the step of obtaining a target gain vector based on the first coding parameter is: A step of obtaining the target gain vector based on the first encoding parameter, the preset maximum value, and the second gain vector corresponding to the preset maximum value, A step of obtaining the second ratio of the first encoding parameter to the preset maximum value, A step of obtaining the target gain vector based on the second ratio and the second gain vector, This includes the steps to obtain, method.

2. The step of obtaining the target gain vector based on the first ratio and the first gain vector is: The step of multiplying the first ratio by the first gain vector in order to obtain the target gain vector. The method according to claim 1, including the method described in claim 1.

3. The step of obtaining the target gain vector based on the second ratio and the second gain vector is: The step of multiplying the second ratio by the second gain vector in order to obtain the target gain vector. The method according to claim 1, including the method described in claim 1.

4. When the value of the first coding parameter is smaller than the preset minimum value, the step of obtaining a target gain vector based on the first coding parameter is: Steps to obtain the target gain vector based on the first coding parameter, the preset minimum value, the first gain vector corresponding to the preset minimum value, the third preset value closest to the preset minimum value, and the third gain vector corresponding to the third preset value. The method according to claim 1, including the method described in claim 1.

5. The step of obtaining a target gain vector based on the first coding parameter is: A step of obtaining the target gain vector based on the first coding parameter, N preset values, and N gain vectors corresponding to the N preset values, wherein N is an integer greater than 2, and the N preset values ​​include the preset minimum value and / or the preset maximum value. The method according to claim 1, including the method described in claim 1.

6. The step of encoding the image based on the target gain vector is: A step of obtaining a first feature map of the image using a neural network, A step of obtaining a second feature map based on the first feature map and the target gain vector, A step of quantizing the second feature map in order to obtain a quantized second feature map, A step of encoding the quantized second feature map in order to obtain a bitstream. The method according to claim 1, including the method described in claim 1.

7. The step of obtaining a second feature map based on the first feature map and the target gain vector is: The step of multiplying the target gain vector by the first feature map. The method according to claim 6, including the method described in claim 6.

8. The aforementioned method, Step of encoding the first encoding parameter into a bitstream. The method according to claim 1, further comprising:

9. A method for image decoding using a neural network, The steps include obtaining a bitstream containing encoded image data, A step of analyzing the bitstream to obtain a first encoding parameter, wherein the value of the first encoding parameter is less than a preset minimum value or greater than a preset maximum value. A step of obtaining a target inverse gain vector based on the first coding parameter, The steps include: acquiring an image based on the target inverse gain vector; Includes, When the value of the first coding parameter is smaller than the preset minimum value, the step of obtaining the target inverse gain vector based on the first coding parameter is: A step of obtaining the target inverse gain vector based on the first coding parameter, the preset minimum value, and the first inverse gain vector corresponding to the preset minimum value, A step of obtaining a first ratio of the first coding parameter to the preset minimum value, A step of obtaining the target inverse gain vector based on the first ratio and the first inverse gain vector, This includes, or further includes, the steps to obtain, When the value of the first coding parameter is greater than a preset maximum value, the step of obtaining a target inverse gain vector based on the first coding parameter is: A step of obtaining the target inverse gain vector based on the first coding parameter, the preset maximum value, and the second inverse gain vector corresponding to the preset maximum value, A step of obtaining the second ratio of the first encoding parameter to the preset maximum value, A step of obtaining the target inverse gain vector based on the second ratio and the second inverse gain vector, This includes the steps to obtain, method.

10. A computer program comprising program code for performing the method described in any one of claims 1 to 8 when executed on one or more processors.

11. An encoding device, Memory containing instructions, A processor coupled to the memory, wherein the processor is configured to execute the instructions to cause the encoding device to perform the encoding method according to any one of claims 1 to 8. An encoding device equipped with the following features.

12. A decoding device, Memory containing instructions, A processor coupled to the memory, wherein the processor is configured to execute the instructions to cause the decoding device to perform the decoding method described in claim 9, and A decoding device equipped with [a specific feature].

13. An encoding device, A receiver unit configured to receive a picture to be encoded or a bitstream to be decoded, A transmission unit coupled to the receiver unit, wherein the transmission unit is configured to transmit the bitstream to a decoder or to transmit the decoded image to a display, A memory coupled to at least one of the receiver unit or the transmission unit, wherein the memory is configured to store instructions, A processor coupled to the memory, wherein the processor is configured to execute the instructions stored in the memory in order to perform the method according to any one of claims 1 to 8 or claim 9; An encoding device equipped with the following features.