Method and apparatus for rate-adaptive neural image compression using an adversarial generator and a non-transitory computer-readable medium
By using an adversarial generator and attention-based generator in neural image compression, adapting anchor model instances to achieve intermediate compression rates solves the problem of poor bit rate control flexibility in the prior art, reducing resource consumption and improving model adaptation efficiency.
Patent Information
- Application Number
- CN202180006080.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-24
- Filing Date
- 2021-07-20
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-20
AI Technical Summary
Existing neural image compression methods are difficult to achieve flexible bit rate control, and usually require training multiple model instances to be adapted to different bit rates, resulting in high consumption of storage and computing resources.
Rate adaptive neural image compression is performed using an adversarial generator, and the anchor model instance is adapted to achieve intermediate compression rates by using a compact adversarial generator or attention-based adversarial generator on the encoder side or decoder side.
Reduces the resource requirement for implementing multi-compression deployment storage, provides a flexible NIC model framework, and can focus on significant information, improving the efficiency of model adaptation.
Smart Images

Figure CN114616576B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 054,648, filed on July 21, 2020, U.S. Provisional Patent Application No. 63 / 054,662, filed on July 21, 2020, U.S. Provisional Patent Application No. 63 / 054,665, filed on July 21, 2020, and U.S. Patent Application No. 17 / 356,722, filed on June 24, 2021, all of which are hereby incorporated by reference in their entirety. Technical field
[0003] This application relates to image processing. In particular, this application relates to a method and apparatus for rate - adaptive neural image compression using an adversarial generator, and a non - transitory computer - readable medium. Background art
[0004] ISO (International Organization for Standardization) / IEC (International Electrotechnical Commission) MPEG (Moving Picture Experts Group) (JTC 1 / SC 29 / WG 11) has been actively looking for potential requirements for standardizing future video coding technologies. ISO / IEC JPEG (Joint Photographic Experts Group) has established the JPEG - AI (Joint Photographic Experts Group - Artificial Intelligence) group, focusing on AI - based end - to - end neural image compression using deep neural networks (DNN). The success of the latest methods has brought increasing industrial interest in advanced neural image and video compression methods.
[0005] For previous Neural Image Compression (NIC) methods, flexible bitrate control remains a challenging issue. Generally, NIC methods may need to train multiple model instances separately for the trade-off between bitrate and distortion (quality of the compressed image) at each desired Rate-Distortion (R-D). All these multiple model instances may need to be stored and deployed on the decoder side to reconstruct images according to different bitrates. For many applications with limited storage resources and computing resources, this can be extremely costly. Summary of the Invention
[0006] According to an embodiment, a method for rate-adaptive neural image compression using an adversarial generator is executed by at least one processor, and the method includes: obtaining first features of an input image using a first part of a first neural network; generating first alternative features based on the obtained first features using a second neural network; and encoding the generated first alternative features using a second part of the first neural network to generate a first encoded representation. The method further includes: compressing the generated first encoded representation; decompressing the compressed representation; and decoding the decompressed representation using a third neural network to reconstruct a first output image.
[0007] According to an embodiment, an apparatus for rate-adaptive neural image compression using an adversarial generator includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to what is indicated by the program code to execute the above method for rate-adaptive neural image compression using an adversarial generator.
[0008] According to an embodiment, a non-transitory computer-readable medium storing instructions, the instructions being executed by at least one processor to cause at least one processor to execute the above method for rate-adaptive neural image compression using an adversarial generator.
[0009] An apparatus for rate-adaptive neural image compression using an adversarial generator includes: an acquisition module configured to obtain first features of an input image using a first part of a first neural network; a generation module configured to generate first alternative features based on the obtained first features using a second neural network; an encoding module configured to encode the generated first alternative features using a second part of the first neural network to generate a first encoded representation; a compression module configured to compress the generated first encoding; a decompression module configured to decompress the compressed representation; and a decoding module configured to decode the decompressed representation using a third neural network to reconstruct a first output image.
[0010] Compared with conventional E2E image compression methods, the method and apparatus for rate-adaptive neural image compression using an adversarial generator provided by this application greatly reduce the deployment storage devices for implementing multiple compression rates; and can adapt to a flexible framework for various types of NIC models. In addition, the attention-based generator can focus on significant information for model adaptation. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a diagram of an environment in which the methods, apparatuses, and systems described herein may be implemented according to an embodiment.
[0012] Figure 2 is Figure 1 a block diagram of example components of one or more devices.
[0013] Figure 3 is a block diagram of a test apparatus for rate-adaptive neural image compression using an encoder-side adversarial generator according to an embodiment.
[0014] Figure 4 is a block diagram of a training apparatus for rate-adaptive neural image compression using an encoder-side adversarial generator according to an embodiment.
[0015] Figure 5 is a block diagram of a test apparatus for rate-adaptive neural image compression using a decoder-side adversarial generator according to an embodiment.
[0016] Figure 6 is a block diagram of a training apparatus for rate-adaptive neural image compression using a decoder-side adversarial generator according to an embodiment.
[0017] Figure 7A 、 Figure 7B and Figure 7C is a block diagram of a test apparatus for rate-adaptive neural image compression using an attention-based adversarial generator according to an embodiment.
[0018] Figure 8A 、 Figure 8B and Figure 8C is a block diagram of a training apparatus for rate-adaptive neural image compression using an attention-based adversarial generator according to an embodiment.
[0019] Figure 9 is a flowchart of a method for rate-adaptive neural image compression using an adversarial generator according to an embodiment.
[0020] Figure 10 is a block diagram of an apparatus for rate-adaptive neural image compression using an adversarial generator according to an embodiment. Detailed implementation manners
[0021] The present disclosure describes methods and apparatuses for compressing an input image through a NIC framework with an adaptive compression rate. Only a few model instances are trained for an anchor compression rate, and a compact adversarial generator is used to achieve intermediate compression rates by adapting the anchor model instances. In addition, an attention-based adversarial generator is used on the encoder side or the decoder side to achieve intermediate compression rates by adapting the anchor model instances.
[0022] Figure 1 FIG. 100 is a diagram of an environment 100 in which the methods, apparatuses, and systems described herein can be implemented, according to an embodiment.
[0023] As Figure 1 shown, the environment 100 may include a user device 110, a platform 120, and a network 130. Devices in the environment 100 may be interconnected via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0024] The user device 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with the platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device. In some implementations, the user device 110 may receive information from the platform 120 and / or send information to the platform 120.
[0025] The platform 120 includes one or more devices as described elsewhere herein. In some implementations, the platform 120 may include a cloud server or a group of cloud servers. In some implementations, the platform 120 may be designed to be modular such that software components can be swapped in or out. Thus, the platform 120 can be easily and / or quickly reconfigured for different uses.
[0026] In some implementations, as shown, the platform 120 may be hosted in a cloud computing environment 122. It should be noted that although the implementations described herein describe the platform 120 as being hosted in the cloud computing environment 122, in some implementations, the platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0027] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 can provide services such as computing, software, data access, storage, etc., without the end user (e.g., the user device 110) having to know the physical location and configuration of the systems and / or devices of the hosting platform 120. As shown, the cloud computing environment 122 can include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").
[0028] The computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, the computing resources 124 can host the platform 120. Cloud resources can include: computing instances executed in the computing resources 124, storage devices provided in the computing resources 124, data transfer devices provided by the computing resources 124, etc. In some implementations, the computing resources 124 can communicate with other computing resources 124 via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0029] As further shown in Figure 1 the computing resources 124 include a set of cloud resources, such as one or more applications ("Application, APP") 124-1, one or more virtual machines ("Virtual Machine, VM") 124-2, virtualized storage ("Virtualized Storage, VS") 124-3, one or more hypervisors ("Hypervisor, HYP") 124-4, etc.
[0030] The application 124-1 includes one or more software applications that can be provided to and / or accessed by the user device 110 and / or the platform 120. The application 124-1 can eliminate the need to install and execute software applications on the user device 110. For example, the application 124-1 can include software associated with the platform 120 and / or any other software that can be provided via the cloud computing environment 122. In some implementations, one application 124-1 can send information to / receive information from one or more other applications 124-1 via the virtual machine 124-2.
[0031] The virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. The virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the usage and correspondence of the virtual machine 124-2 to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (OS). A process virtual machine can execute a single program and can support a single process. In some implementations, the virtual machine 124-2 can execute on behalf of a user (e.g., the user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.
[0032] The virtualized storage device 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resource 124. In some implementations, within the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage, such that the storage system can be accessed without regard to the physical storage or heterogeneous structure. The separation can allow the administrator of the storage system flexibility in how the administrator manages storage for end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location where the file is physically stored. This can enable optimization of storage usage, server consolidation, and / or the performance of uninterrupted file migration.
[0033] The hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to execute simultaneously on a host computer such as the computing resource 124. The hypervisor 124-4 can present a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share the virtualized hardware resources.
[0034] Network 130 includes one or more wired networks and / or wireless networks. For example, network 130 may include a cellular network (e.g., a Fifth Generation (5G) network, a Long-Term Evolution (LTE) network, a Third Generation (3G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-based network, etc. and / or a combination of these or other types of networks.
[0035] Figure 1 The number and arrangement of the devices and networks shown are provided as an example. In fact, compared with Figure 1 the devices and / or networks shown, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or differently arranged devices and / or networks. Additionally, Figure 1 two or more of the devices shown may be implemented within a single device, or Figure 1 a single device shown may be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of environment 100 may perform one or more functions described as being performed by another set of devices of environment 100.
[0036] Figure 2 is Figure 1 a block diagram of example components of one or more of the devices.
[0037] Device 200 may correspond to user device 110 and / or platform 120. As Figure 2 shown, device 200 may include a bus 210, a processor 220, a memory 230, a storage device 240, an input interface 250, an output interface 260, and a communication interface 270.
[0038] The bus 210 includes components that permit communication among the components of the device 200. The processor 220 is implemented in hardware, firmware, or a combination of hardware and software. The processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, the processor 220 includes one or more processors that can be programmed to perform functions. The memory 230 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by the processor 220.
[0039] The storage device 240 stores information and / or software related to the operation and use of the device 200. For example, the storage device 240 can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-transitory computer-readable medium and a corresponding drive.
[0040] The input interface 250 includes components that permit the device 200 to receive information such as via a user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone). Additionally or alternatively, the input interface 250 can include sensors for sensing information (e.g., global positioning system (GPS) components, an accelerometer, a gyroscope, and / or an actuator). The output interface 260 includes components that provide output information from the device 200 (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)).
[0041] The communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 270 may allow the device 200 to receive information from and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0042] The device 200 may perform one or more of the processes described herein. The device 200 may perform these processes in response to the processor 220 executing software instructions stored by a non-transitory computer-readable medium, such as the memory 230 and / or the storage device 240. The computer-readable medium is defined herein as a non-transitory memory device. The memory device includes storage space within a single physical storage device or storage space distributed across multiple physical storage devices.
[0043] Software instructions may be read into the memory 230 and / or the storage device 240 from another computer-readable medium or from another device via the communication interface 270. The software instructions stored in the memory 230 and / or the storage device 240, when executed, may cause the processor 220 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more of the processes described herein. Thus, the implementations described herein are not limited to any particular combination of hardware circuitry and software.
[0044] Figure 2 The number and arrangement of the components shown are provided as an example. In fact, compared with the Figure 2 components shown, the device 200 may include additional components, fewer components, different components, or differently arranged components. Additionally or alternatively, a group of components (e.g., one or more components) of the device 200 may perform one or more functions described as being performed by another group of components of the device 200.
[0045] A method and apparatus for rate-adaptive neural image compression using an adversarial generator will now be described in detail.
[0046] The embodiments described herein include a multi-compression rate NIC framework in which, for several anchor compression rates, only a few NIC model instances are learned and deployed, while other intermediate compression rates are achieved by adapting the anchor model instances using a compact adversarial generator on the encoder side or decoder side or using an attention-based adversarial generator to adapt the anchor model instances. The generator is a compact DNN that can be used as a plug-in component added to any underlying NIC model (e.g., before the NIC model or between two layers of the NIC model), and the generator is designed to generate an alternative of the feature according to the features of the original NIC model (e.g., the input image if placed before the NIC model, or the intermediate feature map if placed between two layers). Thus, the newly generated alternative can obtain the desired compression rate.
[0047] Given an input image x, the objective of the test phase of the NIC workflow can be described as follows. Compute a compressed representation that is compact for storage and transmission Then, based on the compressed representation reconstruct the output image And the reconstructed output image may be similar to the original input image x. The process of computing the compressed representation can be divided into two parts: DNN encoding processing using a test DNN encoder to compute the DNN encoded representation y; and then an encoding process in which y is encoded by the test encoder (usually including quantization and entropy encoding) to generate the compressed representation Thus, the decoding process is divided into two parts: a decoding process in which the compressed representation is decoded by the test decoder (usually including decoding and dequantization) to generate the recovered ; and then a DNN decoding process in which the test DNN decoder uses the recovered representation to reconstruct the image In the present disclosure, there is no limitation on the network structure of the test DNN encoder for DNN encoding or the network structure of the test DNN decoder for DNN decoding. There is also no limitation on the methods for encoding or decoding (quantization methods and entropy encoding methods).
[0048] To learn the NIC model, two competing objectives need to be addressed: better reconstruction quality and less bit consumption. Use a loss function to measure the reconstruction error, which is referred to as the distortion loss, such as Peak Signal-to-Noise Ratio (PSNR) and / or Structural Similarity Index Measure (SSIM). Compute the rate loss In measurement compression representation of the bit consumption. Therefore, the trade-off hyperparameter λ is used to optimize the joint R-D loss:
[0049]
[0050] Training with a large hyperparameter λ produces a compression model with less distortion but more bit consumption, while training with a small hyperparameter λ produces a compression model with more distortion but less bit consumption. Traditionally, for each predefined hyperparameter λ, an NIC model instance will be trained, which will not be applicable to other values of the hyperparameter λ. Therefore, to achieve multiple bitrates for the compression stream, multiple model instances may need to be trained and stored, one model instance for one desired hyperparameter λ.
[0051] In the present disclosure, the rate-adaptive NIC framework uses an additional compact adversarial generator on the encoder side or the decoder side to adapt a model instance trained for one anchor R-D trade-off hyperparameter λ o value to another intermediate hyperparameter λ t value. Therefore, to achieve multi-compression rate NIC, only a few anchor model instances can be trained and deployed for the target anchor R-D trade-off values, while the remaining intermediate R-D trade-off values of interest can be generated by the adversarial generator. Since the adversarial generator is a much smaller compact DNN in terms of both storage and computation, this framework is much more efficient for multi-compression rate NIC compared to the traditional method of training and deploying each model instance for each hyperparameter λ of interest.
[0052] Figure 3 is a block diagram of a test apparatus 300 for rate-adaptive neural image compression using an encoder-side adversarial generator according to an embodiment.
[0053] As Figure 3 shown, the test apparatus 300 includes a test DNN encoder 310, a surrogate generator 320, a test DNN encoder 330, a test encoder 340, a test decoder 350, and a test DNN decoder 360.
[0054] The surrogate generator 320 is an adversarial generator that is an additional component that can be inserted into any existing NIC model. Without loss of generality, there are a test DNN encoder and a test DNN decoder from the NIC DNN, and a model instance M o is trained for the target anchor λ o value. By using the generator to adapt this model instance M o to achieve a virtual NIC model instance M t trained for the target λ tcompression effect. In addition, assume that the generator will be inserted between the i-th layer and the (i + 1)-th layer of the test DNN encoder. The input of the (i + 1)-th layer (the output of the i-th layer) is the feature f. Therefore, the original DNN encoding process of the original NIC model can be divided into two parts, in which the input image x passes through the module of DNN encoding part 1 to calculate the feature f using part 1 of the test DNN encoder 310, and then f passes through the module of DNN encoding part 2 to calculate the DNN encoding representation using part 2 of the test DNN encoder 330
[0055] If the generator is placed before the entire test DNN encoder (i.e., i = 0), then the feature f is the input image x, and part 2 of the test DNN encoder 330 includes the entire original test DNN encoder.
[0056] Using the inserted adversarial generator The feature f passes through the alternative generator 320 to calculate an alternative feature (different or enhanced features that can be used to obtain the desired compression ratio), and (instead of f) passes through part 2 of the test DNN encoder 330 to calculate the DNN encoding representation of the alternative feature Based on The encoding module uses the test encoder 340 to calculate the compressed representation Then, on the decoder side, based on The restored representation can be calculated by using the decoding process of the test decoder 350 Then, the DNN decoding module is based on Uses the test DNN decoder 360 to calculate the reconstructed output image Compressed representation And the reconstructed output image Will have an approximately optimal R-D loss at the target λ t value (i.e., similar to the R-D loss of the virtual model M t trained by optimizing the R-D loss at the λ t case).
[0057] In the present disclosure, there is no restriction on the DNN network structure of the alternative generator 320. The DNN network structure of the alternative generator 320 can be much smaller than the underlying NIC model.
[0058] In an embodiment, for a given target λ t value, select the anchor model instance for adapting the anchor λ o to the target λ t as the one in which λo The model instance closest to λ t In addition, it is worth mentioning that the embodiments of the present disclosure are related to training only one model instance with respect to an anchor λ value, and all other intermediate R-D trade-off values are simulated by different compact generators - one generator for each intermediate λ o value. t
[0059] Figure 4 is a block diagram of a training apparatus 400 for rate-adaptive neural image compression using an encoder-side adversarial generator according to an embodiment.
[0060] As Figure 4 shown, the training apparatus 400 includes a training DNN encoder 405, a surrogate generator 410, a training DNN encoder 415, a training encoder 420, a rate loss generator 425, a training decoder 430, a training DNN decoder 435, a distortion loss generator 440, a training DNN encoder 445, a representation discriminative loss generator 450, a feature discriminative loss generator 455, and a weight update unit 460.
[0061] Pre-train the NIC model as model instance M o to optimize the R-D loss in Equation (1) for the target λ o case. The generator is learned to adapt model M o for another λ of interest t without retraining the underlying NIC model. Similar to the test phase, the corresponding training DNN encoder is divided into 2 parts: training DNN encoder 405 part 1 and training DNN encoder 415 part 2.
[0062] The input training image x (x ∈ S) from the training dataset S first passes through the module of DNN encoding part 1 to calculate the feature f using the training DNN encoder 405 part 1 in the pre-trained model instance M o . Then, the surrogate generator 410 uses the current generator to calculate a surrogate feature (different or enhanced features that can be used to obtain the desired compression rate). Similarly, the surrogate generator 410 is a DNN, and in the embodiment, the surrogate generator 410 calculates a surrogate perturbation based on the input f and the surrogate feature is calculated as Then, the surrogate passes through the module of DNN encoding part 2 to calculate the DNN encoded representation by using the training DNN encoder 415 part 2 in the pre-trained model instance M o Then, the encoding process calculates a compressed representation using the trained encoder 420 Using the rate loss generator 425 calculates the rate loss Then, the decoding module is based on and calculates a decompressed representation by using the trained decoder 430 And the DNN decoding process further generates a reconstructed output image by using the trained DNN decoder 435 The distortion loss generator 440 calculates the distortion loss between the reconstruction and the original input image x The rate loss is related to the bit rate of the encoded representation and, in an implementation, an entropy estimation method is used as the rate loss generator to calculate the rate loss Using the target of interest λ t , the R-D loss of Equation (1) can be calculated as
[0063] Meanwhile, using the original feature f, the module of the DNN encoding part 2 can also use the trained DNN encoder 445 to generate a DNN encoded representation y. Based on both and y, the representation discriminant loss generator 450 calculates the representation discriminant loss by calculating the representation discriminant loss process In an implementation, the representation discriminant loss generator 450 is a DNN that discriminates the encoded feature representation y generated based on the original feature f from the representation generated based on the alternative feature . For example, the representation discriminant loss generator 450 can be a binary DNN classifier that discriminates the representation generated according to the original feature as one class and the representation generated according to the alternative feature as another class. In addition, based on the original feature f and the alternative feature the feature discriminant loss generator 455 can calculate the feature discriminant loss by calculating the feature discriminant loss process In an implementation, the feature discriminant loss generator 455 is a DNN that discriminates the original feature f from the alternative feature . For example, the feature discriminant loss generator 455 can be a binary DNN classifier that discriminates the original feature as one class and the alternative feature as another class. Based on the feature discriminant loss and the representation discriminant loss the weight update unit 460 calculates the adversarial loss as (α is a hyperparameter):
[0064]
[0065] Based on and The weight update unit 460 updates the weight coefficients of the trainable part of the DNN model using gradients through backpropagation optimization.
[0066] In an embodiment, the weight coefficients of the model instance M o (including the trained DNN encoders 405, 415, or 445, the trained encoder 420, the trained decoder 430, and the trained DNN decoder 435) are fixed during the above training phase. Additionally, the rate loss generator 425 is also predetermined and fixed. The weight coefficients of the alternative generator 410, the feature discriminative loss generator 455, and the representation discriminative loss generator 450 can be trained and updated by a Generative Adversarial Network (GAN) training framework through the above training phase. For example, in an embodiment, the R-D loss gradient is used to update the weight coefficients of the alternative generator 410, and the adversarial loss gradient is used to update the weight coefficients of the feature discriminative loss generator 455 and the representation discriminative loss generator 450.
[0067] In the present disclosure, there are no restrictions on the pre-training process for determining the model instance M o and the rate loss generator 425. As an example, in an embodiment, a set of training images S pre is used in the pre-training process, and this set of training images S pre can be the same as or different from the training data set S. For each image x ∈ S pre , the same forward inference calculation is performed through DNN encoding, encoding, decoding, and DNN decoding to calculate the encoded representation and the reconstructed Then, the distortion loss and the rate loss can be calculated. Then, given the pre-training hyperparameter λ pre , the overall R-D loss can be calculated based on Equation (1). The gradient of the overall R-D loss is used to update the weights of the trained DNN encoders 405, 415, or 445, the trained encoder 420, the trained decoder 430, the trained DNN decoder 435, and the rate loss generator 425 through backpropagation.
[0068] It is also worth mentioning that, in the embodiment, the training DNN encoder 405 part 1, the training DNN encoder 415 part 2, and the training DNN decoder 435 are the same as the corresponding test DNN encoder 310 part 1, the test DNN encoder 330 part 2, and the test DNN decoder 360. However, the training encoder 420 and the training decoder 430 are different from the corresponding test encoder 340 and test decoder 350. For example, the test encoder 340 and the test decoder 350 respectively include a general test quantizer and a test entropy encoder, and a general test entropy decoder and a test dequantizer. Each of the training encoder 420 and the training decoder 430 uses a statistical sampler to approximate the effects of the test quantizer and the test dequantizer respectively. The entropy encoder and the entropy decoder are skipped during the training phase.
[0069] Figure 5 It is a block diagram of a test device 500 for rate-adaptive neural image compression using a decoder-side adversarial generator according to an embodiment.
[0070] As Figure 5 shown, the test device 500 includes a test DNN encoder 510, a test encoder 520, a test decoder 530, a test DNN decoder 540, an alternative generator 550, and a test DNN decoder 560.
[0071] The alternative generator 550 is an additional component that can be inserted into any existing NIC model. Without loss of generality, there are a test DNN encoder and a test DNN decoder from the NIC DNN, and the model instance M is trained for the target anchor λ o value. o By using the generator to adapt the model instance M o , to achieve the compression effect of the virtual NIC model instance M t trained for the target λ t value. In addition, it is assumed that the generator will be inserted between the i-th layer and the (i + 1)-th layer of the test DNN decoder. The input of the (i + 1)-th layer (the output of the i-th layer) is the feature f. Therefore, the original DNN decoding process of the original NIC model can be divided into two parts, in which the input recovery representation passes through the module of the DNN decoding part 1 to calculate the feature f using the test DNN decoder 540 part 1, and then f passes through the module of the DNN decoding part 2 to calculate the reconstructed
[0072] If the generator is placed before the entire test DNN decoder (i.e., i = 0), then the feature f is the input recovery representation And the test DNN decoder 560 part 2 includes the entire original test DNN decoder.
[0073] Thus, given an input image x, the DNN encoding process uses the test DNN encoder 510 to compute a DNN encoded representation This DNN encoded representation y is further encoded by the test encoder 520 during the encoding process to generate a compressed representation Then, on the decoder side, the compressed representation is decoded by the test decoder 530 to generate a recovered representation in the decoding module Then, the module of the DNN decoding part 1 is based on using the test DNN decoder 540 part 1 to compute the feature f. With the inserted adversarial generator The feature f is passed through the substitute generator 550 to compute a substitute feature And (instead of f) is passed through the module of the DNN decoding part 2 to compute the reconstructed output image by the test DNN decoder 560 part 2 The compressed representation and the reconstructed output image will have a near-optimal R-D loss at the target λ t value of Equation (1) (i.e., an R-D loss similar to that of the virtual model M t trained by optimizing the R-D loss at the λ t case).
[0074] In the present disclosure, there is no restriction on the DNN network structure of the generator. The DNN network structure of the generator can be much smaller than the underlying NIC model.
[0075] In an embodiment, for a given target λ t value, the anchor model instance selected for adapting the anchor λ o to the target λ t is the model instance where λ o is closest to λ t . Additionally, it is worth mentioning that the embodiments of the present disclosure are about training only one model instance for one anchor λ o value, and simulating all other intermediate R-D trade-off values through different compact generators - one generator for each intermediate λ t .
[0076] Figure 6 is a block diagram of a training apparatus 600 for rate-adaptive neural image compression using an adversarial generator on the decoder side according to an embodiment.
[0077] As Figure 6As shown, the training device 600 includes a training DNN encoder 605, a training encoder 610, a rate loss generator 615, a training decoder 620, a training DNN decoder 625, an alternative generator 630, a training DNN decoder 635, a distortion loss generator 640, a training DNN decoder 645, a reconstruction discrimination loss generator 650, a feature discrimination loss generator 655, and a weight update unit 660.
[0078] Pre-train the NIC model as the model instance M o , to optimize the R-D loss of Equation (1) at the target λ o Case. Make the generator Learn to make the model M o Adapt for another λ of interest t , without retraining the NIC model. Similar to the test phase, the corresponding training DNN decoder is divided into two parts: training DNN decoder 625 part 1 and training DNN decoder 635 part 2.
[0079] The input training image x (x ∈ S) from the training dataset S first passes through the DNN encoding module to calculate the DNN encoding representation based on the training DNN encoder 605 Then, the encoding process uses the training encoder 610 to calculate the compressed representation Based on The rate loss generator 615 calculates the rate loss Then, on the decoder side, the decoding module is based on Calculate the decompressed representation by using the training decoder 620 Then, the decompressed Pass through the module of DNN decoding part 1 to use the trained model instance M o In the training DNN decoder 625 part 1 to calculate the feature f. Then, the alternative generator 630 uses the current generator To calculate the alternative feature Similarly, the alternative generator 630 is a DNN, and in the embodiment, the alternative generator 630 calculates the alternative perturbation δ(f) based on the input f, and the alternative feature Is calculated as Then, the alternative Pass through the module of DNN decoding part 2 to calculate the reconstructed output image by using the trained model instance M o In the training DNN decoder 635 part 2 The distortion loss generator 640 calculates the Distortion loss between the reconstruction and the original input image x Rate loss And the coding representation is related to the bit rate, and in an embodiment, an entropy estimation method is used as the rate loss generator 615 to calculate the rate loss Using the target λ of interest t , the R-D loss of Equation (1) can be calculated as
[0080] Meanwhile, using the original feature f, the module of the DNN decoder 645 part 2 can also calculate the reconstructed output image Based on and both, the reconstruction discrimination loss generator 650 calculates the reconstruction discrimination loss by calculating the reconstruction discrimination loss process In an embodiment, the reconstruction discrimination loss generator 650 is a DNN that discriminates between the reconstructed output image generated based on the original feature f and the reconstructed output image generated based on the alternative feature . For example, the reconstruction discrimination loss generator 650 can be a binary DNN classifier that discriminates the reconstructed output image generated according to the original feature as one class and the reconstructed output image generated according to the alternative feature as another class. In addition, based on the original feature f and the alternative feature the feature discrimination loss generator 655 can calculate the feature discrimination loss by calculating the feature discrimination loss process In an embodiment, the feature discrimination loss generator 655 is a DNN that discriminates between the original feature f and the alternative feature . For example, the feature discrimination loss generator 655 can be a binary DNN classifier that discriminates the original feature as one class and the alternative feature as another class. Based on the feature discrimination loss and the reconstruction discrimination loss the weight update unit 660 calculates the adversarial loss as (α as a hyperparameter):
[0081]
[0082] Based on and the weight update unit 660 updates the weight coefficients of the trainable part of the DNN model using the gradient through backpropagation optimization
[0083] In an embodiment, the model instance M oThe weight coefficients of (including the training DNN encoder 605, the training encoder 610, the training decoder 620, and the training DNN decoder 625, 635, or 645) are fixed during the above training phase. Additionally, the rate loss generator 615 is also predetermined and fixed. The weight coefficients of the alternative generator 630, the feature discriminative loss generator 655, and the reconstruction discriminative loss generator 650 can be trained and updated by the GAN training framework through the above training phase. For example, in an embodiment, the R-D loss gradient is used to update the weight coefficients of the alternative generator 630, and the adversarial loss gradient is used to update the weight coefficients of the feature discriminative loss generator 655 and the reconstruction discriminative loss generator 650.
[0084] In the present disclosure, there are no restrictions on the pre-training process for determining the model instance M o and the rate loss generator 615. As an example, in an embodiment, a set of training images S pre is used in the pre-training process, and this set of training images S pre can be the same as or different from the training dataset S. For each image x ∈ S pre , the same forward inference calculation is performed through DNN encoding, encoding, decoding, DNN decoding to calculate the encoded representation and the reconstructed Then, the distortion loss and the rate loss can be calculated. Then, given the pre-training hyperparameter λ pre , the overall R-D loss can be calculated based on Equation (1). The gradient of the overall R-D loss is used to update the weights of the training DNN encoder 605, the training encoder 610, the training decoder 620, the training DNN decoder 625, 635, or 645, and the rate loss generator 615 through backpropagation.
[0085] It is also worth mentioning that, in an embodiment, the training DNN encoder 605, the training DNN decoder 625 part 1, and the training DNN decoder 635 part 2 are the same as the corresponding test DNN encoder 510, the test DNN decoder 540 part 1, and the test DNN decoder 560 part 2. However, the training encoder 610 and the training decoder 620 are different from the corresponding test encoder 520 and the test decoder 530. For example, the test encoder 520 and the test decoder 530 include a universal test quantizer and a test entropy encoder and a universal test entropy decoder and a test dequantizer, respectively. Each of the training encoder 610 and the training decoder 620 uses a statistical sampler to approximate the effects of the test quantizer and the test dequantizer, respectively. The entropy encoder and the entropy decoder are skipped in the training phase.
[0086] Figure 7A , Figure 7B and Figure 7C is a block diagram of a test setup 700A, 700B, and 700C for rate-adaptive neural image compression using an attention-based adversarial generator, according to an embodiment.
[0087] Attention-based adversarial generators can be used on the encoder side ( Figure 7A and Figure 7B ) or on the decoder side ( Figure 7C ) is an additional component inserted into any existing NIC model. The attention-based adversarial generator uses the attention map generated by the attention model to automatically focus on important information during RD trade-off adaptation. Since the attention-based adversarial generator uses the output of the attention model, Figure 7A and Figure 7B ) or decoder side ( Figure 7C ) requires putting the attention model before the generator.
[0088] Without loss of generality, there is a test DNN encoder and a test DNN decoder from the NIC DNN, and the target anchor λ o Value training model instance M o By using the generator To adapt the model instance M o , with the goal of achieving t The virtual NIC model instance M for training t In addition, assuming that the generator It will be inserted between the i-th layer and the (i+1)-th layer in the test DNN encoder or between the i-th layer and the (i+1)-th layer in the test DNN decoder. The input of the (i+1)-th layer (the output of the i-th layer) is feature f.
[0089] like Figure 7AAs shown, the test device 700A includes a test DNN encoder 705, an alternative generator 710, a test DNN encoder 715, a test encoder 720, a test decoder 725, a test DNN decoder 730, and an attention generator 735.
[0090] When the alternative generator 710 is placed on the encoder side, the original DNN encoding process of the original NIC model can be divided into two parts. In these two parts, the input image x passes through the modules of DNN encoding part 1 to calculate the feature f using part 1 of the test DNN encoder 705, and then f passes through the modules of DNN encoding part 2 to calculate the DNN encoding representation using part 2 of the test DNN encoder 715. If the alternative generator 710 is placed before the entire test DNN encoder (i.e., i = 0), then the feature f is the input image x, and part 2 of the test DNN encoder 715 includes the entire original test DNN encoder.
[0091] As Figure 7B and Figure 7C shown, the test device 700B or 700C includes a test DNN encoder 740, a test encoder 745, a test decoder 750, a test DNN decoder 755, an alternative generator 760, and a test DNN decoder 765. Figure 7B The test device 700B of Figure 7C includes an attention generator 770, and
[0092] the test device 700C of includes an attention generator 775. When the alternative generator 760 is placed on the decoder side, the original DNN decoding process of the original NIC model can be divided into two parts. In these two parts, the input recovery representation passes through the modules of DNN decoding part 1 to calculate the feature f using part 1 of the test DNN decoder 755, and then f passes through the modules of DNN decoding part 2 to calculate the reconstructed If the alternative generator 760 is placed before the entire test DNN decoder (i.e., i = 0), then the feature f is the input recovery representation
[0093] Referring to Figures 7A to 7C , without loss of generality, there is an attention model that is a DNN and will be inserted between the j-th layer and the (j + 1)-th layer in the test DNN encoder or between the j-th layer and the (j + 1)-th layer (j ≤ i) in the test DNN decoder. The input of the (j + 1)-th layer (the output of the j-th layer) is the feature a, and the attention generator generates an attention map based on a. Therefore, when the attention model is placed before the entire test DNN encoder (i.e., j = 0), the feature a is the input image x. When the attention model is placed before the entire test DNN decoder, the feature a is the input recovery representation
[0094] Referring to Figure 7A , given the input image x, for the configuration in which the generator is placed on the encoder side, x passes through the module of the DNN encoding part 1 to calculate the feature f using the test DNN encoder 705 part 1. In addition, the feature a, which is the output of the j-th layer and the input of the (j + 1)-th layer of the test DNN encoder 705 part 1, passes through the attention generator 735 to generate an attention map by using the attention model Then, the feature f and the attention map pass through the substitution generator 710 to calculate the substitution feature And (instead of f) passes through the module of the DNN encoding part 2 to calculate the DNN encoded representation Based on The encoding module uses the test encoder 720 to calculate the compressed representation Then, on the decoder side, based on The restored representation can be calculated using the test decoder 725 through the decoding process Then, the DNN decoding module is based on Uses the test DNN decoder 730 to calculate the reconstructed output image Compressed representation And the reconstructed output image Will have an approximately optimal R-D loss in the case of the target λ t value (i.e., similar to the R-D loss of the virtual model M t trained by optimizing the R-D loss in the case of λ t ).
[0095] Referring to Figure 7B And Figure 7C , for the configuration in which the generator is placed on the decoder side, using the input image x, the DNN encoding process uses the test DNN encoder 740 to calculate the DNN encoded representation This DNN encoded representation y is further encoded by the test encoder 745 during the encoding process to generate the compressed representation Then, on the decoder side, the compressed representation is decoded by the test decoder 750 to generate the restored Then, the module of the DNN decoding part 1 is based on Compute the feature f using the test DNN decoder 755 part 1.
[0096] For Figure 7B the configuration described in
[0097] For Figure 7C the configuration described in
[0098] Then, utilize the inserted adversarial generator feature f and the attention map to compute the surrogate feature through the surrogate generator 760 And (instead of f) through the module of the DNN decoding part 2 to compute the reconstructed output image through the test DNN decoder 765 part 2 compressed representation and the reconstructed output image will have a near-optimal R-D loss at the target λ t value of Equation (1) (i.e., an R-D loss similar to that of the virtual model M t trained by optimizing the R-D loss at the λ t case).
[0099] In this disclosure, there is no restriction on the DNN network structure of the attention model. In an embodiment, the attention map will have the same shape as the feature f, and the larger the value of the attention map, the more important the corresponding feature in f, and the smaller the value of the attention map, the less important the corresponding feature in f.
[0100] In this disclosure, there is no restriction on the DNN network structure of the surrogate generator 710 or 760. In an embodiment, the surrogate generator 710 or 760 is much smaller than the underlying NIC model, and the input feature f and the input attention map are combined, for example, through element-wise multiplication to generate the attention-masked input to the generator DNN.
[0101] In an embodiment, for a given target λ t value, select the anchor λ o to adapt to the target λ tThe anchor model instance as λ o The model instance closest to λ t Moreover, it is worth mentioning that the embodiments of the present disclosure relate to training only one model instance for an anchor λ o value, and simulating all other intermediate R-D trade-off values through different compact generators - one generator for each intermediate λ t value.
[0102] Figure 8A 、 Figure 8B and Figure 8C are block diagrams of training apparatuses 800A, 800B, and 800C for rate-adaptive neural image compression using an attention-based adversarial generator according to an embodiment.
[0103] There exists a NIC model that is pre-trained as a model instance M o to optimize the R-D loss in the case of the target λ o . The generator is learned to adapt the model M o for another λ t of interest without retraining the underlying NIC model. Similar to the corresponding test phase, Figure 8A shows the workflow when the generator is placed on the encoder side, in which the corresponding training DNN encoder is divided into two parts: training DNN encoder 802 part 1 and training DNN encoder 806 part 2. Figure 8B and Figure 8C show the workflow when the generator is placed on the decoder side, in which the corresponding training DNN decoder is divided into two parts: training DNN
[0104] decoder 834 part 1 and training DNN decoder 838 part 2.
[0105] As Figure 8A shown, the training apparatus 800A includes a training DNN encoder 802, an alternative generator
[0106] 804, a training DNN encoder 806, a training encoder 808, a training decoder 810, training
[0107] DNN decoder 812, an attention generator 814, a rate loss generator 816, a distortion loss generator 818, a feature discrimination loss generator 820, a training DNN encoder 822, a representation discrimination loss generator 824, and a weight update unit 826.
[0108] When the generator is placed on the encoder side, given an input training image x (x ∈ S) from the training dataset S, the input training image x first passes through the module of the DNN encoding part 1 to calculate the feature f using the trained DNN encoder 802 part 1 in the pre-trained model instance M o Moreover, the feature a, which is the output of the j-th layer and the input of the (j + 1)-th layer of the trained DNN encoder 802 part 1, passes through the attention generator 814 to generate an attention map by using the attention model Then, the feature f and the attention map pass through the alternative generator 804 to calculate the alternative feature using the current generator The alternative feature In an embodiment, the attention map and the feature f have the same shape, and the attention map and the feature f are combined, for example, by element-wise multiplication to generate the attention-masked input to calculate the alternative perturbation δ(f) by the alternative generator 804, which is based on the attention-masked input. The alternative feature is calculated as Then, the alternative passes through the module of the DNN encoding part 2 to calculate the DNN encoding representation using the trained DNN encoder 806 part 2 in the pre-trained model instance M o Then, the encoding process calculates the compressed representation using the trained encoder 808 The rate loss generator 816 calculates the rate loss using the rate loss generator 816 calculates the rate loss Then, the decoding module calculates the decompressed by using the trained decoder 810 based on and the DNN decoding process further generates the reconstructed The distortion loss generator 818 calculates the reconstruction distortion loss between the reconstructed and the original input image x The rate loss is related to the bit rate of the encoded representation and, in an embodiment, the rate loss generator 816 calculates the rate loss using an entropy estimation method Using the target of interest λ t the R-D loss of Equation (1) can be calculated as
[0109] As Figure 8B and Figure 8CAs shown, the training device 800B or 800C includes a training DNN encoder 828, a training encoder 830, a training decoder 832, a training DNN decoder 834, an alternative generator 836, a training DNN decoder 838, a rate loss generator 842, a distortion loss generator 844, a feature discrimination loss generator 846, a training DNN decoder 848, a reconstruction discrimination loss generator 850, and a weight update unit 852. Figure 8B The training device 800B includes an attention generator 840, and Figure 8C the training device 800C includes an attention generator 854.
[0110] Referring to Figure 8B and Figure 8C When the generator is placed on the decoder side, the input training image x (x ∈ S) from the training dataset S first passes through the DNN encoding module to calculate the DNN encoded representation y based on the training DNN encoder 828. Then, the encoding process uses the training encoder 830 to calculate the compressed representation Based on the rate loss generator 842 calculates the rate loss Then, on the decoder side, the decoding module is based on to calculate the decompressed using the training decoder 832 Then, the decompressed o passes through the module of DNN decoding part 1 to calculate the feature f using the training DNN decoder 834 part 1 in the pre-trained model instance M
[0111] For Figure 8B the configuration described in, the feature a, which is the output of the j-th layer and the input of the (j + 1)-th layer of the training DNN encoder 828, passes through the attention generator 840 to generate an attention map using the attention model Utilizing the inserted adversarial generator The feature f and the attention map pass through the alternative generator 836 to calculate the alternative feature Similar to Figure 8A In the embodiment, the attention map and the feature f have the same shape, and the attention map and the feature f are combined, for example, by element-wise multiplication to generate the attention-masked input to pass through the alternative generator 836, and the alternative generator 836 calculates the alternative perturbation δ(f) based on the attention-masked input. The alternative feature is calculated as Then, the alternative passes through the module of DNN decoding part 2 to pass through the pre-trained model instance Mo The trained DNN decoder 838 part 2 in to compute the reconstructed output image The distortion loss generator 844 computes the reconstructed distortion loss between and the original input image x The rate loss is related to the bit rate of the encoded representation t and, in an implementation, the rate loss generator 842 uses an entropy estimation method to compute the rate loss
[0112] For Figure 8C the configuration of the feature a that is the output of the j-th layer and the input of the (j + 1)-th layer of the training DNN decoder 834 part 1 passes through the attention generator to generate an attention map by using the attention model Figure 8B Similar to the above case, the feature f and the attention map are passed through the substitution generator 836 to compute the substitution feature where, in an implementation, the attention map and the feature f have the same shape, and the attention map and the feature f are combined, for example, by element-wise multiplication to generate the attention-masked input to pass through the substitution generator 836. The substitution generator 836 computes the substitution perturbation δ(f) based on the attention-masked input, and the substitution feature is then computed as The substitution o is then passed through the module of the DNN decoding part 2 to compute the reconstructed output image by using the pre-trained model instance M The distortion loss generator 844 computes the reconstructed distortion loss between and the original input image x The rate loss is related to the bit rate of the encoded representation and, in an implementation, the rate loss generator 842 uses an entropy estimation method to compute the rate loss t Using the target of interest λ
[0113] Referring to Figures 8A to 8C at the same time, based on the original feature f and the substitution feature The feature discrimination loss generator 820 or 846 can calculate the feature discrimination loss by computing the feature discrimination loss process. In an embodiment, the feature discrimination loss generator 820 or 846 is a DNN that discriminates between the original feature f and the alternative feature For example, the feature discrimination loss generator 820 or 846 can be a binary DNN classifier that discriminates the features of the original attention mask as one class and the alternative features as another class.
[0114] In addition, when the generator is on the encoder side as in Figure 8A Using the original feature f, training the DNN encoder 822 part 2 can also generate the DNN encoded representation y. Based on Both and y, the representation discrimination loss generator 824 calculates the representation discrimination loss by computing the representation discrimination loss process In an embodiment, the representation discrimination loss generator 824 is a DNN that discriminates between the encoded feature representation y generated based on the original feature f and the representation Generated based on the alternative feature For example, the representation discrimination loss generator 824 can be a binary DNN classifier that discriminates the representation generated from the original feature as one class and the representation generated from the alternative feature as another class.
[0115] When the generator is on the decoder side as in Figure 8B And Figure 8C Using the original feature f, training the DNN decoder 848 part 2 can also calculate the reconstructed output image Based on And Both, the reconstruction discrimination loss generator 850 calculates the reconstruction discrimination loss by computing the reconstruction discrimination loss process In an embodiment, the reconstruction discrimination loss generator 850 is for the reconstructed output image generated based on the original feature f And the reconstructed output image generated based on the alternative feature Generated For example, the reconstruction discrimination loss generator 850 can be a binary DNN classifier that discriminates the reconstructed output image generated from the original feature as one class and the reconstructed output image generated from the alternative feature as another class.
[0116] When the generator is on the encoder side as shown in Figure 8A Based on the feature discrimination loss And the representation discrimination loss The weight update unit 826 calculates the adversarial loss As (α is a hyperparameter):
[0117]
[0118] and use and The weight update unit 826 updates the weight coefficients of the trainable part of the DNN model using gradient-based backpropagation optimization. In an embodiment, the weight coefficients of the model instance M o (including the training DNN encoder 802 part 1, the training DNN encoder 806 part 2, the training encoder 808, the training decoder 810, and the training DNN decoder 812) are fixed during the above training phase. Additionally, the rate loss generator 816 is also predetermined and fixed. The weight coefficients of the alternative generator 804, the feature discriminative loss generator 820, and the representation discriminative loss generator 824 can be trained and updated by the GAN training framework through the above training phase. For example, in an embodiment, the R-D loss gradient is used to update the weight coefficients of the alternative generator 804, and the adversarial loss gradient is used to update the weight coefficients of the feature discriminative loss generator 820 and the representation discriminative loss generator 824.
[0119] When the generator is on the decoder side as Figure 8B and Figure 8C shown, based on the feature discriminative loss and the reconstruction discriminative loss the weight update unit 852 calculates the adversarial loss as (α being a hyperparameter):
[0120]
[0121] Based on and the weight update unit 852 updates the weight coefficients of the trainable part of the DNN model using gradient-based backpropagation optimization. In an embodiment, the weight coefficients of the model instance M o (including the training DNN encoder 828, the training encoder 830, the training decoder 832, the training DNN decoder 834 part 1, and the training DNN decoder 838 part 2) are fixed during the above training phase. Additionally, the rate loss generator 842 is also predetermined and fixed. The weight coefficients of the alternative generator 836, the feature discriminative loss generator 846, and the reconstruction discriminative loss generator 850 can be trained and updated by the GAN training framework through the above training phase. For example, in an embodiment, the R-D loss gradient is used to update the weight coefficients of the alternative generator 836, and the adversarial loss to update the weight coefficients of the feature discrimination loss generator 846 and the reconstruction discrimination loss generator 850 using the gradient.
[0122] In the present disclosure, there is no limitation on the pre-training process for determining the model instance M o and the rate loss generator 816 or 842. As an example, in an embodiment, a set of training images S pre is used in the pre-training process, and this set of training images S pre can be the same as or different from the training data set S. For each image x ∈ S pre , the same forward inference calculation is performed through DNN encoding, encoding, decoding, and DNN decoding to calculate the encoded representation and the reconstructed Then, the distortion loss and the rate loss can be calculated. Then, given the pre-training hyperparameter λ pre , the overall R-D loss can be calculated based on Equation (1) Using the gradient of the overall R-D loss to update the weights of the training DNN encoders 802, 806 or 828, the training encoders 808 or 830, the training decoders 810 or 832, the training DNN decoders 812, 834 or 838, and the rate loss generator 816 or 842 through backpropagation.
[0123] It is also worth mentioning that, in an embodiment, for Figure 8A the case of the encoder-side alternative generator 804 in Figure 8B and Figure 8CIn the case of the decoder-side alternative generator 836, the trained DNN encoder 828, the trained DNN decoder 834 part 1, and the trained DNN decoder 838 part 2 are the same as the corresponding test DNN encoder 740, the test DNN decoder 755 part 1, and the test DNN decoder 765 part 2. On the other hand, the trained encoder 808 or 830 and the trained decoder 810 or 832 are different from the corresponding test encoder 720 or 745 and the test decoder 725 or 750. For example, the test encoder 720 or 745 and the test decoder 725 or 75 respectively include a general test quantizer and a test entropy encoder, as well as a general test entropy decoder and a test dequantizer. Each of the trained encoder 808 or 830 and the trained decoder 810 or 832 uses a statistical sampler to approximate the effects of the test quantizer and the test dequantizer, respectively. The entropy encoder and the entropy decoder are skipped during the training phase.
[0124] Figure 9 is a flowchart of a method 900 for rate-adaptive neural image compression using an adversarial generator according to an embodiment.
[0125] In some implementations, Figure 9 one or more of the processing blocks in may be executed by the platform 120. In some implementations, Figure 9 one or more of the processing blocks in may be executed by another device or a group of devices such as the user device 110 that is separate from or includes the platform 120.
[0126] As Figure 9 shown, in operation 910, the method 900 includes obtaining first features of an input image using a first part of a first neural network.
[0127] In operation 920, the method 900 includes generating first alternative features using a second neural network based on the obtained first features.
[0128] In operation 930, the method 900 includes encoding the generated first alternative features using a second part of the first neural network to generate a first encoded representation.
[0129] In operation 940, the method 900 includes compressing the generated first encoded representation.
[0130] In operation 950, the method 900 includes decompressing the compressed representation.
[0131] In operation 960, the method 900 includes decoding the decompressed representation using a third neural network to reconstruct a first output image.
[0132] The second neural network can be trained as follows: determining a rate loss of the compressed representation; determining a distortion loss between the input image and the reconstructed first output image; encoding the obtained first features using a third part of the first neural network to generate a second encoded representation; using a fourth neural network to determine a representation discrimination loss between the generated first encoded representation and the generated second encoded representation; using a fifth neural network to determine a feature discrimination loss between the generated first alternative features and the obtained first features; and updating the weight coefficients of the second neural network, the fourth neural network, and the fifth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
[0133] Method 900 may further include: encoding an input image using a first neural network to generate a first encoded representation; obtaining second features from the decompressed representation using a first part of a third neural network; generating second alternative features using a fourth neural network based on the obtained second features; and decoding the generated second alternative features using a second part of the third neural network to reconstruct a first output image.
[0134] The fourth neural network can be trained as follows: determining a rate loss of the compressed representation; determining a distortion loss between the input image and the reconstructed first output image; decoding the obtained second features using a third part of the third neural network to reconstruct a second output image; using a fifth neural network to determine a representation discrimination loss between the reconstructed first output image and the reconstructed second output image; using a sixth neural network to determine a feature discrimination loss between the generated second alternative features and the obtained second features; and updating the weight coefficients of the fourth neural network, the fifth neural network, and the sixth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
[0135] Method 900 may further include: obtaining third features of an input image using a first neural network; and generating an attention map based on the obtained third features. Generating the second alternative features may include: generating the second alternative features using a fourth neural network based on the obtained second features and the generated attention map.
[0136] Method 900 may further include: obtaining third features from the decompressed representation using a first part of the third neural network; and generating an attention map based on the obtained third features. Generating the second alternative features may include: generating the second alternative features using a fourth neural network based on the obtained second features and the generated attention map.
[0137] Method 900 may further include: obtaining second features of the input image using a first part of the first neural network; and generating an attention map based on the obtained second features. Generating the first alternative features may include: using a second neural network to generate the first alternative features based on the obtained first features and the generated attention map.
[0138] Although Figure 9 example blocks of method 900 are shown, in some implementations, method 900 may include additional blocks, fewer blocks, different blocks, or blocks arranged differently compared to the Figure 9 blocks depicted therein. Additionally or alternatively, two or more of the blocks of method 900 may be executed in parallel.
[0139] Figure 10 is a block diagram of an apparatus 1000 for rate-adaptive neural image compression using an adversarial generator according to an embodiment.
[0140] As Figure 10 shown, apparatus 1000 includes a first obtaining code 1010, a first generating code 1020, a first encoding code 1030, a compressing code 1040, a decompressing code 1050, and a first decoding code 1060.
[0141] The first obtaining code 1010 is configured to cause at least one processor to obtain first features of the input image using a first part of the first neural network.
[0142] The first generating code 1020 is configured to cause at least one processor to generate first alternative features using a second neural network based on the obtained first features.
[0143] The first encoding code 1030 is configured to cause at least one processor to encode the generated first alternative features using a second part of the first neural network to generate a first encoded representation.
[0144] The compressing code 1040 is configured to cause at least one processor to compress the generated first encoded representation.
[0145] The decompressing code 1050 is configured to cause at least one processor to decompress the compressed representation.
[0146] The first decoding code 1060 is configured to cause at least one processor to decode the decompressed representation using a third neural network to reconstruct a first output image.
[0147] The second neural network can be trained as follows: determining a rate loss of the compressed representation; determining a distortion loss between the input image and the reconstructed first output image; encoding the obtained first features using a third part of the first neural network to generate a second encoded representation; using a fourth neural network to determine a representation discrimination loss between the generated first encoded representation and the generated second encoded representation; using a fifth neural network to determine a feature discrimination loss between the generated first alternative features and the obtained first features; and updating the weight coefficients of the second neural network, the fourth neural network, and the fifth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
[0148] The apparatus 1000 may further include: a second encoding code configured to cause at least one processor to encode an input image using the first neural network to generate a first encoded representation; a second obtaining code configured to cause at least one processor to obtain second features from the decompressed representation using a first part of the third neural network; a second generating code configured to cause at least one processor to generate second alternative features based on the obtained second features using the fourth neural network; and a second decoding code configured to cause at least one processor to decode the generated second alternative features using a second part of the third neural network to reconstruct a first output image.
[0149] The fourth neural network can be trained as follows: determining a rate loss of the compressed representation; determining a distortion loss between the input image and the reconstructed first output image; decoding the obtained second features using a third part of the third neural network to reconstruct a second output image; using a fifth neural network to determine a representation discrimination loss between the reconstructed first output image and the reconstructed second output image; using a sixth neural network to determine a feature discrimination loss between the generated second alternative features and the obtained second features; and updating the weight coefficients of the fourth neural network, the fifth neural network, and the sixth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
[0150] The apparatus 1000 may further include: a third obtaining code configured to cause at least one processor to obtain third features of the input image using the first neural network; and a third generating code configured to cause at least one processor to generate an attention map based on the obtained third features. The second generating code may further be configured to cause at least one processor to generate second alternative features based on the obtained second features and the generated attention map using the fourth neural network.
[0151] The apparatus 1000 may further include: a third obtaining code configured to cause at least one processor to obtain third features from the decompressed representation using a first part of a third neural network; and a third generating code configured to cause at least one processor to generate an attention map based on the obtained third features. The second generating code may further be configured to cause at least one processor to generate second alternative features based on the obtained second features and the generated attention map using a fourth neural network.
[0152] The apparatus 1000 may further include: a second obtaining code configured to cause at least one processor to obtain second features of an input image using a first part of a first neural network; and a second generating code configured to cause at least one processor to generate an attention map based on the obtained second features. The first generating code 1020 may further be configured to cause at least one processor to generate first alternative features based on the obtained first features and the generated attention map using a second neural network.
[0153] Compared with conventional end-to-end (E2E) image compression methods, the described embodiments have the following new features. A compact adversarial generator is adapted to the anchor rate-distortion (R-D) trade-off λ o value-trained NIC model instance to simulate the compression effect of the intermediate R-D trade-off λ t value. The common generator is offline-trained to adapt the NIC model instance in a data-independent manner such that no online learning or feedback is required for this adaptation.
[0154] Compared with conventional E2E image compression methods, the described embodiments have the following advantages: significantly reducing the deployed storage devices for implementing multi-compression rate compression; and a flexible framework adaptable to various types of NIC models. In addition, the attention-based generator can focus on the significant information for model adaptation.
[0155] These methods can be used alone or in any combination in any order. In addition, each of the methods (or embodiments), encoders, and decoders can be implemented by a processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0156] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure, or may be obtained from the practice of the implementations.
[0157] As used herein, the term "component" is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.
[0158] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual special control hardware or software code used to implement these systems and / or methods does not limit the implementation. Thus, the operations and behavior of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware can be designed based on the description herein to implement the systems and / or methods.
[0159] Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly refer to only one claim, the disclosure of possible implementations includes each dependent claim combined with every other claim in the claim group.
[0160] Unless explicitly described as such, any element, act, or instruction used herein should not be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and can be used interchangeably with "one or more." Additionally, as used herein, the term "group" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and can be used interchangeably with "one or more." The term "one" or similar language is used when referring to only one item. Additionally, as used herein, the terms "having," "has," "with," etc. are intended to be open-ended terms. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on."
Claims
1. A method for rate-adaptive neural image compression using an adversarial generator, characterized in that, The method is executed by at least one processor, and the method includes: Obtaining first features of an input image using a first part of a first neural network; Generating first alternative features based on the obtained first features using a second neural network, where the second neural network is an additional component adversarial generator that can be inserted into an existing NIC model, and is trained based on a rate loss R of a compressed representation and a distortion loss D between the input image and the reconstructed first output image, and is used to process the features to generate alternative features after obtaining the input image features, and by continuously adjusting its own parameters to optimize the overall R-D loss, enabling the generated alternative features to adapt to different compression rate requirements and achieving rate adaptability; Encoding the generated first alternative features using a second part of the first neural network to generate a first encoded representation; Compressing the generated first encoded representation; Decompressing the compressed representation; and Decoding the decompressed representation using a third neural network to reconstruct a first output image.
2. The method according to claim 1, wherein Training the second neural network by: Determining the rate loss of the compressed representation; Determining the distortion loss between the input image and the reconstructed first output image; Encoding the obtained first features using a third part of the first neural network to generate a second encoded representation; Using a fourth neural network to determine a representation discrimination loss between the generated first encoded representation and the generated second encoded representation; Using a fifth neural network to determine a feature discrimination loss between the generated first alternative features and the obtained first features; And Updating the weight coefficients of the second neural network, the fourth neural network, and the fifth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
3. The method according to claim 1, characterized in that, The method further includes: Encoding the input image using the first neural network to generate the first encoded representation; Obtaining second features from the decompressed representation using a first part of the third neural network; Generating second alternative features based on the obtained second features using a fourth neural network; and Decoding the generated second alternative features using a second part of the third neural network to reconstruct the first output image.
4. The method according to claim 3, characterized in that Training the fourth neural network by: Determining the rate loss of the compressed representation; Determining the distortion loss between the input image and the reconstructed first output image; Decoding the obtained second features using a third part of the third neural network to reconstruct a second output image; Using a fifth neural network to determine a representation discrimination loss between the reconstructed first output image and the reconstructed second output image; Using a sixth neural network to determine a feature discrimination loss between the generated second alternative features and the obtained second features; And Updating the weight coefficients of the fourth neural network, the fifth neural network, and the sixth neural network to optimize the determined rate loss, the determined distortion loss, the determined representation discrimination loss, and the determined feature discrimination loss.
5. The method according to claim 3, characterized in that The method further includes: obtaining a third feature of the input image using the first neural network; and generating an attention map based on the obtained third feature, wherein generating the second alternative feature includes: generating the second alternative feature using the fourth neural network based on the obtained second feature and the generated attention map.
6. The method according to claim 3, characterized in that, The method further includes: obtaining a third feature from the decompressed representation using a first part of the third neural network; and generating an attention map based on the obtained third feature, wherein generating the second alternative feature includes: generating the second alternative feature using the fourth neural network based on the obtained second feature and the generated attention map.
7. The method according to claim 1, characterized in that The method further includes: obtaining a second feature of the input image using a first part of the first neural network; and generating an attention map based on the obtained second feature, wherein generating the first alternative feature includes: generating the first alternative feature using the second neural network based on the obtained first feature and the generated attention map.
8. An apparatus for rate-adaptive neural image compression using an adversarial generator, characterized in that, The apparatus includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to what is indicated by the program code to execute the method according to any one of claims 1 to 7.
9. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to execute the method according to any one of claims 1 to 7.
10. An apparatus for rate-adaptive neural image compression using an adversarial generator, characterized in that, The apparatus includes: an obtaining module configured to obtain a first feature of an input image using a first part of a first neural network; a generating module configured to generate a first alternative feature using a second neural network based on the obtained first feature, the second neural network being an additional component adversarial generator insertable into an existing NIC model, and being trained based on a rate loss R of the compressed representation and a distortion loss D between the input image and the reconstructed first output image, for processing the feature to generate an alternative feature after obtaining the input image feature, and optimizing the overall R-D loss by continuously adjusting its own parameters so that the generated alternative feature can adapt to different compression rate requirements and achieve rate adaption; an encoding module configured to encode the generated first alternative feature using a second part of the first neural network to generate a first encoded representation; a compression module configured to compress the generated first encoding; a decompression module configured to decompress the compressed representation; and a decoding module configured to decode the decompressed representation using a third neural network to reconstruct a first output image.
Citation Information
Patent Citations
A method and technical equipment for video processing
WO2018150083A1