Multi-bitrate Neural Image Compression Method, Device, and Storage Medium

By using a single model example and encoding mask method in neural image compression, the problem of difficulty in multi-bit rate control in the prior art is solved, and flexible bit rate management and resource conservation effects are achieved.

CN114631119BActive Publication Date: 2025-06-03TENCENT AMERICA LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180006013.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-04-28
Filing Date
2021-05-03
Publication Date
2025-06-03
Estimated Expiration
2041-05-03

AI Technical Summary

Technical Problem

Existing end-to-end neural image compression methods based on deep neural networks have challenges in flexible bit rate control, requiring training multiple model instances to accommodate trade-offs of different bit rate, resulting in high consumption of storage and computing resources.

Method used

The multi-bit rate neural image compression method is adopted, and the encoding mask is selected based on the hyperparameters, and the mask weight is generated through convolutional operations to encode and decode the input image, thereby realizing image compression at multi-bit rate.

Benefits of technology

This method can realize multi-bit rate image compression without increasing the number of model instances, reducing the consumption of storage and computing resources, and improving compression efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114631119B_ABST
    Figure CN114631119B_ABST
Patent Text Reader

Abstract

A multi-bitrate neural image compression method includes selecting an encoding mask based on hyperparameters and performing a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights. The method further includes encoding an input image using the first masked weights to obtain an encoded representation and encoding the obtained encoded representation to obtain a compressed representation.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference to related applications

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 242,874, filed on April 28, 2021, which is based on and claims the priority of U.S. Provisional Patent Application No. 63 / 045,341, filed on June 29, 2020, and U.S. Provisional Patent Application No. 63 / 087,519, filed on October 5, 2020. The disclosures of these patent applications are incorporated herein by reference in their entirety. Background of the Invention

[0003] Standard groups and companies have been actively looking for potential requirements for future video coding technology standardization. These standard groups and companies have focused on end - to - end neural image compression (NIC) based on artificial intelligence (AI) using deep neural networks (DNNs). The success of this approach has brought increasing industrial benefits to advanced neural image and video compression methods.

[0004] For previous NIC methods, flexible bit - rate control remains a challenging problem. Traditionally, it may involve training multiple model instances, each of the multiple training instances being respectively for each desired trade - off between bit - rate and distortion (the quality of the compressed image). All these multiple model instances may need to be stored and deployed on the decoder side to reconstruct images from different bit - rates. This can be very expensive for many applications with limited storage and computational resources. Summary of the Invention

[0005] According to an embodiment, a method for multi - bit - rate neural image compression is performed by at least one processor, and the method includes selecting an encoding mask based on hyperparameters, and performing a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights. The method further includes encoding an input image using the first masked weights to obtain an encoded representation, and encoding the obtained encoded representation to obtain a compressed representation.

[0006] According to an embodiment, a multi-bitrate neural image compression device includes at least one memory and at least one processor. The at least one memory is configured to store program code, and the at least one processor is configured to read the program code and operate according to the instructions of the program code. The program code includes a first selection code and a first execution code. The first selection code is configured to cause the at least one processor to select an encoding mask based on hyperparameters. The first execution code is configured to cause the at least one processor to perform a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights. The program code further includes a first encoding code and a second encoding code. The first encoding code is configured to cause the at least one processor to encode an input image using the first masked weights to obtain an encoded representation. The second encoding code is configured to cause the at least one processor to encode the obtained encoded representation to obtain a compressed representation.

[0007] According to an embodiment, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor for multi-bitrate neural image compression, cause the at least one processor to select an encoding mask based on hyperparameters and perform a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights. When executed by the at least one processor, the instructions further cause the at least one processor to encode an input image using the first masked weights to obtain an encoded representation and encode the obtained encoded representation to obtain a compressed representation. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a diagram of an environment according to an embodiment in which the methods, devices, and systems described herein can be implemented.

[0009] Figure 2 is Figure 1 a block diagram of example components of one or more devices.

[0010] Figure 3 is a block diagram of a test device for multi-bitrate neural image compression in a test phase according to an embodiment.

[0011] Figure 4A is a block diagram of a training device for multi-bitrate neural image compression in a training phase according to an embodiment.

[0012] Figure 4B is a block diagram of a training device for multi-bitrate neural image compression in a training phase according to an embodiment.

[0013] Figure 4CBlock diagram of a training device for multi-bitrate neural image compression in a training phase according to an embodiment.

[0014] Figure 5 Flowchart of a multi-bitrate neural image compression method according to an embodiment.

[0015] Figure 6 Block diagram of a multi-bitrate neural image compression device according to an embodiment.

[0016] Figure 7 Flowchart of a multi-bitrate neural image decompression method according to an embodiment.

[0017] Figure 8 Block diagram of a multi-bitrate neural image decompression device according to an embodiment. Detailed implementation

[0018] The present invention describes a method and device for compressing an input image using a multi-bitrate neural image compression (NIC) framework, in which only one NIC model instance is used to achieve image compression at multiple bitrates under the guidance of multiple binary masks for different bitrates.

[0019] Figure 1 Diagram of an environment 100 according to an embodiment, in which the methods, devices, and systems described herein can be implemented.

[0020] As Figure 1 shown, the environment 100 may include a user device 110, a platform 120, and a network 130. The devices in the environment 100 may be interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0021] The user device 110 includes one or more devices that are capable of receiving, generating, storing, processing, and / or providing information associated with the platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or a similar device. In some embodiments, the user device 110 may receive information from the platform 120 and / or send information to the platform 120.

[0022] Platform 120 includes one or more devices as described elsewhere in this disclosure. In some embodiments, platform 120 may include a cloud server or a group of cloud servers. In some embodiments, platform 120 may be designed to be modular such that software components can be swapped in or out. Thus, platform 120 can be reconfigured for different uses easily and / or quickly.

[0023] In some embodiments, as shown, platform 120 may be hosted in a cloud computing environment 122. It should be noted that while the embodiments described herein describe platform 120 as being hosted in cloud computing environment 122, in some embodiments, platform 120 may not be cloud-based (i.e., may be implemented outside of a cloud computing environment) or may be partially cloud-based.

[0024] Cloud computing environment 122 includes an environment that hosts platform 120. Cloud computing environment 122 may provide end users (e.g., user device 110) with services such as computing, software, data access, storage, etc. without the need for knowledge of the physical location and configuration of the (one or more) systems and / or (one or more) devices of hosted platform 120. As shown, cloud computing environment 122 may include a group of computing resources 124 (collectively referred to as "the plurality of computing resources 124" and individually as "computing resource 124").

[0025] Computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some embodiments, computing resources 124 may host platform 120. Cloud resources may include computing instances executed in computing resources 124, storage devices provided in computing resources 124, data transfer devices provided by computing resources 124, etc. In some embodiments, computing resources 124 may communicate with other computing resources 124 via a wired connection, a wireless connection, or a combination of wired and wireless connections.

[0026] As Figure 1 Further shown, computing resources 124 include a group of cloud resources such as one or more applications ("APPs") 124-1, one or more virtual machines ("VMs") 124-2, virtualized storage ("VSs") 124-3, or one or more hypervisors ("HYPs") 124-4, etc.

[0027] Application 124-1 includes one or more software applications that can be provided to and / or accessed by user device 110 and / or platform 120. Application 124-1 can eliminate the need to install and execute software applications on user device 110. For example, Application 124-1 can include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some embodiments, one Application 124-1 can send information to and / or receive information from one or more other Applications 124-1 via virtual machine 124-2.

[0028] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs similar to a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the use of virtual machine 124-2 and its correspondence to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (“OS”). A process virtual machine can execute a single program and can support a single process. In some embodiments, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of cloud computing environment 122, such as data management, synchronization, or long-term data transfer.

[0029] Virtualized memory 124-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of computing resources 124. In some embodiments, in the context of a storage system, the types of virtualization can include block-level virtualization and file virtualization. Block-level virtualization can refer to abstracting (or separating) logical storage from physical storage so that the storage system can be accessed without regard to the physical storage or heterogeneous structure. The separation can allow flexibility for the administrator of the storage system in terms of how the administrator manages the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location of the physical storage file. This can enable optimization of storage usage, server consolidation, and / or the performance of uninterrupted file migration.

[0030] Hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., “guest operating systems”) to execute simultaneously on a host computer (e.g., computing resources 124). Hypervisor 124-4 can present a virtual operating platform to the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share the virtualized hardware resources.

[0031] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic based network, etc., and / or a combination of these networks or other types of networks.

[0032] Provide Figure 1 The number and arrangement of the devices and networks shown are for example purposes. In practice, compared to Figure 1 that shown, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with a different arrangement. Additionally, Figure 1 two or more of the devices shown in Figure 1 may be implemented within a single device, or

[0033] Figure 2 is Figure 1 a block diagram of example components of one or more of the devices of

[0034] Device 200 may correspond to user device 110 and / or platform 120. As Figure 2 shown, device 200 may include a bus 210, a processor 220, a memory 230, memory components 240, an input component 250, an output component 260, and a communication interface 270.

[0035] Bus 210 includes components that permit communication among the components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other type of processing component. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 220.

[0036] The memory component 240 stores information and / or software related to the operation and use of the device 200. For example, the memory component 240 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or other types of non-transitory computer-readable media, as well as corresponding drives.

[0037] The input component 250 includes components that allow the device 200 to receive information, for example, via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone). Additionally, or alternatively, the input component 250 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). The output component 260 includes components that provide output information from the device 200 (e.g., a display, a speaker, and / or one or more light-emitting diodes (LEDs)).

[0038] The communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receivers and transmitters) that enable the device 200 to communicate with other devices, for example, via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection. The communication interface 270 may allow the device 200 to receive information from another device and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0039] The device 200 may perform one or more of the processes described in this disclosure. In response to the processor 220 executing software instructions stored by a non-transitory computer-readable medium (e.g., the memory 230 and / or the memory component 240), the device 200 may perform these processes. A computer-readable medium is defined herein as a non-transient memory device. A memory device includes a memory space within a single physical storage device or a memory space distributed across multiple physical storage devices.

[0040] The software instructions may be read into the memory 230 and / or the memory component 240 from another computer-readable medium or from another device via the communication interface 270. When the software instructions stored in the memory 230 and / or the memory component 240 are executed, the software instructions may cause the processor 220 to perform one or more of the processes described in this disclosure. Additionally, or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more of the processes described in this disclosure. Accordingly, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.

[0041] ProvidedFigure 2 The number and arrangement of the components shown are by way of example. In practice, compared with Figure 2 those shown, the device 200 may include additional components, fewer components, different components, or components in a different arrangement. Additionally, or alternatively, a set of components (e.g., one or more components) of the device 200 may perform one or more functions described as being performed by another set of components of the device 200.

[0042] A multi-bitrate neural image compression method and device will now be described in detail.

[0043] The present disclosure presents a multi-bitrate NIC framework for learning and deploying only one NIC model instance, which supports multi-bitrate image compression. A set of binary masks is learned for each target bitrate to guide the decoder in the reconstruction phase to recover the image from different bitrates.

[0044] Figure 3 is a block diagram of a test device 300 for multi-bitrate neural image compression in a test phase according to an embodiment.

[0045] Referring to Figure 3 , the test device 300 includes a test deep neural network (DNN) encoder 310, a test encoder 320, a test decoder 330, and a test DNN decoder 340.

[0046] Given an input image x of size (h, w, c), where h, w, and c are the height, width, and number of channels respectively, the objectives of the test phase of the NIC workflow can be described as follows.

[0047] The test DNN encoder 310 encodes the input image x using a DNN to obtain an encoded representation y.

[0048] The test encoder 320 encodes the obtained encoded representation y to obtain a compressed representation that is compact for storage and transmission. The obtained encoded representation y can be encoded through quantization and entropy coding.

[0049] The test decoder 330 decodes the obtained compressed representation to obtain a recovered representation The obtained compressed representation can be decoded through decoding and inverse quantization.

[0050] The test DNN decoder 340 decodes the obtained recovered representation to reconstruct a reconstructed image The reconstructed image should be similar to the initial input image x.

[0051] There are no restrictions on the network structures of the test DNN encoder 310 and the test DNN decoder 340. Moreover, there are no restrictions on the methods (quantization and entropy coding) used for the test encoder 320 and the test decoder 330.

[0052] To learn the NIC model, it may be necessary to balance two competing expectations: better reconstruction quality and less bit consumption. The loss function is used to measure the reconstruction error, which is called the distortion loss, such as the peak signal-to-noise ratio (PSNR) and / or the structural similarity index measure (SSIM). The bitrate loss is calculated to measure the bit consumption of the compressed representation. Therefore, the trade-off hyperparameter λ is used to optimize the joint bitrate-distortion (R-D) loss:

[0053]

[0054] Training with a large hyperparameter λ results in a compressed model with less distortion but more bit consumption, and vice versa. Traditionally, for each value of the predefined hyperparameter λ, an NIC model instance would be trained, which would not work well for other values of the predefined hyperparameter λ. Therefore, to achieve multiple bitrates for the compressed stream, traditional methods may need to train and store multiple model instances.

[0055] In an embodiment, the multi-bitrate neural image compression method and apparatus use a single trained model instance of the NIC network, and use a set of binary masks to guide the NIC model instance to generate different compressed representations and corresponding reconstructed images, each mask for a different value of the hyperparameter λ.

[0056] Specifically, and respectively represent the sets of weight coefficients of the encoder and decoder parts of the NIC model instance, where, and are the weight coefficients of the j-th layer of the test DNN encoder 310 and the test DNN decoder 340, respectively. λ 1 , …, λ N represent N hyperparameters, and and represent the compressed representation and the reconstructed image corresponding to the hyperparameter λ i . and respectively represent the compressed representation and the reconstructed image corresponding to the hyperparameter λ iThe binary mask of the j-th layer of the test DNN encoder 310 and the test DNN decoder 340. The weights is a five-dimensional (5D) tensor with dimensions (c 1 , k 1 , k 2 , k 3 , c 2 ). The input to the layer is a four-dimensional (4D) tensor A with dimensions (h 1 , w 1 , d 1 , c 1 ), and the output of the layer is a four-dimensional tensor B with dimensions (h 2 , w 2 , d 2 , c 2 ). The dimensions c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , h 2 , w 2 , and d 2 are integers greater than or equal to 1. When any of the dimensions c 1 , k 1 , k 2 , k 3 , c 2 , h 1 , w 1 , d 1 , h 2 , w 2 , and d 2 is the number 1, the corresponding tensor is reduced to a lower dimension. Each term in each tensor is a floating-point number. The parameters h 1 , w 1 , and d 1 (h 2 , w 2 , and d 2 ) are the height, width, and depth of the input tensor A (output tensor B). The parameter c 1 (c 2 ) is the number of input (output) channels. The parameters k 1 , k 2 , and k 3 are the dimensions of the convolutional kernel corresponding to the height, width, and depth axes respectively. The output tensor B is calculated by a convolution operation ⊙ based on the input tensor A, the mask , and the weights . That is, the output tensor B is calculated as the convolution with the mask weights The input tensor A of the convolution, where · is element-wise multiplication. Similarly, for the weights Its output tensor B is calculated by the convolution operation of the input tensor A with the masked weights .

[0057] Referring to Figure 3 , the test DNN encoder 310 only includes one model instance with weights , and the test DNN decoder 340 only includes one model instance with weights . Given the input image x and the target hyperparameter λ i , the test DNN encoder 310 selects a set of encoding masks to calculate the masked weights which are used by the test DNN encoder 310 to calculate the DNN encoded representation y. Then, the test encoder 320 calculates the compressed representation during the encoding process based on the compressed representation The test decoder 330 calculates the recovered representation through the decoding process using the hyperparameter λ i . The test DNN decoder 340 selects a set of decoding masks to calculate the masked weights which are used by the test DNN decoder 340 to calculate the reconstructed image based on the recovered representation

[0058] The weights or (the masks or similarly) can be reshaped to correspond to the convolution of a reshaped input with reshaped weights or to obtain the same output. Specifically, there can be two configurations. First, the 5D weight tensor can be reshaped into a 3D tensor with dimensions (c′ 1 , c′ 2 , k), where c′ 1 ×c′ 2 ×k = c 1 ×c 2 ×k 1 ×k 2 ×k 3 . For example, the configuration can be c′ 1 = c 1 , c′ 2 = c 2 , k = k 1 ×k2 ×k 3 Second, the 5D weight tensor can be reshaped into a 2D matrix with dimensions (c′ 1 , c′ 2 ), where c′ 1 ×c′ 2 = c 1 ×c 2 ×k 1 ×k 2 ×k 3 . For example, the configuration can be c′ 1 = c 1 , c′ 2 = c 2 ×k 1 ×k 2 ×k 3 or c′ 2 = c 2 , c′ 1 = c 1 ×k 1 ×k 2 ×k 3 .

[0059] The desired microstructure of the mask can be designed to match how the matrix multiplication process of the underlying general matrix multiplication (GEMM) for implementing the convolution operation is performed, so that the inference calculation using the masked weight coefficients can be accelerated. For example, a block microstructure can be used for the mask of each layer in the 3D reshaped weight tensor or the 2D reshaped weight matrix (similarly for the masked weight coefficients). For the case of the 3D reshaped weight tensor, the mask can be divided into blocks with dimensions (g i , g o , g k ), and for the case of the 2D reshaped weight matrix, the mask can be divided into blocks with dimensions (g i , g o ). All the terms in the block of the mask will have the same binary value, either 1 or 0. That is, the weight coefficients are masked in a block microstructure manner.

[0060] The goal is to learn a set of microstructure encoding masks and microstructure decoding masks for each mask in the mask and for each hyperparameter of the hyperparameter λ i . A progressive multi-stage training framework can achieve this goal.

[0061] Specifically, assume hyperparameters λ 1 , …, λ iSorted in ascending order and corresponding to masks that generate compressed representations with increasing distortion (decreasing quality) and decreasing bitrate loss (increasing bitrate). Two different training frameworks can be used to learn the model instance and the masks, namely, As Figure 4A shown.

[0062] The overall workflow of the first training framework is as Figure 4A shown.

[0063] Figure 4A is a block diagram of a training device 400A for multi-bitrate neural image compression in the training phase according to an embodiment.

[0064] Referring to Figure 4A , the training device 400A includes a weight update component 410, a pruning component 420, and a weight update component 430.

[0065] Assume that the current goal is to train a mask for the hyperparameter λ i-1 The current model instance has weights And the mask is represented as The aim is to obtain the mask And the updated weights

[0066] In the first step, the weight coefficients in the weights Masked by respectively are fixed or set. For example, if the entry in the mask is 1, the corresponding weight is fixed.

[0067] Then, the weight update component 410 uses the R-D loss of equation (1) for the first hyperparameter λ 1 Through backpropagation, the remaining unmasked weight coefficients in the weights And Are updated to the updated weights And Multiple rounds of iteration can be performed to optimize the R-D loss in this weight update process. For example, until the maximum number of iterations is reached or until the loss converges.

[0068] Thereafter, a microstructural weight pruning process is performed. In this process, using the updated weights And As inputs, for the updated weights And Among the unfixed weight coefficients, the pruning component 420 obtains or calculates the pruning loss L of each microstructure block b (a 3D block for the 3D shaping weight tensor or a 2D block for the 2D shaping weight matrix). s (b) (e.g., the L of the weights in the block 1 or L 2 norm). The pruning component 420 arranges these microstructure blocks in ascending order and prunes these blocks from the top down in the sorted list (i.e., by setting the corresponding weights in the pruned blocks to 0) until a termination criterion is reached.

[0069] For example, assume a validation dataset S val , with updated weights and a mask The NIC model produces a distortion loss As more and more microblocks are pruned, this distortion loss will gradually increase. The termination criterion can be a tolerance percentage threshold that allows the distortion loss to increase.

[0070] The pruning component 420 generates a set of binary pruning masks and where the entries in the mask or being 0 indicates that the corresponding weight or is pruned.

[0071] Then, the weight update component 430 fixes the updated weights and in the additional unfixed weights masked by the masks and and updates the remaining weights in the updated weights and not masked by any of the masks or to optimize the overall R-D loss of Equation (1) for the hyperparameter λ i-1 . Multiple rounds of iteration can be performed to optimize the R-D loss in this weight update process, e.g., until the maximum number of iterations is reached or until the loss converges. Then, the weight update component 430 obtains or calculates the corresponding masks and as follows:[[]]END]] and That is, the unpruned entries in the mask will be additionally set to 1 when masked in , and this unpruned entry is not masked in the mask . In addition, the weight update component 430 outputs the updated weights and Final updated weights and are the final output weights of the learned model instance and

[0072] The overall workflow of the second training framework is as Figure 4B shown

[0073] Figure 4B is a block diagram of a training device 400B for multi-bitrate neural image compression in a training phase according to an embodiment

[0074] Referring to Figure 4B , the training device 400B includes a weight update component 440, a pruning component 450, a weight update component 460, and an inverse pruning weight update component 470

[0075] Assuming an initial set of weights and (e.g., randomly initialized according to some distribution), the weight update component 440 learns the model weight sets tr by using the training data set S 1 and a weight update process using conventional backpropagation, by optimizing the R-D loss of Equation (1) for the hyperparameter λ and

[0076] Thereafter, the pruning component 450 performs a micro-structure pruning process based on the model weights . In this micro-structure pruning process, the pruning component 450 divides each shaped 3D weight tensor or 2D weight matrix into micro-blocks (3D blocks for three-dimensional shaped weight tensors or 2D blocks for two-dimensional shaped two-dimensional weight matrices), and obtains or calculates the pruning loss L s (b) (e.g., the L 1 or L 2 norm of the weights in the block)

[0077] In the following manner, the pruning component 450 sorts these micro-structure blocks in ascending order and prunes these blocks from top to bottom on the sorted list (i.e., by setting the corresponding weights in the pruned blocks to 0), for each hyperparameter in the hyperparameters λ 1 ,..., λ N . Assuming the current weight is , the pruning component 450 obtains the corresponding binary pruning masks and where the entries in the mask or being 0 indicates the weight or The corresponding weights in are pruned. The pruning component 450 further obtains λ i+1 's pruning mask and to obtain the updated weights To achieve this goal, during the pruning process, the pruning component 450 fixes the weights and in the weight coefficient, which is masked to be masked by the mask and pruned, and continue to prune the remaining unpruned microblocks downward in the sorted list until reaching the hyperparameter λ i+1 's termination criterion. For example, assume the validation dataset S val , with the NIC model having weights producing a distortion loss As more and more microblocks are pruned, this distortion loss will gradually increase. The termination criterion can be a tolerance percentage threshold for allowing the distortion loss to increase. Then, the pruning component 450 generates the pruning mask by adding these additionally pruned microblocks to the mask and to generate the pruning mask and

[0078] Then, during the weight update process, the weight update component 460 fixes all these pruned microblocks masked by the mask and and uses conventional backpropagation to optimize the R-D loss of Equation (1) for the hyperparameter λ i+1 to update the remaining unfixed weights to generate the updated weight set By repeating the above pruning and weight update process for each hyperparameter λ 1 , …, λ N the pruning component 450 obtains a set of pruning masks and the weight update component 460 updates the finally updated weight pruning mask Pruning mask and are directly used as the model masks for the hyperparameter λ i and

[0079] Thereafter, in the following manner, the inverse pruning weight update component 470 trains the weights and the model mask through the inverse pruning weight update process and Assume the current weights are obtained For the weights The weight coefficients in the mask and that are masked to 1 are fixed. The weight coefficients that are masked to 1 in the mask and and are masked to 0 in the mask and are filled. These weights can be filled when they are pruned during the pruning process, or can be filled with randomly initialized values. Then, the inverse pruning weight update component 470 updates these newly filled weights by optimizing the R-D loss of Equation (1) for the hyperparameter λ i-1 using conventional backpropagation. This results in updated weights Repeat this process until the final weights Obtain as the final output.

[0080] In an embodiment, a prune-and-grow (PnG) training framework can be used to learn binary masks. Figure 4C The overall workflow of the PnG training framework is given.

[0081] Figure 4C is a block diagram of a training device 400C for multi-bitrate neural image compression in the training phase according to an embodiment.

[0082] Referring to Figure 4C , the training device 400C includes a weight update component 480, a pruning component 485, and a weight update component 490.

[0083] The goal is to learn a set of sparse coding masks and sparse decoding masks where each mask in the masks and is for each hyperparameter λi. The PnG training framework is a progressive multi-stage training framework for achieving this goal.

[0084] Specifically, assume that the hyperparameters λ 1 , …, λ N are arranged in descending order and correspond to masks that generate compressed representations with increasing distortion (decreasing quality) and decreasing bitrate loss (increasing bitrate). Assume that the current goal is to train a mask for the hyperparameter λ i+1 , the current model instance has weights and the mask is represented as The aim is to obtain the mask as well as the updated weights

[0085] In the first step, the weights The weight coefficients in are respectively masked by a mask. For example, if an entry in the mask is 1, the corresponding weight

[0086] will then be fixed. i+1 Then, the weight update component 480 uses the R-D loss of Equation (1) for the first hyperparameter λ and to update the remaining unmasked weight coefficients in and to the updated weights

[0087] through backpropagation. Multiple rounds of iteration will be adopted to optimize the R-D loss in this weight update process. For example, until the maximum number of iterations is reached or until the loss converges Thereafter, the pruning component 485 performs a weight pruning process. Any DNN weight pruning method can be used here, such as the unstructured weight sparsification method [1] or the structured weight pruning method [2]. A sparse regularization loss

[0088]

[0089] The hyperparameter η ≥ 0 balances the importance of the sparse regularization loss, which is usually predefined. The sparse regularization loss aims to promote and a plurality of zero-valued weight coefficients in. For example, each layer can be processed separately:

[0090]

[0091] Each is a sparse loss defined on the weight tensor . For example, a weight tensor of size (c 1 , k 1 , k 2 , k 3 , c 2 ) can be flattened into a vector of size c 1 × k 1 × k 2 × k 3 × c 2 , and the L 0 of the flattened vector can be calculated, L1 , L 2 or L 2,1 norm as the sparsity loss.

[0092] The weight pruning process includes two modules. First, in the pruning module, the updated weights and are used as inputs. For the unfixed weight coefficients in the updated weights and , the pruning component 485 first selects the unimportant weight coefficients (i.e., if pruned, the loss is smaller). Then, the pruning component 485 fixes the previously fixed weights through the mask , and the weight update component 490 updates the remaining weights in the updated weights and by conventional backpropagation to optimize the overall R-D loss of equation (2) for the first hyperparameter λ i+1 . Multiple rounds of iteration will be adopted to optimize the total loss. For example, until the maximum number of iterations is reached or until the loss converges. The pruning component 485 finally outputs a set of binary pruning masks and where an entry in the mask or being zero indicates that the corresponding weight in or is set to zero (pruned). The pruning component 420 generates a set of binary pruning masks and where an entry in the mask or being 0 indicates that the corresponding weight or is pruned.

[0093] Then, the weight update component 490 fixes the additional unfixed weights in the updated weights and , which are masked by the masks and , and updates the remaining weights in the updated weights and by conventional backpropagation to optimize the overall R-D loss of equation (1) for the first hyperparameter λ i+1 , where the remaining weights are not masked by either of the masks or . Multiple rounds of iteration will be adopted to optimize the R-D loss in this weight update process. For example, until the maximum number of iterations is reached or until the loss converges. Then, the weight update component 490 obtains or calculates the corresponding mask and } , according to: and That is, the mask in the mask The unpruned entries that are not masked will be additionally set to 1 when masked. In addition, the weight update component 490 outputs the updated weights and and The final updated weights and are the final output weights of the learned model instance and

[0094] Different patterns of binary masks can be implemented. For example, binary masks can be structurally or non-structurally sparse. That is, the zero entries can be randomly distributed or form some special patterns in the weight tensor. It may be required that all layers in the DNN model have the same sparsity pattern. Each layer of the DNN model can also adopt different sparsity patterns. Three embodiments of the sparsity pattern are given below.

[0095] Non-structural mask

[0096] The binary mask can have randomly distributed zero entries. This is called a non-structural mask. In this case, the unimportant weights are the weights with very small values. For example, the p% weight coefficients with the minimum values in the weight tensor can be selected for pruning.

[0097] Structural mask

[0098] For a 5D weight tensor of size (c 1 , k 1 , k 2 , k 3 , c 2 ), if the weight tensor is reshaped into 3D cubic blocks with the shape (c 1 , c 2 , k 1 × k 2 × k 3 ), then entire columns (along the c 1 axis), rows (along the c 2 axis), and channels (along the k 1 × k 2 × k 3 axis) can be set to zero. For example, the loss (e.g., L 1 or L 2 norm) of each column, row, or channel can be calculated, and the posterior p% of the columns, rows, or channels with the minimum loss can be selected for pruning.

[0099] Micro-structural mask

[0100] The 5D weight tensor of size (c 1 ,k 1 ,k 2 ,k 3 ,c 2 ) can be reshaped into a 4D tensor, a 3D cube, a 2D matrix, or even a 1D vector. Small microstructural weights can be set to zero, e.g., small 4D, 3D, 2D, or 1D blocks, rather than setting entire rows, columns, or channels along any reshaping axis to zero. For example, the loss for each microstructure (e.g., L 1 or L 2 norm) can be calculated, and the posterior p% of the microstructures to be pruned can be selected.

[0101] Comparing the above three embodiments, the unstructured mask can have the least constraint on the weight coefficients and can better maintain the compression performance. However, due to the randomly distributed zero entries, this embodiment may not accelerate the inference calculation. The structured mask can naturally reduce the calculation but has a strong constraint on the weight coefficients. Therefore, the structured mask causes more harm to the compression performance. The microstructural mask is a trade-off between the unstructured mask and the structured mask, and maintaining the balance between compression performance and inference acceleration depends on the specific design of the microstructure and the corresponding hardware computing device.

[0102] Figure 5 is a flowchart of a multi-bitrate neural image compression method 500 according to an embodiment.

[0103] In some embodiments, one or more processing blocks of Figure 5 can be executed by the platform 120. In some embodiments, one or more processing blocks of Figure 5 can be executed by another device or a group of devices (such as the user device 110) that is separate from or includes the platform 120.

[0104] In operation 510, the method 500 includes selecting an encoding mask based on hyperparameters.

[0105] In operation 520, the method 500 includes performing a convolution of the first plurality of weights of the first neural network with the selected encoding mask to obtain first masked weights.

[0106] In operation 530, the method 500 includes encoding the input image using the first masked weights to obtain an encoded representation.

[0107] In operation 540, the method 500 includes encoding the obtained encoded representation to obtain a compressed representation.

[0108] Although Figure 5illustrates exemplary blocks of method 500, but in some embodiments, method 500 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to the blocks depicted in Figure 5 . Additionally, or alternatively, two or more blocks of method 500 may be executed in parallel.

[0109] Figure 6 is a block diagram of a multi-bitrate neural image compression device 600 according to an embodiment.

[0110] As Figure 6 shown, device 600 includes a first selection code 610, a first execution code 620, a first encoding code 630, and a second encoding code 640.

[0111] The first selection code 610 is configured to cause at least one processor to select an encoding mask based on hyperparameters.

[0112] The first execution code 620 is configured to cause at least one processor to perform a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights.

[0113] The first encoding code 630 is configured to cause at least one processor to encode an input image using the first masked weights to obtain an encoded representation.

[0114] The second encoding code 640 is configured to cause the at least one processor to encode the obtained encoded representation to obtain a compressed representation.

[0115] Figure 7 is a flowchart of a multi-bitrate neural image decompression method 700 according to an embodiment.

[0116] In some embodiments, one or more processing blocks of Figure 7 may be executed by platform 120. In some embodiments, one or more processing blocks of Figure 7 may be executed by another device or a group of devices (such as user device 110) that is separate from or includes platform 120.

[0117] As Figure 7 shown, in operation 710, method 700 includes decoding the obtained compressed representation to obtain a recovered representation.

[0118] In operation 720, method 700 includes selecting a decoding mask based on hyperparameters.

[0119] In operation 730, method 700 includes performing a convolution of a second plurality of weights of a second neural network with the selected decoding mask to obtain second masked weights.

[0120] In operation 740, method 700 includes decoding the obtained restored representation using a second mask weight to reconstruct an output image.

[0121] Each of the encoding mask and the decoding mask can be divided into blocks, and each item in each of the respective blocks of the blocks can have the same binary value.

[0122] The first neural network and the second neural network are trained by: updating one or more weights among the first plurality of weights and the second plurality of weights that are not masked by the encoding mask and the decoding mask respectively to minimize a bitrate-distortion loss that is determined based on an input image, an output image, and a compressed representation; pruning one or more weights among the first plurality of weights and the second plurality of weights that are updated and not masked by the encoding mask and the decoding mask respectively to obtain a binary pruning mask indicating which of the first plurality of weights and the second plurality of weights to prune; updating at least one weight among the first plurality of weights and the second plurality of weights that is not masked by the encoding mask, the decoding mask, and the obtained binary pruning mask to minimize the bitrate-distortion loss; and updating the encoding mask and the decoding mask based on the obtained binary pruning mask.

[0123] Pruning can include: determining a pruning loss for each block of the blocks, where the blocks are divided by each of the encoding mask and the decoding mask; sorting the blocks in ascending order based on the determined pruning loss for each block of the blocks; and setting two or more weights among the first plurality of weights and the second plurality of weights until a termination criterion is reached, where the two or more weights correspond to a plurality of blocks from top to bottom in the sorted blocks.

[0124] Each of the encoding mask and the decoding mask can have randomly distributed binary values.

[0125] Each of the encoding mask and the decoding mask can be divided into columns, rows, or channels, and each item in each of the respective columns, rows, or channels of the columns, rows, or channels can have the same binary value.

[0126] Although Figure 7 illustrates exemplary blocks of method 700, in some embodiments, method 700 can include additional blocks, fewer blocks, different blocks, or differently arranged blocks compared to the blocks depicted in Figure 7 . Additionally, or alternatively, two or more blocks of method 700 can be executed in parallel.

[0127] Figure 8 is a block diagram of a multi-bitrate neural image decompression device 800 according to an embodiment.

[0128] As Figure 8As shown, device 800 includes a first decoding code 810, a second selection code 820, a second execution code 830, and a second decoding code 840.

[0129] The first decoding code 810 is configured to cause at least one processor to decode the obtained compressed representation to obtain a recovered representation.

[0130] The second selection code 820 is configured to cause at least one processor to select a decoding mask based on hyperparameters.

[0131] The second execution code 830 is configured to cause at least one processor to perform a convolution of the second plurality of weights of the second neural network with the selected decoding mask to obtain second masked weights.

[0132] The second decoding code 840 is configured to cause at least one processor to decode the obtained recovered representation using the second masked weights to reconstruct an output image.

[0133] Each of the encoding mask and the decoding mask can be divided into blocks, and each item in each of the respective blocks of the blocks can have the same binary value.

[0134] The first neural network and the second neural network are trained by: updating one or more weights in the first plurality of weights and the second plurality of weights that are not masked by the encoding mask and the decoding mask respectively to minimize a bitrate-distortion loss determined based on an input image, an output image, and a compressed representation; pruning one or more weights in the first plurality of weights and the second plurality of weights that are updated and not masked by the encoding mask and the decoding mask respectively to obtain a binary pruning mask indicating which of the first plurality of weights and the second plurality of weights to prune; updating at least one weight in the first plurality of weights and the second plurality of weights that is not masked by the encoding mask, the decoding mask, and the obtained binary pruning mask respectively to minimize the bitrate-distortion loss; and updating the encoding mask and the decoding mask based on the obtained binary pruning mask.

[0135] Pruning can include: determining a pruning loss for each block of the blocks, where the blocks are divided by each of the encoding mask and the decoding mask; arranging the blocks in ascending order based on the determined pruning loss for each block of the blocks; and setting two or more weights in the first plurality of weights and the second plurality of weights until a termination criterion is reached, where the two or more weights correspond to a plurality of blocks from top to bottom in the arranged blocks.

[0136] Each of the encoding mask and the decoding mask can have randomly distributed binary values.

[0137] Each mask in the encoding mask and the decoding mask can be divided into columns, rows, or channels, and each item in each column, row, or channel of the columns, rows, or channels can have the same binary value.

[0138] Compared with previous end-to-end (E2E) image compression methods, the embodiments described herein use only one model instance to achieve multi-bitrate compression with multiple binary masks. Two training frameworks can be used to learn the model instance and the masks, and the model instance and the masks can have a block microstructure. In addition, a prune-and-grow training framework can be used to learn the model instance and general and flexible binary masks.

[0139] Compared with previous E2E image compression methods, the embodiments described herein can greatly reduce the deployment storage to achieve multi-bitrate compression and use a flexible and general framework adaptable to various types of NIC models. The structural mask and the microstructure mask provide an additional benefit of reduced computation.

[0140] The proposed methods can be used alone or in combination in any order. In addition, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0141] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the implementation to the exact forms disclosed. Modifications and variations can be made in light of the above disclosure, or can be obtained from practice of the embodiments.

[0142] As used in this disclosure, the term component is intended to be broadly interpreted as hardware, firmware, or a combination of hardware and software.

[0143] Obviously, the systems and / or methods described herein can be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code for implementing these systems and / or methods does not limit the implementation. Therefore, the operations and behaviors of the systems and / or methods are described herein without reference to specific software code—it should be understood that the software and hardware can be designed to implement the systems and / or methods based on the present disclosure.

[0144] Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification. While each dependent claim listed below may depend directly on only one claim, the disclosure that is possible includes the combination of each dependent claim with every other claim in the claim set.

[0145] Unless explicitly described, no element, act, or instruction used in this disclosure should be construed as critical or essential. Additionally, as used in this disclosure, the term "a" is intended to include one or more items and may be used interchangeably with "one or more." Further, as used in this disclosure, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." The term "a" or similar language is used where only one item is meant. Additionally, as used in this disclosure, the term "has" or similar terms are open-ended terms. Further, the term "based on" means "at least partially based on" unless otherwise explicitly stated.

Claims

1. A multi-bitrate neural image processing method, the method being executed by at least one processor, and the method comprises: selecting an encoding mask based on hyperparameters; performing a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights; encoding an input image using the first masked weights to obtain an encoded representation; and encoding the obtained encoded representation to obtain a compressed representation; wherein the first neural network is trained by: updating one or more weights of the first plurality of weights not masked by the encoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, an output image, and the compressed representation; pruning one or more weights of the updated first plurality of weights not masked by the encoding mask to obtain a binary pruning mask indicating which of the first plurality of weights to prune; the pruning includes: determining a pruning loss for each block into which each mask in the encoding mask is divided; sorting the plurality of blocks in ascending order based on the determined pruning loss for each block of the plurality of blocks; and setting two or more weights of the first plurality of weights corresponding to the plurality of blocks from top to bottom in the sorted plurality of blocks until a termination criterion is reached; updating at least one weight of the first plurality of weights not masked by the encoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and updating the encoding mask based on the obtained binary pruning mask.

2. The method according to claim 1, further comprises: decoding the obtained compressed representation to obtain a recovered representation; selecting a decoding mask based on the hyperparameters; performing a convolution of a second plurality of weights of a second neural network with the selected decoding mask to obtain second masked weights; and decoding the obtained recovered representation using the second masked weights to reconstruct an output image; wherein the second neural network is trained by: updating one or more weights of the second plurality of weights not masked by the decoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, the output image, and the compressed representation; pruning one or more weights of the updated second plurality of weights not masked by the decoding mask to obtain a binary pruning mask indicating which of the second plurality of weights to prune; the pruning includes: determining a pruning loss for each block into which each mask in the decoding mask is divided; sorting the plurality of blocks in ascending order based on the determined pruning loss for each block of the plurality of blocks; and setting two or more weights of the second plurality of weights corresponding to the plurality of blocks from top to bottom in the sorted plurality of blocks until a termination criterion is reached; updating at least one weight of the second plurality of weights not masked by the decoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and Update the decoding mask based on the obtained binary pruning mask.

3. The method according to claim 1, wherein, each item in each of the plurality of blocks has the same binary value.

4. The method according to claim 2, wherein, each of the encoding mask and the decoding mask has randomly distributed binary values.

5. The method according to claim 2, wherein, each of the encoding mask and the decoding mask is divided into columns, rows or channels, and each item in each of the columns, rows or channels of the columns, rows or channels has the same binary value.

6. A multi-bitrate neural image processing device, the device comprising: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code, the program code comprising: a first selection code configured to cause the at least one processor to select an encoding mask based on hyperparameters; a first execution code configured to cause the at least one processor to perform a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain a first masked weight; a first encoding code configured to cause the at least one processor to encode an input image using the first masked weight to obtain an encoded representation; and a second encoding code configured to cause the at least one processor to encode the obtained encoded representation to obtain a compressed representation; wherein the first neural network is trained by: updating one or more weights in the first plurality of weights that are not masked by the encoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, the output image, and the compressed representation; pruning one or more weights in the updated first plurality of weights that are not masked by the encoding mask to obtain a binary pruning mask indicating which of the first plurality of weights to prune; the pruning includes: determining a pruning loss for each of a plurality of blocks into which each of the masks in the encoding mask is divided; arranging the plurality of blocks in ascending order based on the determined pruning loss for each of the plurality of blocks; and setting two or more weights in the first plurality of weights corresponding to the plurality of blocks from top to bottom in the arranged plurality of blocks until a termination criterion is reached; updating at least one weight in the first plurality of weights that is not masked by the encoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and updating the encoding mask based on the obtained binary pruning mask.

7. The device according to claim 6, wherein, the program code further includes: a first decoding code configured to cause the at least one processor to decode the obtained compressed representation to obtain a recovered representation; a second selection code configured to cause the at least one processor to select a decoding mask based on the hyperparameters; A second execution code, configured to cause the at least one processor to perform a convolution of a second plurality of weights of a second neural network with a selected decoding mask to obtain second masked weights; and A second decoding code, configured to cause the at least one processor to decode the obtained recovered representation using the second masked weights to reconstruct an output image; wherein, the second neural network is trained by: Updating one or more weights in the second plurality of weights that are not masked by the decoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, the output image, and the compressed representation; Pruning one or more weights in the updated second plurality of weights that are not masked by the decoding mask to obtain a binary pruning mask indicating which of the second plurality of weights to prune; the pruning includes: determining a pruning loss for each of a plurality of blocks into which each mask in the decoding mask is divided; arranging the plurality of blocks in ascending order based on the determined pruning loss for each of the plurality of blocks; and setting two or more weights in the second plurality of weights corresponding to the plurality of blocks from top to bottom in the arranged plurality of blocks until a termination criterion is reached; Updating at least one weight in the second plurality of weights that is not masked by the decoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and Updating the decoding mask based on the obtained binary pruning mask.

8. The apparatus according to claim 7, wherein, Each item in each of the plurality of blocks has the same binary value.

9. The apparatus according to claim 7, wherein, Each of the encoding mask and the decoding mask has randomly distributed binary values.

10. The apparatus according to claim 7, wherein, Each of the encoding mask and the decoding mask is divided into columns, rows, or channels, and Each item in each of the columns, rows, or channels of the columns, rows, or channels has the same binary value.

11. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor for multi-bitrate neural image compression, cause the at least one processor to: Select an encoding mask based on hyperparameters; Perform a convolution of a first plurality of weights of a first neural network with the selected encoding mask to obtain first masked weights; Encode an input image using the first masked weights to obtain an encoded representation; and Encode the obtained encoded representation to obtain a compressed representation; wherein, the first neural network is trained by: Updating one or more weights in the first plurality of weights that are not masked by the encoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, the output image, and the compressed representation; Prune one or more weights in the updated first plurality of weights that are not masked by the encoding mask to obtain a binary pruning mask indicating which of the first plurality of weights are pruned; the pruning includes: determining a pruning loss for each block into which each mask in the encoding mask is divided; arranging the plurality of blocks in ascending order based on the determined pruning loss for each block of the plurality of blocks; and setting two or more weights in the first plurality of weights corresponding to the plurality of blocks from top to bottom in the arranged plurality of blocks until a termination criterion is reached; Update at least one weight in the first plurality of weights that is not masked by the encoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and Update the encoding mask based on the obtained binary pruning mask.

12. The non-transitory computer-readable medium according to claim 11, wherein, when the instructions are executed by the at least one processor, the instructions further cause the at least one processor to: Decode the obtained compressed representation to obtain a recovered representation; Select a decoding mask based on the hyperparameter; Perform a convolution of a second plurality of weights of a second neural network with the selected decoding mask to obtain second masked weights; and and Decode the obtained recovered representation using the second masked weights to reconstruct an output image; wherein the second neural network is trained by: Updating one or more weights in the second plurality of weights that are not masked by the decoding mask to minimize a bitrate-distortion loss, the bitrate-distortion loss being determined based on the input image, the output image, and the compressed representation; Prune one or more weights in the updated second plurality of weights that are not masked by the decoding mask to obtain a binary pruning mask indicating which of the second plurality of weights are pruned; the pruning includes: determining a pruning loss for each block into which each mask in the decoding mask is divided; arranging the plurality of blocks in ascending order based on the determined pruning loss for each block of the plurality of blocks; and setting two or more weights in the second plurality of weights corresponding to the plurality of blocks from top to bottom in the arranged plurality of blocks until a termination criterion is reached; Update at least one weight in the second plurality of weights that is not masked by the decoding mask and the obtained binary pruning mask to minimize the bitrate-distortion loss; and Update the decoding mask based on the obtained binary pruning mask.

13. The non-transitory computer-readable medium according to claim 12, wherein, Each item in each block of the plurality of blocks has the same binary value.

14. The non-transitory computer-readable medium according to claim 12, wherein, Each of the encoding mask and the decoding mask is divided into columns, rows, or channels, and Each item in each column, row, or channel of the columns, rows, or channels has the same binary value.

Citation Information

Patent Citations

  • Method and apparatus for compressing and accelerating multi-rate neural image compression model through microstructure nested mask and weight homogenization

    CN114556911A